Files
hermes-agent/agent
teknium1 55db0a9847 fix: local sub-64K context refusal stops assuming Ollama
When an OpenAI-compatible local server (llama.cpp, vLLM, ...) serves a
window below MINIMUM_CONTEXT_LENGTH, the startup refusal, the CLI banner
warning and the num_ctx log line were worded for Ollama (/api/show,
model.context_length advice that cannot raise a served window). Name the
served window and the server-agnostic remedies: start the server with a
>=64K context or set model.ollama_num_ctx (honoured on every local
endpoint) to the window it really serves. Hosted routes keep the
model.context_length advice. The floor itself is unchanged.

Part of #87075
2026-09-19 10:21:28 -07:00
..
…
…