docs(config): explicit ollama_num_ctx is never capped by context_length

agent/agent_init.py::_configure_ollama_num_ctx caps only the auto-detected
value (the cap is skipped when an explicit override is set). The salvaged
example wording said context_length caps the explicit value too, which
would send readers chasing a cap that does not apply.
This commit is contained in:
teknium1
2026-09-15 11:39:13 -07:00
committed by Teknium
parent 6c9a665de2
commit 462ca7c112

View File

@@ -113,8 +113,9 @@ model:
#
# ollama_num_ctx: Ollama only — the num_ctx sent on every chat request.
# Hermes auto-detects the model's window and sends it (Ollama otherwise
# defaults to 2048); set this to cap VRAM use. context_length, if set,
# caps it further.
# defaults to 2048); context_length, if set, caps the detected value.
# An explicit ollama_num_ctx is sent as-is (never capped) — set it to
# pin VRAM use.
#
# ollama_num_ctx: 32768
#