docs(config): explicit ollama_num_ctx is never capped by context_length
agent/agent_init.py::_configure_ollama_num_ctx caps only the auto-detected value (the cap is skipped when an explicit override is set). The salvaged example wording said context_length caps the explicit value too, which would send readers chasing a cap that does not apply.
This commit is contained in:
@@ -113,8 +113,9 @@ model:
|
||||
#
|
||||
# ollama_num_ctx: Ollama only — the num_ctx sent on every chat request.
|
||||
# Hermes auto-detects the model's window and sends it (Ollama otherwise
|
||||
# defaults to 2048); set this to cap VRAM use. context_length, if set,
|
||||
# caps it further.
|
||||
# defaults to 2048); context_length, if set, caps the detected value.
|
||||
# An explicit ollama_num_ctx is sent as-is (never capped) — set it to
|
||||
# pin VRAM use.
|
||||
#
|
||||
# ollama_num_ctx: 32768
|
||||
#
|
||||
|
||||
Reference in New Issue
Block a user