docs(vision): document vision.embed_target_bytes and vision.max_calls_per_image

User-visible knobs need a home: the Vision feature page explains why native embeds ride
the session and what each key does (subagent-only default cap, clamp range), the
configuration reference points at it next to auxiliary.vision so the two sections are
not confused, cli-config.yaml.example carries the commented block, and tools/AGENTS.md
names vision_tools_history_budget.py as the single owner of embed-cost policy.
This commit is contained in:
teknium1
2026-09-16 22:10:53 -07:00
committed by Teknium
parent abc56351aa
commit ed7eb5685a
4 changed files with 38 additions and 0 deletions

View File

@@ -831,6 +831,19 @@ prompt_caching:
# uses "google/gemini-3-flash-preview" and Nous uses "gemini-3-flash".
# Other providers pick a sensible default automatically.
#
# Native vision embeds (vision_analyze / browser screenshots when the MAIN model
# is vision-capable) ride conversation history and are re-sent on every later
# API call, so both knobs below bound recurring cost:
#
# vision:
# # Byte budget for one embedded image (clamped 64 KiB..4 MiB). Raise it for
# # dense phone screenshots of tables that read as "unreadable" at 256 KB.
# embed_target_bytes: 262144
# # How often vision_analyze may embed the SAME image (region crops included)
# # per session. Unset = 3 inside delegated subagents, unlimited for the main
# # agent; a number applies everywhere; 0 = unlimited.
# max_calls_per_image: 3
#
# auxiliary:
# # Image analysis: vision_analyze tool + browser screenshots
# vision:

View File

@@ -84,6 +84,14 @@ client (`mcp_tool_*.py`: config, discovery, transport, registration, content, er
never an `elif` on a backend name (root shape rules). Remote-backend file visibility problems are
fixed at the mount, not by adding a tool.
**Native vision embeds are history, not one-shot payloads.** `vision_tools.py::_vision_analyze_native`
(and the browser screenshot twins in `browser_tool_vision.py` / `browser_use_cli.py`) bake the image
into a tool result that is re-sent on every later API call. Size and repeat policy live in
`vision_tools_history_budget.py` (config section `vision`: `embed_target_bytes`, `max_calls_per_image`);
the repeat counter is keyed on (session id, resolved source) so region crops share their file's count,
and its default cap applies only inside `agent.delegation_context.is_delegated_child_process_context()`.
Put new embed-cost rules there, never a second counter in a tool.
**Every spawn goes through one env builder.** `environments/local.py::build_subprocess_env` (+
`hermes_constants.apply_subprocess_home_env`, `env_passthrough.py::resolve_passthrough_value`) is
how a terminal, `execute_code`, background process, delegation child, ACP or MCP stdio child gets

View File

@@ -1574,6 +1574,10 @@ Each entry supports the same three knobs as any auxiliary task config:
`fallback_chain` is available on any auxiliary task — `compression`, `vision`, `approval`, `skills_hub`, `mcp`, etc.
### Native vision embed budgets (top-level `vision:`)
Separate from `auxiliary.vision` (which picks the describer model): when the *main* model is vision-capable, `vision_analyze` and browser screenshots embed real pixels into tool results that are re-sent every later turn. `vision.embed_target_bytes` (default `262144`, clamped 64 KiB..4 MiB) sizes one embed; `vision.max_calls_per_image` caps how often the same image may be embedded per session (unset = 3 inside delegated subagents, unlimited for the main agent; `0` = unlimited). See [Vision → Native embeds ride the session](/user-guide/features/vision#native-embeds-ride-the-session-visionembed_target_bytes-and-visionmax_calls_per_image).
### Limiting auxiliary concurrency
`max_concurrency` caps in-flight LLM calls for auxiliary tasks such as `compression` and `title_generation` across the whole process. `auxiliary.vision.max_concurrency` is excluded: it already controls only vision's CPU-bound image encode/resize workers, not LLM requests. This is most useful when:

View File

@@ -212,3 +212,16 @@ Which auxiliary model handles the text-description path is configurable under `a
The `vision_analyze` tool itself follows the same routing. When the active main model is vision-capable **and** its provider supports image content inside tool results (currently the Anthropic, OpenAI, Azure-OpenAI, and Gemini 3.x stacks), `vision_analyze` short-circuits the auxiliary describer and returns the raw image pixels as a multimodal tool-result envelope. The main model sees the image natively on its next turn — no aux call, no text-summary information loss, no extra latency.
For text-only main models (or providers whose tool-result channel doesn't carry images), `vision_analyze` falls back to the legacy path: it asks the configured auxiliary vision model to describe the image and returns the description as plain text. Either way the calling tool signature is the same — the tool decides which path to take at runtime based on the active model.
### Native embeds ride the session: `vision.embed_target_bytes` and `vision.max_calls_per_image`
A native `vision_analyze` result bakes the image into the tool result, and that result is re-sent on every later API call of the session. Two `config.yaml` keys bound the recurring cost:
```yaml
vision:
embed_target_bytes: 262144 # per-embed byte budget; clamped 64 KiB..4 MiB (default 256 KB)
max_calls_per_image: 3 # unset = 3 inside delegated subagents, unlimited for the main agent
```
- **`embed_target_bytes`** — images above the budget (or wider than 1568 px) are downscaled to a JPEG that fits. 256 KB keeps ordinary screenshots cheap; dense phone screenshots of tables can come out unreadable at that size, so raise it (say `1048576`) when the model keeps calling figures "unreadable". Browser screenshots delivered natively use the same budget.
- **`max_calls_per_image`** — how often the *same* image (region crops of it included; local paths compare by resolved path) may be embedded per session. Once the cap is hit the tool returns `"vision_analyze refused: this image has already been loaded into context N time(s) …"` instead of another embed, so the model answers from what it already sees. Left unset, only delegated `delegate_task` subagents are capped (at 3): they run unattended and cannot be steered mid-loop from the CLI, and a re-load loop there once burned 158 calls on five files. Set a number to cap every session, or `0` for unlimited everywhere.