docs(vision): document vision.embed_target_bytes and vision.max_calls_per_image
User-visible knobs need a home: the Vision feature page explains why native embeds ride the session and what each key does (subagent-only default cap, clamp range), the configuration reference points at it next to auxiliary.vision so the two sections are not confused, cli-config.yaml.example carries the commented block, and tools/AGENTS.md names vision_tools_history_budget.py as the single owner of embed-cost policy.
This commit is contained in:
@@ -831,6 +831,19 @@ prompt_caching:
|
||||
# uses "google/gemini-3-flash-preview" and Nous uses "gemini-3-flash".
|
||||
# Other providers pick a sensible default automatically.
|
||||
#
|
||||
# Native vision embeds (vision_analyze / browser screenshots when the MAIN model
|
||||
# is vision-capable) ride conversation history and are re-sent on every later
|
||||
# API call, so both knobs below bound recurring cost:
|
||||
#
|
||||
# vision:
|
||||
# # Byte budget for one embedded image (clamped 64 KiB..4 MiB). Raise it for
|
||||
# # dense phone screenshots of tables that read as "unreadable" at 256 KB.
|
||||
# embed_target_bytes: 262144
|
||||
# # How often vision_analyze may embed the SAME image (region crops included)
|
||||
# # per session. Unset = 3 inside delegated subagents, unlimited for the main
|
||||
# # agent; a number applies everywhere; 0 = unlimited.
|
||||
# max_calls_per_image: 3
|
||||
#
|
||||
# auxiliary:
|
||||
# # Image analysis: vision_analyze tool + browser screenshots
|
||||
# vision:
|
||||
|
||||
@@ -84,6 +84,14 @@ client (`mcp_tool_*.py`: config, discovery, transport, registration, content, er
|
||||
never an `elif` on a backend name (root shape rules). Remote-backend file visibility problems are
|
||||
fixed at the mount, not by adding a tool.
|
||||
|
||||
**Native vision embeds are history, not one-shot payloads.** `vision_tools.py::_vision_analyze_native`
|
||||
(and the browser screenshot twins in `browser_tool_vision.py` / `browser_use_cli.py`) bake the image
|
||||
into a tool result that is re-sent on every later API call. Size and repeat policy live in
|
||||
`vision_tools_history_budget.py` (config section `vision`: `embed_target_bytes`, `max_calls_per_image`);
|
||||
the repeat counter is keyed on (session id, resolved source) so region crops share their file's count,
|
||||
and its default cap applies only inside `agent.delegation_context.is_delegated_child_process_context()`.
|
||||
Put new embed-cost rules there, never a second counter in a tool.
|
||||
|
||||
**Every spawn goes through one env builder.** `environments/local.py::build_subprocess_env` (+
|
||||
`hermes_constants.apply_subprocess_home_env`, `env_passthrough.py::resolve_passthrough_value`) is
|
||||
how a terminal, `execute_code`, background process, delegation child, ACP or MCP stdio child gets
|
||||
|
||||
@@ -1574,6 +1574,10 @@ Each entry supports the same three knobs as any auxiliary task config:
|
||||
|
||||
`fallback_chain` is available on any auxiliary task — `compression`, `vision`, `approval`, `skills_hub`, `mcp`, etc.
|
||||
|
||||
### Native vision embed budgets (top-level `vision:`)
|
||||
|
||||
Separate from `auxiliary.vision` (which picks the describer model): when the *main* model is vision-capable, `vision_analyze` and browser screenshots embed real pixels into tool results that are re-sent every later turn. `vision.embed_target_bytes` (default `262144`, clamped 64 KiB..4 MiB) sizes one embed; `vision.max_calls_per_image` caps how often the same image may be embedded per session (unset = 3 inside delegated subagents, unlimited for the main agent; `0` = unlimited). See [Vision → Native embeds ride the session](/user-guide/features/vision#native-embeds-ride-the-session-visionembed_target_bytes-and-visionmax_calls_per_image).
|
||||
|
||||
### Limiting auxiliary concurrency
|
||||
|
||||
`max_concurrency` caps in-flight LLM calls for auxiliary tasks such as `compression` and `title_generation` across the whole process. `auxiliary.vision.max_concurrency` is excluded: it already controls only vision's CPU-bound image encode/resize workers, not LLM requests. This is most useful when:
|
||||
|
||||
@@ -212,3 +212,16 @@ Which auxiliary model handles the text-description path is configurable under `a
|
||||
The `vision_analyze` tool itself follows the same routing. When the active main model is vision-capable **and** its provider supports image content inside tool results (currently the Anthropic, OpenAI, Azure-OpenAI, and Gemini 3.x stacks), `vision_analyze` short-circuits the auxiliary describer and returns the raw image pixels as a multimodal tool-result envelope. The main model sees the image natively on its next turn — no aux call, no text-summary information loss, no extra latency.
|
||||
|
||||
For text-only main models (or providers whose tool-result channel doesn't carry images), `vision_analyze` falls back to the legacy path: it asks the configured auxiliary vision model to describe the image and returns the description as plain text. Either way the calling tool signature is the same — the tool decides which path to take at runtime based on the active model.
|
||||
|
||||
### Native embeds ride the session: `vision.embed_target_bytes` and `vision.max_calls_per_image`
|
||||
|
||||
A native `vision_analyze` result bakes the image into the tool result, and that result is re-sent on every later API call of the session. Two `config.yaml` keys bound the recurring cost:
|
||||
|
||||
```yaml
|
||||
vision:
|
||||
embed_target_bytes: 262144 # per-embed byte budget; clamped 64 KiB..4 MiB (default 256 KB)
|
||||
max_calls_per_image: 3 # unset = 3 inside delegated subagents, unlimited for the main agent
|
||||
```
|
||||
|
||||
- **`embed_target_bytes`** — images above the budget (or wider than 1568 px) are downscaled to a JPEG that fits. 256 KB keeps ordinary screenshots cheap; dense phone screenshots of tables can come out unreadable at that size, so raise it (say `1048576`) when the model keeps calling figures "unreadable". Browser screenshots delivered natively use the same budget.
|
||||
- **`max_calls_per_image`** — how often the *same* image (region crops of it included; local paths compare by resolved path) may be embedded per session. Once the cap is hit the tool returns `"vision_analyze refused: this image has already been loaded into context N time(s) …"` instead of another embed, so the model answers from what it already sees. Left unset, only delegated `delegate_task` subagents are capped (at 3): they run unattended and cannot be steered mid-loop from the CLI, and a re-load loop there once burned 158 calls on five files. Set a number to cap every session, or `0` for unlimited everywhere.
|
||||
|
||||
Reference in New Issue
Block a user