diff --git a/cli-config.yaml.example b/cli-config.yaml.example index f48c109c77..ba88f515f3 100644 --- a/cli-config.yaml.example +++ b/cli-config.yaml.example @@ -831,6 +831,19 @@ prompt_caching: # uses "google/gemini-3-flash-preview" and Nous uses "gemini-3-flash". # Other providers pick a sensible default automatically. # +# Native vision embeds (vision_analyze / browser screenshots when the MAIN model +# is vision-capable) ride conversation history and are re-sent on every later +# API call, so both knobs below bound recurring cost: +# +# vision: +# # Byte budget for one embedded image (clamped 64 KiB..4 MiB). Raise it for +# # dense phone screenshots of tables that read as "unreadable" at 256 KB. +# embed_target_bytes: 262144 +# # How often vision_analyze may embed the SAME image (region crops included) +# # per session. Unset = 3 inside delegated subagents, unlimited for the main +# # agent; a number applies everywhere; 0 = unlimited. +# max_calls_per_image: 3 +# # auxiliary: # # Image analysis: vision_analyze tool + browser screenshots # vision: diff --git a/tools/AGENTS.md b/tools/AGENTS.md index 90f7972bcd..75a973a354 100644 --- a/tools/AGENTS.md +++ b/tools/AGENTS.md @@ -84,6 +84,14 @@ client (`mcp_tool_*.py`: config, discovery, transport, registration, content, er never an `elif` on a backend name (root shape rules). Remote-backend file visibility problems are fixed at the mount, not by adding a tool. +**Native vision embeds are history, not one-shot payloads.** `vision_tools.py::_vision_analyze_native` +(and the browser screenshot twins in `browser_tool_vision.py` / `browser_use_cli.py`) bake the image +into a tool result that is re-sent on every later API call. Size and repeat policy live in +`vision_tools_history_budget.py` (config section `vision`: `embed_target_bytes`, `max_calls_per_image`); +the repeat counter is keyed on (session id, resolved source) so region crops share their file's count, +and its default cap applies only inside `agent.delegation_context.is_delegated_child_process_context()`. +Put new embed-cost rules there, never a second counter in a tool. + **Every spawn goes through one env builder.** `environments/local.py::build_subprocess_env` (+ `hermes_constants.apply_subprocess_home_env`, `env_passthrough.py::resolve_passthrough_value`) is how a terminal, `execute_code`, background process, delegation child, ACP or MCP stdio child gets diff --git a/website/docs/user-guide/configuration.md b/website/docs/user-guide/configuration.md index f527592035..5eeabdfc33 100644 --- a/website/docs/user-guide/configuration.md +++ b/website/docs/user-guide/configuration.md @@ -1574,6 +1574,10 @@ Each entry supports the same three knobs as any auxiliary task config: `fallback_chain` is available on any auxiliary task — `compression`, `vision`, `approval`, `skills_hub`, `mcp`, etc. +### Native vision embed budgets (top-level `vision:`) + +Separate from `auxiliary.vision` (which picks the describer model): when the *main* model is vision-capable, `vision_analyze` and browser screenshots embed real pixels into tool results that are re-sent every later turn. `vision.embed_target_bytes` (default `262144`, clamped 64 KiB..4 MiB) sizes one embed; `vision.max_calls_per_image` caps how often the same image may be embedded per session (unset = 3 inside delegated subagents, unlimited for the main agent; `0` = unlimited). See [Vision → Native embeds ride the session](/user-guide/features/vision#native-embeds-ride-the-session-visionembed_target_bytes-and-visionmax_calls_per_image). + ### Limiting auxiliary concurrency `max_concurrency` caps in-flight LLM calls for auxiliary tasks such as `compression` and `title_generation` across the whole process. `auxiliary.vision.max_concurrency` is excluded: it already controls only vision's CPU-bound image encode/resize workers, not LLM requests. This is most useful when: diff --git a/website/docs/user-guide/features/vision.md b/website/docs/user-guide/features/vision.md index 628c7ab69e..f8b4c38d3a 100644 --- a/website/docs/user-guide/features/vision.md +++ b/website/docs/user-guide/features/vision.md @@ -212,3 +212,16 @@ Which auxiliary model handles the text-description path is configurable under `a The `vision_analyze` tool itself follows the same routing. When the active main model is vision-capable **and** its provider supports image content inside tool results (currently the Anthropic, OpenAI, Azure-OpenAI, and Gemini 3.x stacks), `vision_analyze` short-circuits the auxiliary describer and returns the raw image pixels as a multimodal tool-result envelope. The main model sees the image natively on its next turn — no aux call, no text-summary information loss, no extra latency. For text-only main models (or providers whose tool-result channel doesn't carry images), `vision_analyze` falls back to the legacy path: it asks the configured auxiliary vision model to describe the image and returns the description as plain text. Either way the calling tool signature is the same — the tool decides which path to take at runtime based on the active model. + +### Native embeds ride the session: `vision.embed_target_bytes` and `vision.max_calls_per_image` + +A native `vision_analyze` result bakes the image into the tool result, and that result is re-sent on every later API call of the session. Two `config.yaml` keys bound the recurring cost: + +```yaml +vision: + embed_target_bytes: 262144 # per-embed byte budget; clamped 64 KiB..4 MiB (default 256 KB) + max_calls_per_image: 3 # unset = 3 inside delegated subagents, unlimited for the main agent +``` + +- **`embed_target_bytes`** — images above the budget (or wider than 1568 px) are downscaled to a JPEG that fits. 256 KB keeps ordinary screenshots cheap; dense phone screenshots of tables can come out unreadable at that size, so raise it (say `1048576`) when the model keeps calling figures "unreadable". Browser screenshots delivered natively use the same budget. +- **`max_calls_per_image`** — how often the *same* image (region crops of it included; local paths compare by resolved path) may be embedded per session. Once the cap is hit the tool returns `"vision_analyze refused: this image has already been loaded into context N time(s) …"` instead of another embed, so the model answers from what it already sees. Left unset, only delegated `delegate_task` subagents are capped (at 3): they run unattended and cannot be steered mid-loop from the CLI, and a re-load loop there once burned 158 calls on five files. Set a number to cap every session, or `0` for unlimited everywhere.