docs(config): align cli-config.yaml.example with the code defaults and document six missing keys

Two default claims in the example contradict the code:

- agent.verify_on_stop: the example says the default is "auto"; the default
  in config_defaults.py is False, agent/verification_stop.py treats "auto"
  as the explicit opt-in for the surface-aware mode, and the website
  already says off. The comment and the sample value now say so.
- display.show_reasoning: the example marks false as the default and ships
  show_reasoning: false; DEFAULT_CONFIG and cli.py's own defaults are True
  ("Default ON ... with this off the user stares at a spinner"). Anyone who
  copied the example silently turned reasoning display off. The marker and
  the sample value now match the code.

Six keys the runtime reads but the example never mentioned:

- security.protected_instruction_files / protected_instruction_extra_patterns
  (tools/file_tools.py) — a write guard for AGENTS.md/CLAUDE.md/SOUL.md and
  friends with no documentation anywhere.
- browser.restrict_evaluate / browser.allow_unsafe_evaluate — both in
  DEFAULT_CONFIG, only the former mentioned in the website browser page.
- tts.delivery_profiles.<platform>.{max_file_bytes,safety_ratio}
  (tools/tts_tool.py) — per-platform audio upload limits over the built-in
  discord/telegram/default table.
- gateway restart_after_turn_timeout (config_defaults.py, 1800) — the
  in-band restart knob its own sibling comment tells readers to prefer.
- model.ollama_num_ctx (agent/agent_init.py) — the documented-in-code VRAM
  cap for Ollama's num_ctx.

display.streaming was left alone on purpose: the CLI builds its config from
its own defaults (streaming: True) rather than DEFAULT_CONFIG (False), so
the example's "(default)" marker is correct for the surface it documents.
This commit is contained in:
John Paul Soliva
2026-09-02 22:18:53 +09:00
committed by Teknium
parent 1cfa892db1
commit f20acf5d9f

View File

@@ -111,6 +111,13 @@ model:
#
# context_length: 131072
#
# ollama_num_ctx: Ollama only — the num_ctx sent on every chat request.
# Hermes auto-detects the model's window and sends it (Ollama otherwise
# defaults to 2048); set this to cap VRAM use. context_length, if set,
# caps it further.
#
# ollama_num_ctx: 32768
#
# Output-token limits are provider-owned, not user configuration. Native
# protocols requiring a limit receive an internal value from Hermes.
@@ -514,6 +521,11 @@ terminal:
# approval:
# transport: builtin # Or an explicitly enabled plugin transport name
# transport_fallback: deny # Set builtin to opt into fallback on transport failure
# # Refuse agent writes to instruction files that steer the agent itself
# # (AGENTS.md, CLAUDE.md, SOUL.md, ...). Set false to allow them; add
# # fnmatch patterns on the basename to protect more files.
# protected_instruction_files: true
# protected_instruction_extra_patterns: []
# =============================================================================
# Browser Tool Configuration
@@ -534,6 +546,14 @@ browser:
# (same per-page budget as web_extract).
# snapshot_threshold: 15000
# JavaScript evaluation guardrails (both default false). restrict_evaluate
# opts into a denylist that blocks sensitive primitives (cookies, storage,
# clipboard, network, form values) in browser_console(expression=...) — use
# it when the agent drives a logged-in profile. allow_unsafe_evaluate is the
# legacy override that switches the denylist back off even when restrict is on.
# restrict_evaluate: false
# allow_unsafe_evaluate: false
# =============================================================================
# Tool Loop Guardrails
# =============================================================================
@@ -1130,6 +1150,14 @@ agent:
# restart_drain_timeout like before.
# cron_drain_timeout: 30
# In-band restart wait (seconds) for active turns to finish BEFORE stop()
# begins. /restart and SIGUSR1 refuse new work, then wait up to this cap for
# in-flight agent, cron and API runs to complete so the requesting turn is
# not cut off by restart_drain_timeout. 0 = enter stop()/drain immediately.
# The default (30 min) is a safety valve for wedged agents, not a target
# latency; raise it for long unattended turns. Env: HERMES_RESTART_AFTER_TURN_TIMEOUT.
# restart_after_turn_timeout: 1800
# Upper bound (seconds) a submitted prompt waits for the deferred agent
# build (MCP discovery, model metadata, skills scan) before failing with a
# visible error. The wait is patient — the message is delivered as soon as
@@ -1146,12 +1174,13 @@ agent:
# api_max_retries: 3
# After the agent edits code without fresh passing verification, nudge it to
# verify before finishing. The default "auto" enables it on interactive
# coding surfaces (CLI, TUI, desktop) and programmatic callers, and disables
# it on conversational messaging surfaces (Telegram, Discord, etc.) where the
# verification summary would reach a human as chat noise. Set true or false to
# force it on or off; the HERMES_VERIFY_ON_STOP env var (1/0) takes precedence.
# verify_on_stop: auto
# verify before finishing. Off by default (false). Set true to force it on
# everywhere, or "auto" for the surface-aware mode: on for interactive
# coding surfaces (CLI, TUI, desktop) and programmatic callers, off on
# conversational messaging surfaces (Telegram, Discord, etc.) where the
# verification summary would reach a human as chat noise. The
# HERMES_VERIFY_ON_STOP env var (1/0) takes precedence.
# verify_on_stop: false
# Standing operator instructions for the coding posture (when Hermes is in a
# code workspace). Appended to the coding brief as an extra system block, so
@@ -1481,6 +1510,13 @@ platform_toolsets:
# tts:
# provider: "gemini"
# speed: 1.0 # global speed multiplier (provider-specific overrides this)
# # Per-platform audio upload limits. Built-in defaults: discord 10 MiB,
# # telegram 50 MiB, everything else 10 MiB, each with safety_ratio 0.85
# # (the fraction of max_file_bytes actually targeted). Override per platform:
# delivery_profiles:
# telegram:
# max_file_bytes: 52428800
# safety_ratio: 0.85
# gemini:
# model: "gemini-3.1-flash-tts-preview"
# voice: "Kore"
@@ -1793,9 +1829,9 @@ display:
# Show model reasoning/thinking before each response.
# When enabled, a dim box shows the model's thought process above the response.
# Toggle at runtime with /reasoning show or /reasoning hide.
# true: Show the reasoning box
# false: Hide reasoning (default)
show_reasoning: false
# true: Show the reasoning box (default)
# false: Hide reasoning
show_reasoning: true
# Stream tokens to the terminal as they arrive instead of waiting for the
# full response. The response box opens on first token and text appears