~/.cache/hermes-pytest failed seven files: a dot-dir ancestor made the hidden-dir search tests
see every fixture as hidden, and the longer root pushed the PulseAudio/voice AF_UNIX test
sockets past sun_path. /var/tmp is the FHS disk-backed temp root (never tmpfs), non-hidden,
and shorter than the old /tmp root. test_tool_result_storage asserts STORAGE_DIR instead of a
literal, and the zh-Hans bot-mode mirror follows the English code block it must copy.
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
Hermes and everything it launches (browser profiles, PTY probes, skill scripts,
tempfile defaults in delegated code) wrote to the system temp dir, which is a
RAM-backed tmpfs on most Linux hosts and containers and fills under agent load.
- hermes_constants.get_scratch_dir(): HERMES_HOME/cache/scratch (0700), entries
older than 72h pruned once per process / once per hour across processes.
- apply_scratch_tmp_env(env) / export_scratch_tmp_env(): TMPDIR/TMP/TEMP point at
the scratch dir when the user or OS has not set them; a value Hermes itself
exported (== HERMES_SCRATCH_DIR) is re-derived for a re-homed process or a
child served under another profile, so profiles never share scratch.
- hermes_bootstrap runs the export on import (every entry point); hermes_cli.main
re-runs it after --profile resolution; the subprocess HOME contract
(apply_subprocess_home_env) and the routed-home rewrites in code_execution_env
and served_profile_child_env apply it to child envs.
- The runtime-environment prompt block names the scratch dir so the model stops
reaching for the system temp dir by reflex; hermes doctor reports the dir, its
size and whether a user TMPDIR overrides it.
The client-direct STT path in the Desktop (voice-client-direct.ts) issued a
bare fetch to the provider with no AbortSignal, so a slow or wedged
transcription endpoint left the microphone stuck on "transcribing" forever;
the gateway's own transcription client already has a deadline.
- tools/voice_client_config.py::_resolve_stt_client_config adds `timeout_s`
(from `stt.openai.timeout`, default 60; groq/deepinfra riders inherit) to
the direct STT config the gateway hands the Desktop.
- voice-client-direct.ts routes all three direct STT fetches
(openai-multipart, xai-stt, elevenlabs-stt) through sttFetch, which aborts
at that deadline and surfaces "Transcription timed out after Ns".
- Docs: desktop.md dictation paragraph notes the shared budget.
Part of #112939
What: the `initOnSessionStart` setup-schema description and the Honcho docs page
now state that in `recallMode: tools` the eager init runs synchronously during
agent construction, that it should stay false for Desktop or a local Honcho that
may be down, and that `timeout` in honcho.json caps each SDK call (both keys
added to the Full Config Reference table).
Why: #51492 reports Desktop `session.resume` / `prompt.submit` timeouts when a
local Honcho is down with `initOnSessionStart: true`. The synchronous eager path
is intentional design (#51562 was closed for that reason:
"ready-before-first-tool-call semantics"), so the fix for users is knowing the
trade-off and the two knobs that bound it.
What: a "Fast tiers behind a gateway or proxy" subsection under Fast Mode in the
configuration guide plus a commented example in cli-config.yaml.example showing
`providers.<name>.extra_body: {service_tier: priority}`.
Why: agent.service_tier / `/fast` deliberately reach only first-party billing
endpoints (hermes_cli/models.py::_fast_mode_route_supported, c7e2e0b779), so
gateway and proxy users asked how to get a priority tier (#78097). The
per-provider extra_body is merged into every chat-completions request for that
endpoint (agent/agent_init.py::_merge_custom_provider_extra_body →
agent/fast_mode.py::effective_request_overrides), which already gives them a
supported path; it was just undocumented next to Fast Mode.
What: a new opt-in gateway runtime-footer field `served_model` rendered as
`alias → served`. It is populated from the `x-litellm-model-id` response header
(fallback `x-litellm-model-api-base`) that routing proxies send on every chat
completion, captured through an httpx response hook installed on the agent's
OpenAI client (agent/served_model.py, wired in create_openai_client), and from
Hermes' own provider fallback (primary runtime model → active model) when no
header is present. The turn result carries `requested_model` / `served_model`;
gateway/run_turn.py passes them to the footer. Off unless listed in
`display.runtime_footer.fields`; the default field set renders exactly as before.
Why: behind a routing proxy (or during a silent Hermes fallback) every reply
shows the configured alias, so operators cannot see which deployment actually
served a request (#54864). The SDK's parsed objects drop response headers, so
the capture has to sit on the transport.
The provider error card's "Switch provider" button navigated to
Settings → Models, which only changes the default provider/model for NEW
sessions — the failed chat kept its broken provider, so the user had to
find the composer pill on their own to actually recover the turn.
The button now calls requestModelMenuToggle(), the same bus request the
`composer.modelPicker` hotkey uses: it opens the composer pill's live
model menu (pane under the pointer, else the active composer), whose
picks go through model.switch on THIS session. When no chat surface is
on screen (requestModelMenuToggle returns false) it falls back to the
Settings → Models deep link as before. No new RPC.
Also refreshes the ErrorRecoveryPlan.switchProvider doc comment and the
Desktop user-guide bullet describing the button.
Tests: two invariant vitests on the error card (menu opened, no
navigation / menu unavailable → Settings deep link); both fail on base
where the click always navigates.
Part of #95066