Under a container terminal backend (docker/modal/...), terminal.cwd names a path
inside the sandbox (/workspace). _completion_cwd host-validated it with
os.path.isdir, failed, and fell back to the gateway's own cwd ($HOME for the
desktop). Every host file then looked "inside the workspace", so file.attach
never staged it into the bind-mounted attachments/ dir and the agent was handed
a host path that does not exist in the container.
- _completion_cwd keeps an absolute container cwd unvalidated when the bound
backend is non-local (matching _terminal_task_cwd), including a named
profile's own container terminal.cwd and a config-only (unbridged) backend.
- _session_is_local_backend reads the effective backend (env or config), so a
container cwd is never "healed" to its nearest host ancestor (/).
- map_cache_path_to_container also matches through a symlinked HERMES_HOME:
@file: expansion passes resolved paths, the mount roots keep the configured
spelling, so staged attachments leaked their host path.
Covers the desktop attachment path of #103147. Typed @file: refs to binary
files outside the mounted dirs still render the host path.
Re-review of the trust-anchor follow-up:
- Cron runs and Kanban workers no longer anchor trust. The cronjob tool
lets the model pick a job's workdir, which becomes the run's session
cwd, and kanban_create lets it pick a task workspace, which becomes the
worker's launch dir and TERMINAL_CWD. Either one let an agent-cloned
repo become trusted. Both now rely on lsp.trusted_workspaces only.
- A repository at or above $HOME never anchors, not only one at $HOME
exactly.
- Operator roots are recorded only where a tool thread enters
(enabled_for) and when the service is created. _trusted() is now a pure
read. Before, it took _state_lock while two of its callers already held
it, so a change in the roots could deadlock the LSP loop.
- Client lookups go through _live_key(). A multi-root client that
started while its root was untrusted is still found and released after
the root becomes trusted, instead of being orphaned.
- `hermes lsp status` stops listing roots that have since become trusted.
The lint gate reads the cwd the same way the linter subprocess does.
- The path-list parser returns None for a malformed value, and each key
logs its own warning. The _node_modules_trees docstring no longer
mentions Vue.
The trust anchor was os.getcwd(), which is the user's project only for a
plain `hermes` launched inside a repo. `hermes -w` sets TERMINAL_CWD to its
new worktree without a chdir, the Desktop/TUI backend runs from the install
tree or $HOME with the project bound as the session cwd, the gateway runs
from HERMES_HOME, and cron binds a per-job workdir. All of them lost
pyright's project venv, rust-analyzer/gopls/jdtls/... and the .ts/.rs lint
fallbacks inside the user's own repository.
operator_workspace_roots() now adds the worktree of resolve_agent_cwd()
(session cwd, then TERMINAL_CWD) to the launch dir's; only surfaces set
those, the agent's `cd` moves the terminal env's cwd, not them. $HOME is
never an anchor: a dotfiles repo there would trust every directory below it.
The service remembers every anchor it has seen, because the loop thread
that spawns servers has no session context of its own.
- The lint gate checks only the linter's cwd: npx and rustup resolve the
toolchain from there (subprocess cwd=env.cwd), not from the file's dir.
The try/except that wrapped code which cannot fail is gone.
- Trust is computed once per spawn and handed to ServerContext.
- Multi-root servers (pyright) share one process across trusted roots
only; an untrusted root gets its own, so whichever root spawned first no
longer decides the interpreter for the others.
Pinning settings per server left every other server free to run project
code: rust-analyzer runs cargo check (build.rs, proc-macros) on each
didSave despite the disabled init options, and jdtls, kotlin-language-server,
elixir-ls, zls, clojure-lsp and haskell-language-server evaluate build
files when they start.
In an untrusted workspace only UNTRUSTED_SAFE_SERVERS now start (pyright,
typescript, vue, svelte with their pinned settings; bash, yaml,
dockerfile, intelephense, clangd without --query-driver). Every other
built-in or user-declared server is skipped at the spawn chokepoint and
in enabled_for, logged once per root and listed by hermes lsp status.
The rust-analyzer init tweak is gone: it never starts untrusted now.
On a local backend the npx tsc and rustfmt --check fallbacks are skipped
the same way, since npx resolves the repo's node_modules/.bin/tsc (or its
.npmrc registry) and rustup honours its rust-toolchain.toml.
The boot-failure report added a fourth hand-written <dm>.live.json join; it
must agree with the path _admit_live_dm pins, so all four sites now share
_live_intent_file.
On Windows, hermes_bootstrap relaunches onto the store interpreter with
subprocess.call and exits with the child's status. The child is the same
delivery runner and has already printed its own outcome, so the parent
treating a nonzero status as "could not activate" printed a second JSON
line that could turn a definite failed/rejected into "ambiguous, do not
resend". The relaunch now raises RelaunchExit (a SystemExit subclass, so
every existing handler and the process exit code are unchanged) and the
runner skips its boot-failure report for it.
Also parses the runner argv in one stdlib-only helper, _runner_argv, used
by both _delivery_main and the boot-failure report, so the two can no
longer disagree about where the DM file is.
Narrows the previous commit to the one case where the outcome is really
unknown. When the sender has already pinned a live intent (<dm>.live.json)
it may have admitted the DM and told the model not to resend; a runner
that then fails hermes_bootstrap now prints the same ambiguous payload
(same delivery id, "Do not resend") that _run_delivery prints for an
admission failure. Without an intent nothing was handed over, so the plain
repair hint stays the truth and a resend remains correct.
The payload is shared with _run_delivery via _live_outcome_unknown instead
of a second classifier, and the evidence-probing variants and the separate
test module are dropped in favour of one real -I -S subprocess test plus an
explicit negative assertion on the existing no-intent test.
The --wait-reply waiter is stdlib only and its stdout is the sender's
wake-up. Booting it through hermes_bootstrap meant an install whose
committed environment cannot be activated killed the waiter with the
repair hint on stderr before it printed anything, so the sender never
learned the delivery outcome. Only the delivery lane boots now.
Port the failure-mode test from #125754: a --run-delivery runner that
cannot activate exits with the `hermes pm repair` remedy and leaves the
DM untouched, instead of the ambiguous "No module named ..." outcome.
Co-authored-by: woriwka-ai <247827632+woriwka-ai@users.noreply.github.com>
message_agent spawns its background runner as
`<sys.executable> tools/bot_mode_dm.py --run-delivery ...`, and the relay's
reply waiter (bot_relay.waiter_command) as `... bot_mode_dm.py --wait-reply ...`.
Both enter through the shared `if __name__ == "__main__":` block, which only
put the repo root on sys.path. Under a PM-managed install sys.executable is the
bare store interpreter: dependencies are activated in-process at boot by
hermes_bootstrap (pm.environments.activate_dependencies), and the terminal
backend strips the Hermes-owned PYTHONPATH, so the child has no third-party
packages. Live, every local Bot Chat DM returned
`{"status": "ambiguous", "error": "Live admission outcome unknown: No module
named 'ruamel'. Do not resend."}` from _run_delivery -> _admit_live_dm.
The __main__ block now imports hermes_bootstrap after putting the repo root on
sys.path (run as a script path, sys.path[0] is tools/). This is the same pattern
as tui_gateway/compute_host.py and agent/transports/hermes_tools_mcp_server.py:
only the script entry bootstraps, so importing tools.bot_mode_dm as a library
stays side-effect free and the module's top-level imports stay stdlib-only.
hermes_bootstrap rather than a bare activate_dependencies() call, because the
runner is a real entry point and needs what every entry point gets: dependency
activation plus early recovery of an interrupted update, the Windows UTF-8 stdio
bootstrap (its stdout carries the teammate's reply as the completion
notification), and import-path hardening against a cwd package shadowing
Hermes modules. The runner argv is safe for the bootstrap: command_argv() reads
`--run-delivery`/`--wait-reply` as flags, so it is neither `update --post-swap`
nor `pm repair`. prepare_launch acts only on a self-updated install whose
dependencies are behind or whose interpreter is not the store Python, and its
relaunch re-runs the same script path with the same argv (runpy.run_path), so
the runner keeps its contract there too.
tests/tools/test_bot_mode_dm_entry.py launches the real script under `-I -S`
with a PM-committed dependency record in a temp HERMES_HOME as the only route to
the packages: the --run-delivery case fails on the parent commit with the exact
live `No module named 'ruamel'` error and passes here; the --wait-reply case
drives the argv waiter_command builds through the same entry.
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales
* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter
ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.
* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs
* chore(tui): split en catalog siblings by lane (slash sibling)
* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports
* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts
* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list
* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes
* feat(plugins): report language-pack layers in the mid-run activation summary
* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n
StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.
* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter
- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
layers partial packs (nested or flat dotted) over bundled/en via
mergeTranslations; a string over a function-valued en entry becomes a
positional {0}/{1} formatter; $appLocaleVersion bumps so translators
re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
registry as source 'backend' (method-not-found is silent); re-synced on
socket open, display.language change and profile switch. A saved pack-only
language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).
* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)
Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.
* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()
- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces
* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)
* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()
Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.
locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).
* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()
- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions
* i18n(platforms): route Google Chat and Teams user-facing text through t()
Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.
Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.
* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()
LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.
* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)
Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.
* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)
_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.
* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)
RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.
* i18n(cli): live-work dock, subagent monitor and render copy through t()
cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.
* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()
* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)
get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.
* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output
* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time
Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).
* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)
cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.
* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()
- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
at call time; category labels via slash.category.*; help/alias/usage suffixes
via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
keeps its English constant; callers use history_unreadable() ->
gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.
* i18n(telegram): route adapter chat copy through t()
Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.
Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.
* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()
- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor
* i18n(discord): route adapter chat copy through t()
Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().
* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)
* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()
- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
_DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
Updating/Generating) are one full template per variant; plurals use
<key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).
* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English
* test(i18n): pin Telegram/Discord adapter catalog wiring
Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.
* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge
* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint
* i18n(tr): translate bundled catalog + tui pack
* i18n(ja): translate bundled catalog + tui pack
* i18n(ko): translate bundled catalog + tui pack
* i18n(zh): translate bundled catalog + tui pack
* i18n(fr): translate bundled catalog + tui pack
* i18n(af): translate bundled catalog + tui pack
* i18n(uk): translate bundled catalog + tui pack
* i18n(ar): translate bundled catalog + tui pack
* i18n(pt): translate bundled catalog + tui pack
* i18n(it): translate bundled catalog + tui pack
* i18n(es): translate bundled catalog + tui pack
* i18n(zh-hant): translate bundled catalog + tui pack
* i18n(ru): translate bundled catalog + tui pack
* i18n(hu): translate bundled catalog + tui pack
* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)
* i18n(de): translate bundled catalog + tui pack
* i18n(ga): translate bundled catalog + tui pack
* test(i18n): fixture matches _normalize_lang(lang, home) signature
* i18n(tui): scaffold userMessages/slashCmd en siblings
* i18n(tui): wire secure prompts + content tables
* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays
Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.
* i18n(tui): wire slash ops/wake replies
* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)
* i18n(tui): wire slash core/debug/setup replies
* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)
* i18n(tui): wire slash session/topup/subscription replies
* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)
* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys
* i18n(tui): wire userMessages copy through the userMessages namespace
* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json
* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)
Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).
* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)
* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)
* tui: i18n-export-en script (English templates for pack translators)
* docs(i18n): bundled TUI packs are the bottom layer of the tui surface
* i18n(ru): translate TUI pack
* i18n(ar): translate TUI pack
* i18n(es): translate TUI pack
* i18n(pt): translate TUI pack
* i18n(ko): translate TUI pack
* i18n(de): translate TUI pack
* i18n(ja): translate TUI pack
* i18n(fr): translate TUI pack
* i18n(tr): translate TUI pack
* i18n(it): translate TUI pack
* i18n(zh): translate TUI pack
* i18n(zh-hant): translate TUI pack
* i18n(hu): translate TUI pack
* i18n(uk): translate TUI pack
1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.
Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).
* i18n(ga): translate TUI pack
* i18n(af): translate TUI pack
* plugin_guard: locale catalogs in language packs step down the agent-config family
A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.
* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)
* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header
Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.
* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel
- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)
* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})
* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property
* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal
* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export
The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.
* test: unbreak two main-red timing tests the PR merge-ref inherits
- test_local_runtime racing fake publishes the modern state record (legacy pid-only
records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
a fixed 0.35s, which a loaded CI runner does not always meet
* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)
* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak
SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.
* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings
* chore(i18n): regenerate desktop key catalog after main sync
---------
Co-authored-by: Teknium <teknium@nousresearch.com>
andrexibiza (review on 60bdd5fbe3): skipping an unreadable plugin dir,
or an unreadable or unparsable flat manifest, replaced the home's
cached secret set with the partial result. A secret declared a moment
ago could then reach children while its manifest was unreadable. And
because the partial result was cached under the same file signature,
recovery could keep hitting it.
platform_manifest_secret_scan() now reports whether the scan was
complete. A partial scan unions in the names this home already had,
and is cached with no stamp, so the next spawn rescans and a recovery
is seen at once. A deleted plugin still releases its names, because
that scan is complete. An unreadable plugins/platforms/ manifest still
raises.
Per-commit review follow-ups:
- Dot directories are scanned again, matching plugins_discovery. A
platform plugin that discovery loads from plugins/.x now also has
its secrets declared.
- Only PermissionError means a plugin directory cannot be searched. A
symlink loop is treated as missing, so plugin.yml is still checked,
as discovery does. A plugin.yaml that is not a regular file is
ignored rather than failing every strict read.
- Dropped a redundant union: the home set is already stripped in
Tier 1.
- Tests now pin the dunder skip, the unlistable platforms root, and a
stamp change on an edit that keeps the mtime.
Per-commit review follow-ups on the per-home plugin declarations:
- The manifest stamp walked each plugin directory with pathlib
is_dir/exists/stat on every spawn, and _scrub_credentials took it
twice. It is now one scandir per root and one stat per candidate,
taken once per scrub.
- The stamp uses utils.file_signature, so a manifest replaced with its
mtime preserved (cp -p, rsync -t) still invalidates the cache.
- An unsearchable plugin directory, including an unlistable flat
plugins/ root, is reported once as a directory entry. The stamp walk
stays quiet; the warning comes from the manifest read, which runs
only on a cache miss.
- The dot/dunder skip now matches plugins_discovery.scan_directory, and
the manifest source argument is a Literal.
- Tests: the unsearchable-dir case uses a non-dunder dir, so both
guards are pinned. New tests cover the bundled read failing closed,
manifest deletion releasing its names, and a symlinked home alias
sharing one cache entry.
Phase 2 review follow-ups:
- OpenViking overlaid the bound profile's whole .env after the scrub,
so that profile's bot, dashboard and relay tokens reached the server
again. The overlaid env is now scrubbed a second time for Tier 1.
Provider keys still pass, and they still come from the bound profile.
- The strict per-home manifest read raised on any unreadable entry
under <home>/plugins, which failed every child spawn for that
profile. It is now strict only where the manifest is known to be a
platform's: bundled or plugins/platforms/. Dunder and dot
directories, unsearchable plugin directories and unreadable flat
plugins/* manifests are skipped with a warning, because they can't
load as plugins either.
- The per-home cache is keyed on hermes_home_key(), so a symlinked
alias no longer gets a second entry.
- The modal snapshot test swapped hermes_cli.config for a stub, which
the policy import can no longer use. HERMES_HOME already points the
real module at the test home, so the stub is gone.
- The env_passthrough docstring now describes declared names rather
than the dropped prefix rule.
/simplify-code follow-ups on the declared-secret policy:
- The policy and the manifest classifier each defined which env-name
suffixes make a platform variable a secret, and they disagreed on
_JSON. PLATFORM_SECRET_ENV_SUFFIXES in hermes_cli/config.py is now
the only definition.
- The gateway env-override table was walked by hand a fourth time,
through private gateway.config_env names. The walk is now
profile_channels.config_env_table_keys(), shared with
declared_channel_env_keys.
- The bundled platform manifests were parsed twice at import (about
16 ms). The config injection reads them once, strictly, and keeps
their secret names. The policy re-reads only if that read hit an
I/O error.
- The per-home cache was keyed on directory mtimes, so an in-place
plugin.yaml edit went unseen until restart. It is now keyed on
every manifest file's mtime.
- The core-declared snapshot is a module-level constant rather than
a global set inside the injector. The two manifest-source booleans
are now one source argument. The Tier 1 set is upper-cased once
at import, not on every spawn.
Review follow-ups on the declared-secret policy:
- A profile's user-installed platform plugins declared secrets into
one process-wide set, read from the launch home at import. Under
multiplex that stripped profile B's same-named user variable and
missed profile A's own declarations. Bundled manifests stay
process-wide; user manifests are now read per bound home, cached
until that home's plugin dirs change, and are Tier 1 for that
profile only.
- The declaration scans no longer fail open. A registry error is no
longer swallowed into an empty set, non-string required_env entries
are skipped rather than breaking the set, and a manifest that
cannot be read raises instead of vanishing from the policy. A
malformed manifest still declares nothing.
- A plugin manifest can no longer reclassify a core-declared name
such as OPENAI_API_KEY; the same rule the config form applies.
- The dashboard basic-auth password and signing secret, the OIDC
client secret and the drain bearer move to Tier 1. Credentialed
CLIs (claude, codex) no longer receive them.
- openviking-server starts from served_profile_child_env, so it gets
the bound profile's provider keys, never the launch profile's.
With no bound profile under multiplex the start is refused.
- NOUS_API_KEY and QWEN_API_KEY are listed statically. Discovering
provider plugins while the policy module imported re-mirrored them
over a plugin's own auth registry entry
(tests/providers/test_auth_registry_import_order.py).
Two more spawns inherited the gateway's full environment. The bot
desktop launcher copied os.environ minus the display variables; the
agent drives that desktop and its dock opens xfce4-terminal, so every
provider key and bot token was one click away. The Telegram gmail-triage
buttons ran HERMES_HOME/scripts/gmail-triage/*.sh with no env at all, a
path the agent can write. The desktop now builds from
served_profile_child_env() like the agent browser (keeping the user's
HOME); the triage scripts use build_subprocess_env like cron and
webhook-filter scripts.
Review of the adapter-secret salvage found Hermes-owned secrets that no
declared source covers, so they reached terminal and execute_code
children and skill passthrough accepted them: the dashboard basic-auth
password and signing secret (enough to forge dashboard sessions; the
session token beside them is already Tier 1), the dashboard drain and
OIDC client secrets, HERMES_ANON_API_SECRET (a provider-category entry,
which the OPTIONAL_ENV_VARS loop never blocks) and the Google Meet
realtime key. Add them to the static blocklist.
The shape rule (a platform prefix plus _TOKEN/_SECRET/_PASSWORD/_KEY) also
matched variables Hermes never reads: the platform list holds plain words
(LOCAL, GATEWAY, WEBHOOK, SLACK), so a user's SLACK_USER_TOKEN,
LOCAL_LLM_API_KEY or GATEWAY_API_KEY vanished from the terminal and
terminal.env_passthrough could not bring them back. The prefix census it
leaned on also failed open (an unreadable plugins dir cached an empty set).
Adapter secrets now come from what adapters declare: password entries of
the messaging OPTIONAL_ENV_VARS (built-ins plus every platform plugin
manifest) and secret-named keys of the gateway env-override table, both
Tier 1 and refused by passthrough; plus the secret-named required_env of
adapters registered in the current profile scope, read per spawn without
loading deferred adapters, Tier 2 only because required_env is an
unchecked setup list. Secrets nothing declared get a manifest entry
(TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN, A2A_PUSH_SECRET,
TEAMS_GRAPH_ACCESS_TOKEN, TEAMS_INCOMING_WEBHOOK_URL) or join the
policy's read-in-code list (QQ_STT_API_KEY, the two MSGRAPH names).
Skill and config passthrough names were checked against the managed-credential policy only when accepted. A plugin adapter registering later claims <PREFIX>_*_SECRET, but is_env_passthrough() and get_all_passthrough() kept returning the stale approval, so terminal, background and execute_code children (and scope-only additions) still received the secret. Both now re-apply the refusal when the allowlist is consumed.
An inheriting child (claude/codex/gemini) kept TELEGRAM_WEBHOOK_SECRET, WHATSAPP_CLOUD_ACCESS_TOKEN and the Graph secrets because the shape rule sat only in the provider tier; it now strips with the bot tokens. The shape rule reads the per-call adapter census, so a plugin adapter registered after import, in the bound profile only, is covered. The OAuth provider scan takes bundled profiles only, so the process-wide blocklist no longer freezes the home bound at import.
The child-env blocklist is derived from the provider registry and
OPTIONAL_ENV_VARS, but many gateway adapters read their secrets straight
from the environment without listing them there: WHATSAPP_CLOUD_ACCESS_TOKEN,
WHATSAPP_CLOUD_APP_SECRET, WEIXIN_TOKEN, YUANBAO_APP_SECRET, FEISHU_ENCRYPT_KEY,
TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN and others. They reached
terminal, background/PTY, cron-script and hermes_subprocess_env children,
and a skill could register them as env passthrough, while the documented
bot tokens next to them were stripped. OAuth provider profiles (nous,
qwen-oauth) also accept a pasted key (NOUS_API_KEY, QWEN_API_KEY), but the
registry mirror copies env_vars only for api_key profiles, so those passed
through too.
Match adapter secrets by shape, the way authorization gates already are: a
built-in or bundled adapter prefix plus a _TOKEN/_SECRET/_PASSWORD/_KEY
suffix, so a new adapter secret is covered without a second edit. Add every
provider profile's env_vars regardless of auth_type, and list the two
Microsoft Graph secrets no adapter prefix owns.
_run_foreground recorded outcome=timeout whenever a backend exception's
text contained "timeout" (an SSH connect timeout, a sandbox API timeout).
That command never reached an exit status, and the doc says timeout comes
only from Hermes' own deadline flag (hermes_timed_out), which
terminal_outcome() already reads. Owner ruling: never from exception text.
The exception path now records nothing; the model-facing result (exit 124)
is unchanged.
Probe (code-read finding; the new test drives the real terminal_tool with a
backend raising "... Connection timeout"):
before: terminal.outcome {backend: local, command_kind: shell_builtin, outcome: timeout} x1
after: no terminal.outcome row; tool result still exit_code 124
test_backend_exception_mentioning_timeout_is_not_a_terminal_timeout is RED on
base, green here.
- hermes.tool_unavailable.count: model called a shipped built-in (BUILTIN_TOOL_NAMES) not enabled in
the session; tool name + catalog provider/model. Unknown names stay in v4 unknown_tool quality.
- hermes.provider_setup.count: started/completed/abandoned/failed per provider and surface
(cli_setup, cli_model, tui, desktop, dashboard) with a closed failure_class. Pending marker per flow;
a dead or stale marker is reported abandoned at the next start (v4 process-marker pattern).
- hermes.feature_adoption.count: once per feature per install at first real use, derived from existing
counters in the subscriber (metric -> feature table) plus Bot Mode / Projects hooks; bucketed by the
owning profile's install age.
- hermes.feature_disabled.count: turning off a default-on toolset/skill/plugin/platform/setting (and the
re-enable back to default), diffed at the config write chokepoints; once per (kind,name,event)/day.
Surface comes from the user entry point; setup/migrations record nothing.
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):
- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
matcher reports the strategy that landed (or no_match / ambiguous) into a
context-local probe opened only around the patch/write_file handlers, so we
learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
the guardrail already warns/blocks/halts and at turn end for the iteration
budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
next_outcome. One row per failed tool call, resolved against the model's
next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
a table lookup on the first program word, never the text. Hermes' own
deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
backend result, so a command's own `exit 124` reads nonzero, not timeout.
Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
truncation only from the structured finish reason; empty/reasoning_only
from the normalized message; issue=none per response as the denominator
(model_route counts attempts and files empty replies as failures).
Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
- memory tool: one row per operation (a batch counts each op); gate/validation
refusals are "rejected", store-level non-application "failed". Memory provider
tools are counted in MemoryManager.handle_tool_call with the op read from the
action arg or tool-name verb.
- curator: run_curator_review records one row per pass (dry run = skipped) with
the before/after diff bucketed; a scheduled pass whose claim is held by another
process counts as skipped. The home is captured before the review thread starts.
- delegate_task: _run_batch opens the call, each joined unit folds its results in,
and the last unit emits the call's single row, so group-split background calls
are not counted per unit.
- terminal / execute_code / browser: counted per call that reached the backend;
the backend is resolved from the owning profile at call time (terminal plan
env_type, code local/remote, browser CDP > Camofox > cloud provider > engine, or
"extension" when the extension controller served the call). Guard refusals and
Hermes' own _host_local commands are not counted.
Tests read rows back from the real store; each call-site group is red with its
source reverted.
Every path that adds an MCP server now emits one
hermes.extension.install.count event through record_extension_install:
- catalog installs (hermes mcp install, the picker, the dashboard route and
its background CLI action) via mcp_catalog.install_entry
- catalog installs from connector cards and the agent's catalog tool via
_CatalogBackend.install / start_install_oauth (OAuth: success at commit,
failed when the flow cannot start; an abandoned browser step is a cancel)
- custom servers from `hermes mcp add`, POST /api/mcp/servers and the
mcp.add RPC (source url for http, local for stdio, name None)
A reinstall or overwrite of an already-configured server is not counted,
so the metric measures new installs rather than config churn. Catalog
names pass through raw; the contract reports non-catalog names as custom.
* feat(vercel): start fresh sandboxes from a managed image instead of the deprecated runtime
Vercel deprecated Sandbox runtimes (node24/node22/python3.13) in Aug 2026 in favour of
images, and rejects runtime+image together and runtime with a snapshot source. New
terminal.vercel_image (default vercel/sandbox/universal:latest, Node 24 + Python 3.14)
picks the image for fresh sandboxes; a pinned terminal.vercel_runtime still works, wins
over the image and logs a deprecation warning; snapshot restores send neither.
Setup wizard prompts for the image, dashboard exposes both keys, status/config show the
effective choice, TERMINAL_VERCEL_IMAGE bridges config to the tool like its siblings.
* feat(config): migration 49 drops the seeded node24 Vercel runtime pin
Every pre-49 config.yaml carries terminal.vercel_runtime: node24 (the template default) and
the setup wizard mirrored it into .env as TERMINAL_VERCEL_RUNTIME. Both are the default
copied, not a choice, so the migration drops them and fresh sandboxes follow vercel_image;
node22 / python3.13 pins are the user's and survive. Persisted sandboxes are unaffected:
a snapshot restore never sends a runtime or an image.
* fix(modal): keep persistent-sandbox snapshots past the SDK's 30-day TTL
modal>=1.5 gives Sandbox.snapshot_filesystem() a default ttl of 30 days, so an idle
persistent Modal sandbox silently lost its filesystem and restarted from the base image.
Pass ttl=None (retain until deleted) and bump the modal extra from 1.3.4 (no ttl
parameter; legacy RPC) to 1.5.5 so the kwarg exists on every install.
* fix(modal): drop the dead modal.Mount credential-mount block
modal.Mount left the public API in modal 1.0, so _modal.Mount.from_local_file raised
AttributeError into the surrounding except on every sandbox start and the block never
mounted anything. The FileSyncManager created right after already uploads the same
credential, skills and cache files (iter_sync_files), so delete the duplicate; the test
fake stops exporting a Mount the real SDK does not have.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
The warm_command/release_command hook thread was a bare threading.Thread, so under
the multiplexer resolve_passthrough_value() saw no secret scope, raised
UnscopedSecretError, and the best-effort hook swallowed it at debug level: every
command provider with a non-empty env_passthrough silently never warmed or released.
Bind the thread with ctx_bound (copy_context().run), the same seam the keep-warm
timer already uses. Regression test: the hook thread resolves the bound profile's
passthrough value.
Docs (review minor): the multiplex isolation table now names per-profile slash-command
gating and the fail-closed empty-admin policy for a served profile with no cached config.
run_command_provider forwarded declared env_passthrough keys from os.environ,
which under the multiplexer is the LAUNCH profile's .env. A served profile's
command TTS/STT subprocess (voice-note transcription, speech, warm/release
hooks) got the launch profile's credential and never its own. Resolve each key
with resolve_passthrough_value, as the terminal and code-execution spawns do.
(cherry picked from commit 22a299d5f562e1d174fa5d37590cd46361a1383a)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
A script that calls hermes_tools.write_file/patch and never prints the
returned {"error": ...} got status=success with empty output, so the
model believed the write happened. In a cross-agent bench this cost 3 of
18 Opus runs 4-6 extra turns (write_file refused by the read-before-write
guard, then re-probing and re-emitting whole files). The result now
carries tool_errors for that cell, and the in-script write_file doc says
existing files must be read first.
`hermes logs --since` and `--level` passed every line that had no leading
timestamp. A traceback's frames are written without one, so an old error
printed its frames without the header that was filtered out, and
`--component` dropped the frames of a matching record. `hermes logs` now
reads each unstamped line as part of the record above it: the line gets
that record's verdict for every filter, in the tail read and in `-f`.
Lines before the first stamp in the read window have an unknown time and
level, so they are dropped when `--since` or `--level` is set.
mcp-stderr.log had no parseable stamp at all: the banner started with
`=====` and server output was copied raw, so `hermes logs mcp --since`
printed the whole file. The stderr tee already reads each server's
stderr in a thread, so it now writes one line at a time, each prefixed
with the asctime-shaped local stamp the Python logs use. The banner
starts with the same stamp. The stamp comes from new public
`timestamp()`/`stamp_line()` in hermes_cli/stderr_timestamp.py, which
stays stdlib-only. The desktop MCP log view accepts both banner shapes.
A test now requires a real writer's sample line for every LOG_FILES
entry to parse with `_parse_line_timestamp`. The docs no longer say the
`timezone` key changes log timestamps; log lines use the machine's
local time.
Trim the salvaged detector/delivery pair to the salvage bar and fix the
delivery half. The PR resolved the refresh selection as
platform_toolsets.<surface> (desktop/tui), a key no config path writes,
so _get_platform_tools fell back to the constructed hermes-<surface>
composite that resolves to 0 tools (a cold resume went 30 -> 1 tool).
Delivery now goes through the builder new desktop/TUI sessions use
(tui_gateway.server._load_enabled_toolsets(platform) +
_load_disabled_toolsets) as an explicit refresh_agent_mcp_tools
override, so the refreshed set equals a fresh session's on that surface;
only desktop/tui agents are long-lived, every other surface builds a
fresh agent per process/turn and is left alone. Drops the
toolsets.py/tui_gateway policy relocation and the raw-value fallback in
the fingerprint; keeps the detector (platform_toolsets +
agent.disabled_toolsets replace the dead tools.enabled_toolsets key)
and the tools re-pin after the refresh.
Tests: 2 invariants in tests/agent/test_bot_chat_toolset_refresh.py
(desktop -> builder consulted, disabled override passed, re-pinned;
cli -> untouched) and the existing fingerprint axis test now edits
the key `hermes tools enable/disable` writes.
The capability epoch watched tools.enabled_toolsets, a key no surface
writes; real hermes tools enable/disable edits (platform_toolsets.* plus
agent.disabled_toolsets) never flipped it. Even when stale, the refresh
rebuilt only the prompt and never tools[]. Watch the real keys and
rebuild + re-pin the tool snapshot on Bot Chat capability refresh.
Closes#124211
Review finding (major): ProcessRegistry.restore_completions() is once-per-process but read
async_delegation._db_path() (ContextVar-aware) under whatever scope the FIRST consumer ran in.
In the TUI gateway both first consumers (the session notification poller and the prompt_turn
drain) run inside _session_profile_runtime_scope(session), so under multi-profile
`hermes serve` the first session's profile ledger was replayed and the LAUNCH profile's
undelivered completions were never replayed for the life of the process (origin/main's
import-time replay always covered the launch profile).
Fix: restore_completions() clears the hermes-home override for the duration of the replay
(set_hermes_home_override(None) + reset), so the once-per-process replay always reads the
launch ledger whoever gets there first; the caller's scope is restored afterwards. Smaller
than a per-home restored-set plus a tui_gateway boot hook: it restores exactly main's
invariant with no new boot seam, and secondaries stay where they were (gateway
_restore_secondary_completion_ledgers). Docstring states the invariant.
Tests:
- tests/tools/test_process_registry_lazy_restore.py::test_first_drain_under_secondary_scope_replays_the_launch_ledger
(red on the PR head: replayed profiles/b/state.db; green now)
- tests/gateway/test_multiplex_unserved_shared_ingress.py::test_boot_replays_the_launch_ledger_before_secondaries_and_watchers
(minor: pins the gateway boot hook — launch-scope replay in _start_secondary_profiles, i.e.
before the secondary bind and before _async_delegation_watcher, the only production consumer
that reads completion_queue.get_nowait() without drain_notifications)
WHAT
- tools/process_registry.py: ProcessRegistry.__init__ no longer calls
restore_undelivered_completions(); a new once-per-process
restore_completions() does, invoked by the first consumer:
drain_notifications() (CLI process_loop / TUI prompt turn), the gateway
startup (run_startup._start_secondary_profiles, right before the secondary
ledgers are replayed) and the TUI session notification poller.
- tools/async_delegation.py: restore_undelivered_completions() returns 0 when
<HERMES_HOME>/state.db does not exist (a replay must not create or migrate
the ledger); _connect() creates the parent via mkdir_under_hermes_home()
instead of a bare mkdir so a late writer cannot resurrect a deleted
(tombstoned) or missing named profile.
- gateway/run_notifications.py: docstring no longer claims the import restores
the launch ledger.
WHY
`process_registry = ProcessRegistry()` runs at module import, and model_tools
imports it transitively (tools.registry -> close_terminal_tool), so
`python -c 'import model_tools'` (hermes doctor, any tool-registry consumer)
under a fresh/typo HERMES_HOME created the full profile skeleton + state.db,
recreated a tombstoned `profiles/.deleted/<name>` profile, and ran
reconcile_state_schema() against an existing store — bypassing the
assert_named_profile_home_live / mkdir_under_hermes_home guards (#97128,
#112592). Restoring on first consume keeps the replay for every process that
actually drains the queue while an import performs no state.db I/O.
Salvaged from #123348 (@webtecnica: lazy restore + missing-ledger guard) and
#123298 (@Wenfengcheng: mkdir_under_hermes_home in _connect), trimmed to the
minimal shape (no completion_queue property, no read-only probe).
Fixes#123265
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
(cherry picked from commit b77c4ff49062d2936271ded80821b48b0c11d368)
MAJOR — `_builtin_gateway_liveness` trusted a bare-epoch ticker heartbeat for
~200 s after its writer died, for every profile, and the multiplexer pid gate
had been dropped: a killed serve/Desktop ticker read as alive and
`hermes -p X cron run` queued a primary-routed run "for the gateway's next
tick" with no ticker present.
- `cron/jobs.py::record_ticker_heartbeat` stamps `<epoch> <pid>`;
`get_ticker_heartbeat_age` parses the first field (legacy bare stamps still
yield an age); new `ticker_heartbeat_writer_alive` requires the stamped pid
to be alive and treats a bare stamp as NOT proof by itself.
- The heartbeat-only rung is now `fresh AND (served-by-multiplexer OR writer
alive)`: the multiplexer record proves the host process for a served named
profile (and covers a stale-code multiplexer still writing bare stamps), so
only the in-process serve/Desktop ticker relies on the heartbeat, and then
only with a live writer. `cron status`'s in-process-ticker rung applies the
same rule.
MINOR (a) — `_hand_off_primary_routed_run` gated on `is True`; an unknown
(None) probe returns an error naming the uncertainty instead of "queued …
runs and delivers it".
MINOR (b) — `_run_claimed_job` wraps a shared-bot satellite's resolved map in
`SharedRouteAdapters(primary, _primary_profile_routes_for_current_home())`,
the same grant the ticker's `tick_adapters_for` makes, instead of handing it
the full primary adapter map.
Tests (each red on the previous head a18977090be):
- test_cron_satellite_diagnostics.py::test_in_process_ticker_heartbeat_counts_only_while_its_writer_lives
- test_cronjob_run_primary_routed.py::test_routed_run_without_a_serving_gateway_fails_before_the_turn[None]
- test_cronjob_run_immediate.py::test_execute_job_now_grants_a_shared_bot_satellite_only_its_routed_targets
A multiplexed satellite profile with no platforms.<p> credential of its own
posts through the primary's bot via a root gateway.profile_routes entry.
The delivery preflight lets such a job through (#97476) on the assumption
that the primary gateway's live adapters send it, which holds on a
scheduler tick but not for a manual run: `hermes -p <profile> cron run`
executed the whole agent turn in the CLI process, which has no sender for
the route, then failed delivery with "platform 'telegram' not
configured/enabled" and overwrote last_status with delivery_failed.
cronjob(action='run') now checks, before the in-process claim and before
the background dispatch, whether any delivery platform of a runnable job
is reachable only through the primary route (routed to this profile, no
connected credential here). If so:
- a gateway that serves the profile is live (or liveness is unknown):
queue the run with trigger_job so that gateway's ticker runs and
delivers it; job status is untouched and `cron run` prints "It will run
on the next scheduler tick";
- no gateway serves it: fail fast before the agent turn, like the
relay-fronted forward does when its api_server is unreachable.
Runs inside the gateway process (its live adapter delivers, #89302),
paused jobs (trigger_job would resume them), local delivery, and profiles
with their own credential keep the in-process path unchanged.
Fixes#120330
(cherry picked from commit 5a0db1c4e37de865f23034937b93d38fa87414f4)
Review feedback: the previous try/except restored runner.adapters on any
resolution error, which is the cross-profile misdelivery this fix
prevents. Let the error propagate to _run_claimed_job's handler, which
marks the run failed and surfaces the message. Runners without
_adapters_for_profile (shims, tests) still keep runner.adapters.
Adds a regression test where resolution raises: run not fired, marked
failed, error surfaced.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 183547ea6b3798874dc8a98f943ef98ee710911b)
A manual `cronjob(action="run")` fired from a secondary profile's agent
resolved delivery through `runner.adapters`, which is the default
profile's map. Under `gateway.multiplex_profiles` HERMES_HOME is
overridden to the owner profile mid-turn, so the result left through the
default profile's bot (a Telegram DM arrived from the wrong bot while the
run was correctly recorded on the secondary profile's cron).
Resolve the owner profile from HERMES_HOME and use
`runner._adapters_for_profile()`, the same fail-closed path notifications
and goal loops already use. Runners without that method (tests, older
shims) keep the previous behaviour.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 350cfc1436504b92d5f9e924f43f81b232beefd7)
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.
Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
Under multiplex a served secondary profile's remote MCP server whose config has a
`${VAR}` Authorization header was brought up at gateway boot with the placeholder
unresolved (HTTP 401) and retried that same rendering forever.
Mechanism: `_owner_scope_home()` trusted any caller-bound secret scope, and the
gateway's boot-time `_profile_runtime_scope` binding is a SNAPSHOT taken before the
profile's external secret source (`secrets.command`) may have answered. The run
task copies that context, and `_refresh_remote_config` re-read config.yaml on every
probe but rendered it under the same frozen mapping, so a parked server sent the
literal `Bearer ${VAR}` every ~5 min for the life of the process.
- `_owner_scope_home()` now returns the owner home (the scope's stamped home, else
the registry scope) even when a scope is bound: MCP rebuilds the owner's scope
fresh, which retries hydration (cached once it succeeds). Single-profile
processes (scope key None) are unchanged.
- `MCPServerTask.run()` binds that fresh owner scope around the rebuild-time config
refresh, so a reconnect heals once the source hydrates.
- `_require_rendered_remote` fails closed: a remote `url`/`headers` still carrying
a `${VAR}` after rendering raises naming the variable instead of sending it.
- Salvaged from #119097 (@JoaoMarcos44): the connect-time re-render, which also
covers a lazy server whose config was rendered at boot; its mechanism alone was a
no-op for the reported topology because the owner-scope install was skipped
whenever a (frozen) scope was already bound.
Row L1/L3, topology T2. Tests: A→B→A under set_multiplex_active(True) with two
temp homes, red on origin/main (literal placeholder sent), green on head.
Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
Independent-review minors on #126157 (#123989 class):
- session.create is a plain @method, so tui_gateway/session_workdir.py::_register_session_cwd
wrote the cwd record under the RAW session key while the scoped turn read
profile:<p>:<key> and missed it until the first `cd`. The writer now binds the
session's own profile_home around register_task_env_overrides.
- tools/terminal_tool.py::_task_env_overrides stayed raw-keyed, so two profiles
registering the same task id (docker_image / cwd) still collided. Writer, clear
and both readers (_has_isolation_overrides, resolve_task_overrides) now use
_qualify_task_key; the isolation-override branch of _resolve_container_task_id
returns the qualified key so the env cache and the override record agree.
- gateway/platforms/api_server.py::_derive_chat_session_id prefixed the seed for
the NAMED LAUNCH profile addressed via /p/<launch>/, so prefixed and
un-prefixed requests on one profile derived different ids. The prefix is now
skipped when the routed name is the process's launch profile
(_names_launch_profile vs get_routing_process_hermes_home()).
Tests (red on 3720519199c, green here):
tests/tui_gateway/test_session_cwd_profile_key.py::test_secondary_profile_session_cwd_is_found_inside_its_scope
tests/gateway/test_api_server.py::TestDeriveChatSessionId::test_launch_profile_prefix_keeps_the_unprefixed_id
Fold the two salvaged fixes to the class and the salvage bar:
- tools/terminal_tool.py: `_qualify_task_key` reuses `_routed_home_task_key`
instead of a duplicate qualifier; branch 2 (session-isolated sandboxes) AND
branch 3's `session:<key>` (SSH-style backends) carry the routed profile, so a
session id two profiles share (header-less API fingerprint, a DM chat id served
by two bots) never resolves to one `_active_environments` slot. Cwd records
keep the symmetric qualification. No routed home → raw key, unchanged.
- gateway/platforms/api_server.py: `_derive_chat_session_id` namespaces only
named routed profiles, so default/standalone ids are byte-identical and live
Open WebUI conversations survive the upgrade.
- Tests trimmed to two invariants: A→B→A across two temp profile homes under
multiplex (distinct sandbox keys, alias follows the qualified parent, cwd
isolated per profile), and derived-id namespacing without moving default.
Both fail on 5912ed81ed, pass here.