Commit Graph

5659 Commits

Author SHA1 Message Date
Austin Pickett
bf84561b9d fix(tui_gateway): keep a container backend's terminal.cwd as the session workspace
Under a container terminal backend (docker/modal/...), terminal.cwd names a path
inside the sandbox (/workspace). _completion_cwd host-validated it with
os.path.isdir, failed, and fell back to the gateway's own cwd ($HOME for the
desktop). Every host file then looked "inside the workspace", so file.attach
never staged it into the bind-mounted attachments/ dir and the agent was handed
a host path that does not exist in the container.

- _completion_cwd keeps an absolute container cwd unvalidated when the bound
  backend is non-local (matching _terminal_task_cwd), including a named
  profile's own container terminal.cwd and a config-only (unbridged) backend.
- _session_is_local_backend reads the effective backend (env or config), so a
  container cwd is never "healed" to its nearest host ancestor (/).
- map_cache_path_to_container also matches through a symlinked HERMES_HOME:
  @file: expansion passes resolved paths, the mount roots keep the configured
  spelling, so staged attachments leaked their host path.

Covers the desktop attachment path of #103147. Typed @file: refs to binary
files outside the mounted dirs still render the host path.
2026-09-28 19:41:57 -04:00
kshitijk4poor
da2b64304b fix(lsp): no automatic trust for scheduled work; lock-free trust lookups
Re-review of the trust-anchor follow-up:

- Cron runs and Kanban workers no longer anchor trust. The cronjob tool
  lets the model pick a job's workdir, which becomes the run's session
  cwd, and kanban_create lets it pick a task workspace, which becomes the
  worker's launch dir and TERMINAL_CWD. Either one let an agent-cloned
  repo become trusted. Both now rely on lsp.trusted_workspaces only.
- A repository at or above $HOME never anchors, not only one at $HOME
  exactly.
- Operator roots are recorded only where a tool thread enters
  (enabled_for) and when the service is created. _trusted() is now a pure
  read. Before, it took _state_lock while two of its callers already held
  it, so a change in the roots could deadlock the LSP loop.
- Client lookups go through _live_key(). A multi-root client that
  started while its root was untrusted is still found and released after
  the root becomes trusted, instead of being orphaned.
- `hermes lsp status` stops listing roots that have since become trusted.
  The lint gate reads the cwd the same way the linter subprocess does.
- The path-list parser returns None for a malformed value, and each key
  logs its own warning. The _node_modules_trees docstring no longer
  mentions Vue.
2026-09-29 03:28:46 +05:30
kshitijk4poor
e929702021 fix(lsp): trust the workspace the session was opened in, not only the process cwd
The trust anchor was os.getcwd(), which is the user's project only for a
plain `hermes` launched inside a repo. `hermes -w` sets TERMINAL_CWD to its
new worktree without a chdir, the Desktop/TUI backend runs from the install
tree or $HOME with the project bound as the session cwd, the gateway runs
from HERMES_HOME, and cron binds a per-job workdir. All of them lost
pyright's project venv, rust-analyzer/gopls/jdtls/... and the .ts/.rs lint
fallbacks inside the user's own repository.

operator_workspace_roots() now adds the worktree of resolve_agent_cwd()
(session cwd, then TERMINAL_CWD) to the launch dir's; only surfaces set
those, the agent's `cd` moves the terminal env's cwd, not them. $HOME is
never an anchor: a dotfiles repo there would trust every directory below it.
The service remembers every anchor it has seen, because the loop thread
that spawns servers has no session context of its own.

- The lint gate checks only the linter's cwd: npx and rustup resolve the
  toolchain from there (subprocess cwd=env.cwd), not from the file's dir.
  The try/except that wrapped code which cannot fail is gone.
- Trust is computed once per spawn and handed to ServerContext.
- Multi-root servers (pyright) share one process across trusted roots
  only; an untrusted root gets its own, so whichever root spawned first no
  longer decides the interpreter for the others.
2026-09-29 03:28:46 +05:30
John Paul Soliva
59ec281140 fix(lsp): deny by default in an untrusted workspace and skip the toolchain lint fallbacks there
Pinning settings per server left every other server free to run project
code: rust-analyzer runs cargo check (build.rs, proc-macros) on each
didSave despite the disabled init options, and jdtls, kotlin-language-server,
elixir-ls, zls, clojure-lsp and haskell-language-server evaluate build
files when they start.

In an untrusted workspace only UNTRUSTED_SAFE_SERVERS now start (pyright,
typescript, vue, svelte with their pinned settings; bash, yaml,
dockerfile, intelephense, clangd without --query-driver). Every other
built-in or user-declared server is skipped at the spawn chokepoint and
in enabled_for, logged once per root and listed by hermes lsp status.
The rust-analyzer init tweak is gone: it never starts untrusted now.

On a local backend the npx tsc and rustfmt --check fallbacks are skipped
the same way, since npx resolves the repo's node_modules/.bin/tsc (or its
.npmrc registry) and rustup honours its rust-toolchain.toml.
2026-09-29 03:28:46 +05:30
kshitijk4poor
2ec7703012 refactor(bot-mode): one helper names the live-delivery intent file
The boot-failure report added a fourth hand-written <dm>.live.json join; it
must agree with the path _admit_live_dm pins, so all four sites now share
_live_intent_file.
2026-09-29 03:01:04 +05:30
kshitijk4poor
9979e56cc1 fix(bot-mode): a Windows relaunch never adds a second delivery outcome
On Windows, hermes_bootstrap relaunches onto the store interpreter with
subprocess.call and exits with the child's status. The child is the same
delivery runner and has already printed its own outcome, so the parent
treating a nonzero status as "could not activate" printed a second JSON
line that could turn a definite failed/rejected into "ambiguous, do not
resend". The relaunch now raises RelaunchExit (a SystemExit subclass, so
every existing handler and the process exit code are unchanged) and the
runner skips its boot-failure report for it.

Also parses the runner argv in one stdlib-only helper, _runner_argv, used
by both _delivery_main and the boot-failure report, so the two can no
longer disagree about where the DM file is.
2026-09-29 03:01:04 +05:30
kshitijk4poor
be0a4147aa fix(bot-mode): only a live-admitted DM stays ambiguous when the runner cannot boot
Narrows the previous commit to the one case where the outcome is really
unknown. When the sender has already pinned a live intent (<dm>.live.json)
it may have admitted the DM and told the model not to resend; a runner
that then fails hermes_bootstrap now prints the same ambiguous payload
(same delivery id, "Do not resend") that _run_delivery prints for an
admission failure. Without an intent nothing was handed over, so the plain
repair hint stays the truth and a resend remains correct.

The payload is shared with _run_delivery via _live_outcome_unknown instead
of a second classifier, and the evidence-probing variants and the separate
test module are dropped in favour of one real -I -S subprocess test plus an
explicit negative assertion on the existing no-intent test.
2026-09-29 03:01:04 +05:30
Neomail2
875635bb18 fix(bot-mode): preserve delivery uncertainty on bootstrap exit
(cherry picked from commit 5fc1e15e4ddec41a388933a2cc77d786d1039a50)
2026-09-29 03:01:04 +05:30
Brooklyn Nicholson
be9fd2a691 fix(bot-mode): keep the reply waiter off dependency activation
The --wait-reply waiter is stdlib only and its stdout is the sender's
wake-up. Booting it through hermes_bootstrap meant an install whose
committed environment cannot be activated killed the waiter with the
repair hint on stderr before it printed anything, so the sender never
learned the delivery outcome. Only the delivery lane boots now.

Port the failure-mode test from #125754: a --run-delivery runner that
cannot activate exits with the `hermes pm repair` remedy and leaves the
DM untouched, instead of the ambiguous "No module named ..." outcome.

Co-authored-by: woriwka-ai <247827632+woriwka-ai@users.noreply.github.com>
2026-09-29 03:01:04 +05:30
Jervaise
6c9420706e fix(bot-mode): the delivery runner and reply waiter activate Hermes dependencies when spawned on the bare store interpreter
message_agent spawns its background runner as
`<sys.executable> tools/bot_mode_dm.py --run-delivery ...`, and the relay's
reply waiter (bot_relay.waiter_command) as `... bot_mode_dm.py --wait-reply ...`.
Both enter through the shared `if __name__ == "__main__":` block, which only
put the repo root on sys.path. Under a PM-managed install sys.executable is the
bare store interpreter: dependencies are activated in-process at boot by
hermes_bootstrap (pm.environments.activate_dependencies), and the terminal
backend strips the Hermes-owned PYTHONPATH, so the child has no third-party
packages. Live, every local Bot Chat DM returned
`{"status": "ambiguous", "error": "Live admission outcome unknown: No module
named 'ruamel'. Do not resend."}` from _run_delivery -> _admit_live_dm.

The __main__ block now imports hermes_bootstrap after putting the repo root on
sys.path (run as a script path, sys.path[0] is tools/). This is the same pattern
as tui_gateway/compute_host.py and agent/transports/hermes_tools_mcp_server.py:
only the script entry bootstraps, so importing tools.bot_mode_dm as a library
stays side-effect free and the module's top-level imports stay stdlib-only.
hermes_bootstrap rather than a bare activate_dependencies() call, because the
runner is a real entry point and needs what every entry point gets: dependency
activation plus early recovery of an interrupted update, the Windows UTF-8 stdio
bootstrap (its stdout carries the teammate's reply as the completion
notification), and import-path hardening against a cwd package shadowing
Hermes modules. The runner argv is safe for the bootstrap: command_argv() reads
`--run-delivery`/`--wait-reply` as flags, so it is neither `update --post-swap`
nor `pm repair`. prepare_launch acts only on a self-updated install whose
dependencies are behind or whose interpreter is not the store Python, and its
relaunch re-runs the same script path with the same argv (runpy.run_path), so
the runner keeps its contract there too.

tests/tools/test_bot_mode_dm_entry.py launches the real script under `-I -S`
with a PM-committed dependency record in a temp HERMES_HOME as the only route to
the packages: the --run-delivery case fails on the parent commit with the exact
live `No module named 'ruamel'` error and passes here; the --wait-reply case
drives the argv waiter_command builds through the same entry.
2026-09-29 03:01:04 +05:30
Teknium
9bcbe7b5df feat(i18n): pluggable, layered language packs across core, Desktop and TUI (#126296)
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales

* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter

ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.

* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs

* chore(tui): split en catalog siblings by lane (slash sibling)

* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports

* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts

* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list

* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes

* feat(plugins): report language-pack layers in the mid-run activation summary

* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n

StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.

* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter

- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
  the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
  layers partial packs (nested or flat dotted) over bundled/en via
  mergeTranslations; a string over a function-valued en entry becomes a
  positional {0}/{1} formatter; $appLocaleVersion bumps so translators
  re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
  registry as source 'backend' (method-not-found is silent); re-synced on
  socket open, display.language change and profile switch. A saved pack-only
  language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
  tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
  from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).

* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)

Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.

* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()

- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
  ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
  accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
  properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces

* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)

* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()

Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.

locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).

* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()

- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions

* i18n(platforms): route Google Chat and Teams user-facing text through t()

Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.

Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.

* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()

LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.

* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)

Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.

* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)

_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.

* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)

RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.

* i18n(cli): live-work dock, subagent monitor and render copy through t()

cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.

* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()

* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)

get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.

* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output

* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time

Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).

* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)

cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.

* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()

- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
  at call time; category labels via slash.category.*; help/alias/usage suffixes
  via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
  bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
  lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
  status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
  approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
  bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
  heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
  sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
  keeps its English constant; callers use history_unreadable() ->
  gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.

* i18n(telegram): route adapter chat copy through t()

Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.

Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.

* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()

- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor

* i18n(discord): route adapter chat copy through t()

Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().

* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)

* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()

- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
  existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
  via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
  _DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
  Updating/Generating) are one full template per variant; plurals use
  <key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
  translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).

* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English

* test(i18n): pin Telegram/Discord adapter catalog wiring

Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.

* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge

* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint

* i18n(tr): translate bundled catalog + tui pack

* i18n(ja): translate bundled catalog + tui pack

* i18n(ko): translate bundled catalog + tui pack

* i18n(zh): translate bundled catalog + tui pack

* i18n(fr): translate bundled catalog + tui pack

* i18n(af): translate bundled catalog + tui pack

* i18n(uk): translate bundled catalog + tui pack

* i18n(ar): translate bundled catalog + tui pack

* i18n(pt): translate bundled catalog + tui pack

* i18n(it): translate bundled catalog + tui pack

* i18n(es): translate bundled catalog + tui pack

* i18n(zh-hant): translate bundled catalog + tui pack

* i18n(ru): translate bundled catalog + tui pack

* i18n(hu): translate bundled catalog + tui pack

* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)

* i18n(de): translate bundled catalog + tui pack

* i18n(ga): translate bundled catalog + tui pack

* test(i18n): fixture matches _normalize_lang(lang, home) signature

* i18n(tui): scaffold userMessages/slashCmd en siblings

* i18n(tui): wire secure prompts + content tables

* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays

Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.

* i18n(tui): wire slash ops/wake replies

* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)

* i18n(tui): wire slash core/debug/setup replies

* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)

* i18n(tui): wire slash session/topup/subscription replies

* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)

* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys

* i18n(tui): wire userMessages copy through the userMessages namespace

* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json

* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)

Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).

* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)

* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)

* tui: i18n-export-en script (English templates for pack translators)

* docs(i18n): bundled TUI packs are the bottom layer of the tui surface

* i18n(ru): translate TUI pack

* i18n(ar): translate TUI pack

* i18n(es): translate TUI pack

* i18n(pt): translate TUI pack

* i18n(ko): translate TUI pack

* i18n(de): translate TUI pack

* i18n(ja): translate TUI pack

* i18n(fr): translate TUI pack

* i18n(tr): translate TUI pack

* i18n(it): translate TUI pack

* i18n(zh): translate TUI pack

* i18n(zh-hant): translate TUI pack

* i18n(hu): translate TUI pack

* i18n(uk): translate TUI pack

1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.

Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).

* i18n(ga): translate TUI pack

* i18n(af): translate TUI pack

* plugin_guard: locale catalogs in language packs step down the agent-config family

A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.

* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)

* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header

Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.

* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel

- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
  and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
  agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)

* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})

* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property

* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal

* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export

The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.

* test: unbreak two main-red timing tests the PR merge-ref inherits

- test_local_runtime racing fake publishes the modern state record (legacy pid-only
  records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
  a fixed 0.35s, which a loaded CI runner does not always meet

* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)

* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak

SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.

* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings

* chore(i18n): regenerate desktop key catalog after main sync

---------

Co-authored-by: Teknium <teknium@nousresearch.com>
2026-09-28 14:16:18 -07:00
kshitijk4poor
a0e20b60d9 fix(child-env): a partial plugin scan keeps the denials it already knew
andrexibiza (review on 60bdd5fbe3): skipping an unreadable plugin dir,
or an unreadable or unparsable flat manifest, replaced the home's
cached secret set with the partial result. A secret declared a moment
ago could then reach children while its manifest was unreadable. And
because the partial result was cached under the same file signature,
recovery could keep hitting it.

platform_manifest_secret_scan() now reports whether the scan was
complete. A partial scan unions in the names this home already had,
and is cached with no stamp, so the next spawn rescans and a recovery
is seen at once. A deleted plugin still releases its names, because
that scan is complete. An unreadable plugins/platforms/ manifest still
raises.
2026-09-29 02:06:58 +05:30
kshitijk4poor
9e26d34475 fix(child-env): manifest scan skips only what plugin discovery skips
Per-commit review follow-ups:

- Dot directories are scanned again, matching plugins_discovery. A
  platform plugin that discovery loads from plugins/.x now also has
  its secrets declared.
- Only PermissionError means a plugin directory cannot be searched. A
  symlink loop is treated as missing, so plugin.yml is still checked,
  as discovery does. A plugin.yaml that is not a regular file is
  ignored rather than failing every strict read.
- Dropped a redundant union: the home set is already stripped in
  Tier 1.
- Tests now pin the dunder skip, the unlistable platforms root, and a
  stamp change on an edit that keeps the mtime.
2026-09-29 02:06:58 +05:30
kshitijk4poor
b1206115a2 perf(child-env): one scandir manifest stamp per scrub, file-signature keyed
Per-commit review follow-ups on the per-home plugin declarations:

- The manifest stamp walked each plugin directory with pathlib
  is_dir/exists/stat on every spawn, and _scrub_credentials took it
  twice. It is now one scandir per root and one stat per candidate,
  taken once per scrub.
- The stamp uses utils.file_signature, so a manifest replaced with its
  mtime preserved (cp -p, rsync -t) still invalidates the cache.
- An unsearchable plugin directory, including an unlistable flat
  plugins/ root, is reported once as a directory entry. The stamp walk
  stays quiet; the warning comes from the manifest read, which runs
  only on a cache miss.
- The dot/dunder skip now matches plugins_discovery.scan_directory, and
  the manifest source argument is a Literal.
- Tests: the unsearchable-dir case uses a non-dunder dir, so both
  guards are pinned. New tests cover the bundled read failing closed,
  manifest deletion releasing its names, and a symlinked home alias
  sharing one cache entry.
2026-09-29 02:06:58 +05:30
kshitijk4poor
8ffcd7488f fix(child-env): openviking drops the profile overlay's bot tokens; unprovable manifests don't block spawns
Phase 2 review follow-ups:

- OpenViking overlaid the bound profile's whole .env after the scrub,
  so that profile's bot, dashboard and relay tokens reached the server
  again. The overlaid env is now scrubbed a second time for Tier 1.
  Provider keys still pass, and they still come from the bound profile.
- The strict per-home manifest read raised on any unreadable entry
  under <home>/plugins, which failed every child spawn for that
  profile. It is now strict only where the manifest is known to be a
  platform's: bundled or plugins/platforms/. Dunder and dot
  directories, unsearchable plugin directories and unreadable flat
  plugins/* manifests are skipped with a warning, because they can't
  load as plugins either.
- The per-home cache is keyed on hermes_home_key(), so a symlinked
  alias no longer gets a second entry.
- The modal snapshot test swapped hermes_cli.config for a stub, which
  the policy import can no longer use. HERMES_HOME already points the
  real module at the test home, so the stub is gone.
- The env_passthrough docstring now describes declared names rather
  than the dropped prefix rule.
2026-09-29 02:06:58 +05:30
kshitijk4poor
1ede0b805b refactor(child-env): one secret-suffix rule, one env-table walk, one bundled manifest read
/simplify-code follow-ups on the declared-secret policy:

- The policy and the manifest classifier each defined which env-name
  suffixes make a platform variable a secret, and they disagreed on
  _JSON. PLATFORM_SECRET_ENV_SUFFIXES in hermes_cli/config.py is now
  the only definition.
- The gateway env-override table was walked by hand a fourth time,
  through private gateway.config_env names. The walk is now
  profile_channels.config_env_table_keys(), shared with
  declared_channel_env_keys.
- The bundled platform manifests were parsed twice at import (about
  16 ms). The config injection reads them once, strictly, and keeps
  their secret names. The policy re-reads only if that read hit an
  I/O error.
- The per-home cache was keyed on directory mtimes, so an in-place
  plugin.yaml edit went unseen until restart. It is now keyed on
  every manifest file's mtime.
- The core-declared snapshot is a module-level constant rather than
  a global set inside the injector. The two manifest-source booleans
  are now one source argument. The Tier 1 set is upper-cased once
  at import, not on every spawn.
2026-09-29 02:06:58 +05:30
kshitijk4poor
3a6406137d fix(child-env): per-profile plugin secrets, fail-closed declaration scans, dashboard auth in Tier 1
Review follow-ups on the declared-secret policy:

- A profile's user-installed platform plugins declared secrets into
  one process-wide set, read from the launch home at import. Under
  multiplex that stripped profile B's same-named user variable and
  missed profile A's own declarations. Bundled manifests stay
  process-wide; user manifests are now read per bound home, cached
  until that home's plugin dirs change, and are Tier 1 for that
  profile only.
- The declaration scans no longer fail open. A registry error is no
  longer swallowed into an empty set, non-string required_env entries
  are skipped rather than breaking the set, and a manifest that
  cannot be read raises instead of vanishing from the policy. A
  malformed manifest still declares nothing.
- A plugin manifest can no longer reclassify a core-declared name
  such as OPENAI_API_KEY; the same rule the config form applies.
- The dashboard basic-auth password and signing secret, the OIDC
  client secret and the drain bearer move to Tier 1. Credentialed
  CLIs (claude, codex) no longer receive them.
- openviking-server starts from served_profile_child_env, so it gets
  the bound profile's provider keys, never the launch profile's.
  With no bound profile under multiplex the start is refused.
- NOUS_API_KEY and QWEN_API_KEY are listed statically. Discovering
  provider plugins while the policy module imported re-mirrored them
  over a plugin's own auth registry entry
  (tests/providers/test_auth_registry_import_order.py).
2026-09-29 02:06:58 +05:30
kshitijk4poor
43800d55f2 fix(child-env): the bot desktop and gmail-triage scripts start from the scrubbed env
Two more spawns inherited the gateway's full environment. The bot
desktop launcher copied os.environ minus the display variables; the
agent drives that desktop and its dock opens xfce4-terminal, so every
provider key and bot token was one click away. The Telegram gmail-triage
buttons ran HERMES_HOME/scripts/gmail-triage/*.sh with no env at all, a
path the agent can write. The desktop now builds from
served_profile_child_env() like the agent browser (keeping the user's
HOME); the triage scripts use build_subprocess_env like cron and
webhook-filter scripts.
2026-09-29 02:06:58 +05:30
kshitijk4poor
6c2a53367f fix(terminal-env): block Hermes' own dashboard, anon and Meet secrets
Review of the adapter-secret salvage found Hermes-owned secrets that no
declared source covers, so they reached terminal and execute_code
children and skill passthrough accepted them: the dashboard basic-auth
password and signing secret (enough to forge dashboard sessions; the
session token beside them is already Tier 1), the dashboard drain and
OIDC client secrets, HERMES_ANON_API_SECRET (a provider-category entry,
which the OPTIONAL_ENV_VARS loop never blocks) and the Google Meet
realtime key. Add them to the static blocklist.
2026-09-29 02:06:58 +05:30
kshitijk4poor
ee942a7bbf fix(terminal-env): adapter secrets are the declared ones, not any platform-named variable
The shape rule (a platform prefix plus _TOKEN/_SECRET/_PASSWORD/_KEY) also
matched variables Hermes never reads: the platform list holds plain words
(LOCAL, GATEWAY, WEBHOOK, SLACK), so a user's SLACK_USER_TOKEN,
LOCAL_LLM_API_KEY or GATEWAY_API_KEY vanished from the terminal and
terminal.env_passthrough could not bring them back. The prefix census it
leaned on also failed open (an unreadable plugins dir cached an empty set).

Adapter secrets now come from what adapters declare: password entries of
the messaging OPTIONAL_ENV_VARS (built-ins plus every platform plugin
manifest) and secret-named keys of the gateway env-override table, both
Tier 1 and refused by passthrough; plus the secret-named required_env of
adapters registered in the current profile scope, read per spawn without
loading deferred adapters, Tier 2 only because required_env is an
unchecked setup list. Secrets nothing declared get a manifest entry
(TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN, A2A_PUSH_SECRET,
TEAMS_GRAPH_ACCESS_TOKEN, TEAMS_INCOMING_WEBHOOK_URL) or join the
policy's read-in-code list (QQ_STT_API_KEY, the two MSGRAPH names).
2026-09-29 02:06:58 +05:30
John Paul Soliva
6abcf4d86a fix(terminal-env): a passthrough accepted before an adapter owned the name stops forwarding it
Skill and config passthrough names were checked against the managed-credential policy only when accepted. A plugin adapter registering later claims <PREFIX>_*_SECRET, but is_env_passthrough() and get_all_passthrough() kept returning the stale approval, so terminal, background and execute_code children (and scope-only additions) still received the secret. Both now re-apply the refusal when the allowlist is consumed.
2026-09-29 02:06:58 +05:30
John Paul Soliva
dca8684c59 fix(terminal-env): adapter secrets are Tier 1, owners are read per call
An inheriting child (claude/codex/gemini) kept TELEGRAM_WEBHOOK_SECRET, WHATSAPP_CLOUD_ACCESS_TOKEN and the Graph secrets because the shape rule sat only in the provider tier; it now strips with the bot tokens. The shape rule reads the per-call adapter census, so a plugin adapter registered after import, in the bound profile only, is covered. The OAuth provider scan takes bundled profiles only, so the process-wide blocklist no longer freezes the home bound at import.
2026-09-29 02:06:58 +05:30
John Paul Soliva
66019e5597 fix(terminal-env): strip adapter secrets and OAuth-profile keys from child envs
The child-env blocklist is derived from the provider registry and
OPTIONAL_ENV_VARS, but many gateway adapters read their secrets straight
from the environment without listing them there: WHATSAPP_CLOUD_ACCESS_TOKEN,
WHATSAPP_CLOUD_APP_SECRET, WEIXIN_TOKEN, YUANBAO_APP_SECRET, FEISHU_ENCRYPT_KEY,
TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN and others. They reached
terminal, background/PTY, cron-script and hermes_subprocess_env children,
and a skill could register them as env passthrough, while the documented
bot tokens next to them were stripped. OAuth provider profiles (nous,
qwen-oauth) also accept a pasted key (NOUS_API_KEY, QWEN_API_KEY), but the
registry mirror copies env_vars only for api_key profiles, so those passed
through too.

Match adapter secrets by shape, the way authorization gates already are: a
built-in or bundled adapter prefix plus a _TOKEN/_SECRET/_PASSWORD/_KEY
suffix, so a new adapter secret is covered without a second edit. Add every
provider profile's env_vars regardless of auth_type, and list the two
Microsoft Graph secrets no adapter prefix owns.
2026-09-29 02:06:58 +05:30
teknium1
6cc2f4ab92 fix(telemetry): terminal.outcome=timeout comes only from Hermes' deadline
_run_foreground recorded outcome=timeout whenever a backend exception's
text contained "timeout" (an SSH connect timeout, a sandbox API timeout).
That command never reached an exit status, and the doc says timeout comes
only from Hermes' own deadline flag (hermes_timed_out), which
terminal_outcome() already reads. Owner ruling: never from exception text.
The exception path now records nothing; the model-facing result (exit 124)
is unchanged.

Probe (code-read finding; the new test drives the real terminal_tool with a
backend raising "... Connection timeout"):
  before: terminal.outcome {backend: local, command_kind: shell_builtin, outcome: timeout} x1
  after:  no terminal.outcome row; tool result still exit_code 124
test_backend_exception_mentioning_timeout_is_not_a_terminal_timeout is RED on
base, green here.
2026-09-28 12:43:03 -07:00
teknium1
c13a4abc8d feat(telemetry): v5 signals — tool_unavailable, provider_setup, feature_adoption, feature_disabled
- hermes.tool_unavailable.count: model called a shipped built-in (BUILTIN_TOOL_NAMES) not enabled in
  the session; tool name + catalog provider/model. Unknown names stay in v4 unknown_tool quality.
- hermes.provider_setup.count: started/completed/abandoned/failed per provider and surface
  (cli_setup, cli_model, tui, desktop, dashboard) with a closed failure_class. Pending marker per flow;
  a dead or stale marker is reported abandoned at the next start (v4 process-marker pattern).
- hermes.feature_adoption.count: once per feature per install at first real use, derived from existing
  counters in the subscriber (metric -> feature table) plus Bot Mode / Projects hooks; bucketed by the
  owning profile's install age.
- hermes.feature_disabled.count: turning off a default-on toolset/skill/plugin/platform/setting (and the
  re-enable back to default), diffed at the config write chokepoints; once per (kind,name,event)/day.
  Surface comes from the user entry point; setup/migrations record nothing.
2026-09-28 12:43:03 -07:00
teknium1
8d8836ddb1 feat(telemetry): harness-accuracy shared metrics for the agent loop
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):

- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
  matcher reports the strategy that landed (or no_match / ambiguous) into a
  context-local probe opened only around the patch/write_file handlers, so we
  learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
  the guardrail already warns/blocks/halts and at turn end for the iteration
  budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
  next_outcome. One row per failed tool call, resolved against the model's
  next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
  a table lookup on the first program word, never the text. Hermes' own
  deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
  backend result, so a command's own `exit 124` reads nonzero, not timeout.
  Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
  truncation only from the structured finish reason; empty/reasoning_only
  from the normalized message; issue=none per response as the denominator
  (model_route counts attempts and files empty replies as failures).

Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
2026-09-28 12:43:03 -07:00
teknium1
25c1b008c8 feat(telemetry): wire memory, curator, delegation and backend call sites
- memory tool: one row per operation (a batch counts each op); gate/validation
  refusals are "rejected", store-level non-application "failed". Memory provider
  tools are counted in MemoryManager.handle_tool_call with the op read from the
  action arg or tool-name verb.
- curator: run_curator_review records one row per pass (dry run = skipped) with
  the before/after diff bucketed; a scheduled pass whose claim is held by another
  process counts as skipped. The home is captured before the review thread starts.
- delegate_task: _run_batch opens the call, each joined unit folds its results in,
  and the last unit emits the call's single row, so group-split background calls
  are not counted per unit.
- terminal / execute_code / browser: counted per call that reached the backend;
  the backend is resolved from the owning profile at call time (terminal plan
  env_type, code local/remote, browser CDP > Camofox > cloud provider > engine, or
  "extension" when the extension controller served the call). Guard refusals and
  Hermes' own _host_local commands are not counted.

Tests read rows back from the real store; each call-site group is red with its
source reverted.
2026-09-28 12:43:03 -07:00
teknium1
98151e7cb7 feat(telemetry): record MCP server installs as extension-install events
Every path that adds an MCP server now emits one
hermes.extension.install.count event through record_extension_install:

- catalog installs (hermes mcp install, the picker, the dashboard route and
  its background CLI action) via mcp_catalog.install_entry
- catalog installs from connector cards and the agent's catalog tool via
  _CatalogBackend.install / start_install_oauth (OAuth: success at commit,
  failed when the flow cannot start; an abandoned browser step is a cancel)
- custom servers from `hermes mcp add`, POST /api/mcp/servers and the
  mcp.add RPC (source url for http, local for stdio, name None)

A reinstall or overwrite of an already-configured server is not counted,
so the metric measures new installs rather than config churn. Catalog
names pass through raw; the contract reports non-catalog names as custom.
2026-09-28 12:43:03 -07:00
Teknium
7c799a6565 feat(vercel): fresh sandboxes use a managed image (universal:latest); runtime presets deprecated, migration 49 (#126741)
* feat(vercel): start fresh sandboxes from a managed image instead of the deprecated runtime

Vercel deprecated Sandbox runtimes (node24/node22/python3.13) in Aug 2026 in favour of
images, and rejects runtime+image together and runtime with a snapshot source. New
terminal.vercel_image (default vercel/sandbox/universal:latest, Node 24 + Python 3.14)
picks the image for fresh sandboxes; a pinned terminal.vercel_runtime still works, wins
over the image and logs a deprecation warning; snapshot restores send neither.

Setup wizard prompts for the image, dashboard exposes both keys, status/config show the
effective choice, TERMINAL_VERCEL_IMAGE bridges config to the tool like its siblings.

* feat(config): migration 49 drops the seeded node24 Vercel runtime pin

Every pre-49 config.yaml carries terminal.vercel_runtime: node24 (the template default) and
the setup wizard mirrored it into .env as TERMINAL_VERCEL_RUNTIME. Both are the default
copied, not a choice, so the migration drops them and fresh sandboxes follow vercel_image;
node22 / python3.13 pins are the user's and survive. Persisted sandboxes are unaffected:
a snapshot restore never sends a runtime or an image.
2026-09-28 12:33:01 -07:00
Teknium
3f871425af fix(modal): persistent sandbox snapshots no longer expire after 30 days (modal 1.5.5, ttl=None) (#126740)
* fix(modal): keep persistent-sandbox snapshots past the SDK's 30-day TTL

modal>=1.5 gives Sandbox.snapshot_filesystem() a default ttl of 30 days, so an idle
persistent Modal sandbox silently lost its filesystem and restarted from the base image.
Pass ttl=None (retain until deleted) and bump the modal extra from 1.3.4 (no ttl
parameter; legacy RPC) to 1.5.5 so the kwarg exists on every install.

* fix(modal): drop the dead modal.Mount credential-mount block

modal.Mount left the public API in modal 1.0, so _modal.Mount.from_local_file raised
AttributeError into the surrounding except on every sandbox start and the block never
mounted anything. The FileSyncManager created right after already uploads the same
credential, skills and cache files (iter_sync_files), so delete the duplicate; the test
fake stops exporting a Mount the real SDK does not have.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 12:31:03 -07:00
teknium1
a2cdaa7388 fix(review): run TTS warm/release command hooks under the caller's secret scope (#120481)
The warm_command/release_command hook thread was a bare threading.Thread, so under
the multiplexer resolve_passthrough_value() saw no secret scope, raised
UnscopedSecretError, and the best-effort hook swallowed it at debug level: every
command provider with a non-empty env_passthrough silently never warmed or released.
Bind the thread with ctx_bound (copy_context().run), the same seam the keep-warm
timer already uses. Regression test: the hook thread resolves the bound profile's
passthrough value.

Docs (review minor): the multiplex isolation table now names per-profile slash-command
gating and the fail-closed empty-admin policy for a served profile with no cached config.
2026-09-28 12:23:25 -07:00
John Paul Soliva
1df4af417f fix(tts): resolve command-provider env_passthrough through the profile secret scope
run_command_provider forwarded declared env_passthrough keys from os.environ,
which under the multiplexer is the LAUNCH profile's .env. A served profile's
command TTS/STT subprocess (voice-note transcription, speech, warm/release
hooks) got the launch profile's credential and never its own. Resolve each key
with resolve_passthrough_value, as the terminal and code-execution spawns do.

(cherry picked from commit 22a299d5f562e1d174fa5d37590cd46361a1383a)
2026-09-28 12:23:25 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
teknium1
3e1bfb2935 fix(execute_code): report in-script tool errors the script ignored
A script that calls hermes_tools.write_file/patch and never prints the
returned {"error": ...} got status=success with empty output, so the
model believed the write happened. In a cross-agent bench this cost 3 of
18 Opus runs 4-6 extra turns (write_file refused by the read-before-write
guard, then re-probing and re-emitting whole files). The result now
carries tool_errors for that cell, and the in-script write_file doc says
existing files must be read first.
2026-09-28 10:11:33 -07:00
funky-xamarin
7a09408e7c fix(execute-code): preserve PM runtime dependencies for own kernel 2026-09-28 06:42:00 -07:00
alt-glitch
fc4dbe32df fix(logs): hermes logs --since/--level handle unstamped lines and MCP output
`hermes logs --since` and `--level` passed every line that had no leading
timestamp. A traceback's frames are written without one, so an old error
printed its frames without the header that was filtered out, and
`--component` dropped the frames of a matching record. `hermes logs` now
reads each unstamped line as part of the record above it: the line gets
that record's verdict for every filter, in the tail read and in `-f`.
Lines before the first stamp in the read window have an unknown time and
level, so they are dropped when `--since` or `--level` is set.

mcp-stderr.log had no parseable stamp at all: the banner started with
`=====` and server output was copied raw, so `hermes logs mcp --since`
printed the whole file. The stderr tee already reads each server's
stderr in a thread, so it now writes one line at a time, each prefixed
with the asctime-shaped local stamp the Python logs use. The banner
starts with the same stamp. The stamp comes from new public
`timestamp()`/`stamp_line()` in hermes_cli/stderr_timestamp.py, which
stays stdlib-only. The desktop MCP log view accepts both banner shapes.

A test now requires a real writer's sample line for every LOG_FILES
entry to parse with `_parse_line_timestamp`. The docs no longer say the
`timezone` key changes log timestamps; log lines use the machine's
local time.
2026-09-28 19:02:20 +05:30
teknium1
10708bc117 fix(bot-chat): rebuild a stale Bot Chat's tools through the surface builder (#124211, salvage #124307)
Trim the salvaged detector/delivery pair to the salvage bar and fix the
delivery half. The PR resolved the refresh selection as
platform_toolsets.<surface> (desktop/tui), a key no config path writes,
so _get_platform_tools fell back to the constructed hermes-<surface>
composite that resolves to 0 tools (a cold resume went 30 -> 1 tool).

Delivery now goes through the builder new desktop/TUI sessions use
(tui_gateway.server._load_enabled_toolsets(platform) +
_load_disabled_toolsets) as an explicit refresh_agent_mcp_tools
override, so the refreshed set equals a fresh session's on that surface;
only desktop/tui agents are long-lived, every other surface builds a
fresh agent per process/turn and is left alone. Drops the
toolsets.py/tui_gateway policy relocation and the raw-value fallback in
the fingerprint; keeps the detector (platform_toolsets +
agent.disabled_toolsets replace the dead tools.enabled_toolsets key)
and the tools re-pin after the refresh.

Tests: 2 invariants in tests/agent/test_bot_chat_toolset_refresh.py
(desktop -> builder consulted, disabled override passed, re-pinned;
cli -> untouched) and the existing fingerprint axis test now edits
the key `hermes tools enable/disable` writes.
2026-09-28 06:06:17 -07:00
finn763
fcda937c5d fix(bot-chat): deliver toolset changes to the canonical Bot Chat
The capability epoch watched tools.enabled_toolsets, a key no surface
writes; real hermes tools enable/disable edits (platform_toolsets.* plus
agent.disabled_toolsets) never flipped it. Even when stale, the refresh
rebuilt only the prompt and never tools[]. Watch the real keys and
rebuild + re-pin the tool snapshot on Bot Chat capability refresh.
Closes #124211
2026-09-28 06:06:17 -07:00
teknium1
d6c9191368 fix(review): replay the launch ledger regardless of the first consumer's profile scope (#123265)
Review finding (major): ProcessRegistry.restore_completions() is once-per-process but read
async_delegation._db_path() (ContextVar-aware) under whatever scope the FIRST consumer ran in.
In the TUI gateway both first consumers (the session notification poller and the prompt_turn
drain) run inside _session_profile_runtime_scope(session), so under multi-profile
`hermes serve` the first session's profile ledger was replayed and the LAUNCH profile's
undelivered completions were never replayed for the life of the process (origin/main's
import-time replay always covered the launch profile).

Fix: restore_completions() clears the hermes-home override for the duration of the replay
(set_hermes_home_override(None) + reset), so the once-per-process replay always reads the
launch ledger whoever gets there first; the caller's scope is restored afterwards. Smaller
than a per-home restored-set plus a tui_gateway boot hook: it restores exactly main's
invariant with no new boot seam, and secondaries stay where they were (gateway
_restore_secondary_completion_ledgers). Docstring states the invariant.

Tests:
- tests/tools/test_process_registry_lazy_restore.py::test_first_drain_under_secondary_scope_replays_the_launch_ledger
  (red on the PR head: replayed profiles/b/state.db; green now)
- tests/gateway/test_multiplex_unserved_shared_ingress.py::test_boot_replays_the_launch_ledger_before_secondaries_and_watchers
  (minor: pins the gateway boot hook — launch-scope replay in _start_secondary_profiles, i.e.
  before the secondary bind and before _async_delegation_watcher, the only production consumer
  that reads completion_queue.get_nowait() without drain_notifications)
2026-09-28 05:37:55 -07:00
webtecnica
34c2dc5cb1 fix(delegation): restore durable completions on first consume, not on import (#123265, salvage #123348, #123298)
WHAT
- tools/process_registry.py: ProcessRegistry.__init__ no longer calls
  restore_undelivered_completions(); a new once-per-process
  restore_completions() does, invoked by the first consumer:
  drain_notifications() (CLI process_loop / TUI prompt turn), the gateway
  startup (run_startup._start_secondary_profiles, right before the secondary
  ledgers are replayed) and the TUI session notification poller.
- tools/async_delegation.py: restore_undelivered_completions() returns 0 when
  <HERMES_HOME>/state.db does not exist (a replay must not create or migrate
  the ledger); _connect() creates the parent via mkdir_under_hermes_home()
  instead of a bare mkdir so a late writer cannot resurrect a deleted
  (tombstoned) or missing named profile.
- gateway/run_notifications.py: docstring no longer claims the import restores
  the launch ledger.

WHY
`process_registry = ProcessRegistry()` runs at module import, and model_tools
imports it transitively (tools.registry -> close_terminal_tool), so
`python -c 'import model_tools'` (hermes doctor, any tool-registry consumer)
under a fresh/typo HERMES_HOME created the full profile skeleton + state.db,
recreated a tombstoned `profiles/.deleted/<name>` profile, and ran
reconcile_state_schema() against an existing store — bypassing the
assert_named_profile_home_live / mkdir_under_hermes_home guards (#97128,
#112592). Restoring on first consume keeps the replay for every process that
actually drains the queue while an import performs no state.db I/O.

Salvaged from #123348 (@webtecnica: lazy restore + missing-ledger guard) and
#123298 (@Wenfengcheng: mkdir_under_hermes_home in _connect), trimmed to the
minimal shape (no completion_queue property, no read-only probe).

Fixes #123265

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
(cherry picked from commit b77c4ff49062d2936271ded80821b48b0c11d368)
2026-09-28 05:37:55 -07:00
teknium1
17957f5014 fix(review): ticker heartbeat proves liveness only while its writer is alive; routed hand-off never trusts an unknown probe; satellite manual runs get the ticker's routed grant
MAJOR — `_builtin_gateway_liveness` trusted a bare-epoch ticker heartbeat for
~200 s after its writer died, for every profile, and the multiplexer pid gate
had been dropped: a killed serve/Desktop ticker read as alive and
`hermes -p X cron run` queued a primary-routed run "for the gateway's next
tick" with no ticker present.

- `cron/jobs.py::record_ticker_heartbeat` stamps `<epoch> <pid>`;
  `get_ticker_heartbeat_age` parses the first field (legacy bare stamps still
  yield an age); new `ticker_heartbeat_writer_alive` requires the stamped pid
  to be alive and treats a bare stamp as NOT proof by itself.
- The heartbeat-only rung is now `fresh AND (served-by-multiplexer OR writer
  alive)`: the multiplexer record proves the host process for a served named
  profile (and covers a stale-code multiplexer still writing bare stamps), so
  only the in-process serve/Desktop ticker relies on the heartbeat, and then
  only with a live writer. `cron status`'s in-process-ticker rung applies the
  same rule.

MINOR (a) — `_hand_off_primary_routed_run` gated on `is True`; an unknown
(None) probe returns an error naming the uncertainty instead of "queued …
runs and delivers it".

MINOR (b) — `_run_claimed_job` wraps a shared-bot satellite's resolved map in
`SharedRouteAdapters(primary, _primary_profile_routes_for_current_home())`,
the same grant the ticker's `tick_adapters_for` makes, instead of handing it
the full primary adapter map.

Tests (each red on the previous head a18977090be):
- test_cron_satellite_diagnostics.py::test_in_process_ticker_heartbeat_counts_only_while_its_writer_lives
- test_cronjob_run_primary_routed.py::test_routed_run_without_a_serving_gateway_fails_before_the_turn[None]
- test_cronjob_run_immediate.py::test_execute_job_now_grants_a_shared_bot_satellite_only_its_routed_targets
2026-09-28 05:28:05 -07:00
Denis Dresvyanskiy
7e849e5548 fix(cron): hand a primary-routed satellite's manual run to the gateway ticker
A multiplexed satellite profile with no platforms.<p> credential of its own
posts through the primary's bot via a root gateway.profile_routes entry.
The delivery preflight lets such a job through (#97476) on the assumption
that the primary gateway's live adapters send it, which holds on a
scheduler tick but not for a manual run: `hermes -p <profile> cron run`
executed the whole agent turn in the CLI process, which has no sender for
the route, then failed delivery with "platform 'telegram' not
configured/enabled" and overwrote last_status with delivery_failed.

cronjob(action='run') now checks, before the in-process claim and before
the background dispatch, whether any delivery platform of a runnable job
is reachable only through the primary route (routed to this profile, no
connected credential here). If so:

- a gateway that serves the profile is live (or liveness is unknown):
  queue the run with trigger_job so that gateway's ticker runs and
  delivers it; job status is untouched and `cron run` prints "It will run
  on the next scheduler tick";
- no gateway serves it: fail fast before the agent turn, like the
  relay-fronted forward does when its api_server is unreachable.

Runs inside the gateway process (its live adapter delivers, #89302),
paused jobs (trigger_job would resume them), local delivery, and profiles
with their own credential keep the in-process path unchanged.

Fixes #120330

(cherry picked from commit 5a0db1c4e37de865f23034937b93d38fa87414f4)
2026-09-28 05:28:05 -07:00
teknium1
e790ef4e31 fix(cron): trim owner-adapter resolution comment to the why (#124248, salvage #116316) 2026-09-28 05:28:05 -07:00
Fabio Ito
7d7166be5e fix(cron): do not fall back to the default adapters when owner-profile resolution fails
Review feedback: the previous try/except restored runner.adapters on any
resolution error, which is the cross-profile misdelivery this fix
prevents. Let the error propagate to _run_claimed_job's handler, which
marks the run failed and surfaces the message. Runners without
_adapters_for_profile (shims, tests) still keep runner.adapters.

Adds a regression test where resolution raises: run not fired, marked
failed, error surfaced.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 183547ea6b3798874dc8a98f943ef98ee710911b)
2026-09-28 05:28:05 -07:00
Fabio Ito
8be50587d4 fix(cron): deliver manual runs through the owner profile's adapters under multiplex
A manual `cronjob(action="run")` fired from a secondary profile's agent
resolved delivery through `runner.adapters`, which is the default
profile's map. Under `gateway.multiplex_profiles` HERMES_HOME is
overridden to the owner profile mid-turn, so the result left through the
default profile's bot (a Telegram DM arrived from the wrong bot while the
run was correctly recorded on the secondary profile's cron).

Resolve the owner profile from HERMES_HOME and use
`runner._adapters_for_profile()`, the same fail-closed path notifications
and goal loops already use. Runners without that method (tests, older
shims) keep the previous behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 350cfc1436504b92d5f9e924f43f81b232beefd7)
2026-09-28 05:28:05 -07:00
teknium1
01d85137bd fix(review): MCP owner rebuild keeps the launch profile's env-only credentials (#119092 review)
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.

Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
2026-09-28 05:24:26 -07:00
teknium1
c8756dbd61 fix(mcp): remote headers render under the owner's fresh secret scope at connect and every reconnect (#119092, salvage #119097)
Under multiplex a served secondary profile's remote MCP server whose config has a
`${VAR}` Authorization header was brought up at gateway boot with the placeholder
unresolved (HTTP 401) and retried that same rendering forever.

Mechanism: `_owner_scope_home()` trusted any caller-bound secret scope, and the
gateway's boot-time `_profile_runtime_scope` binding is a SNAPSHOT taken before the
profile's external secret source (`secrets.command`) may have answered. The run
task copies that context, and `_refresh_remote_config` re-read config.yaml on every
probe but rendered it under the same frozen mapping, so a parked server sent the
literal `Bearer ${VAR}` every ~5 min for the life of the process.

- `_owner_scope_home()` now returns the owner home (the scope's stamped home, else
  the registry scope) even when a scope is bound: MCP rebuilds the owner's scope
  fresh, which retries hydration (cached once it succeeds). Single-profile
  processes (scope key None) are unchanged.
- `MCPServerTask.run()` binds that fresh owner scope around the rebuild-time config
  refresh, so a reconnect heals once the source hydrates.
- `_require_rendered_remote` fails closed: a remote `url`/`headers` still carrying
  a `${VAR}` after rendering raises naming the variable instead of sending it.
- Salvaged from #119097 (@JoaoMarcos44): the connect-time re-render, which also
  covers a lazy server whose config was rendered at boot; its mechanism alone was a
  no-op for the reported topology because the owner-scope install was skipped
  whenever a (frozen) scope was already bound.

Row L1/L3, topology T2. Tests: A→B→A under set_multiplex_active(True) with two
temp homes, red on origin/main (literal placeholder sent), green on head.

Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
2026-09-28 05:24:26 -07:00
JoaoMarcos44
c418294eca fix(mcp): reinterpolate config after owner secret hydration 2026-09-28 05:24:26 -07:00
teknium1
124b5c7d45 fix(review): key cwd/override records by the session's own profile home; keep /p/<launch>/ ids stable
Independent-review minors on #126157 (#123989 class):

- session.create is a plain @method, so tui_gateway/session_workdir.py::_register_session_cwd
  wrote the cwd record under the RAW session key while the scoped turn read
  profile:<p>:<key> and missed it until the first `cd`. The writer now binds the
  session's own profile_home around register_task_env_overrides.
- tools/terminal_tool.py::_task_env_overrides stayed raw-keyed, so two profiles
  registering the same task id (docker_image / cwd) still collided. Writer, clear
  and both readers (_has_isolation_overrides, resolve_task_overrides) now use
  _qualify_task_key; the isolation-override branch of _resolve_container_task_id
  returns the qualified key so the env cache and the override record agree.
- gateway/platforms/api_server.py::_derive_chat_session_id prefixed the seed for
  the NAMED LAUNCH profile addressed via /p/<launch>/, so prefixed and
  un-prefixed requests on one profile derived different ids. The prefix is now
  skipped when the routed name is the process's launch profile
  (_names_launch_profile vs get_routing_process_hermes_home()).

Tests (red on 3720519199c, green here):
  tests/tui_gateway/test_session_cwd_profile_key.py::test_secondary_profile_session_cwd_is_found_inside_its_scope
  tests/gateway/test_api_server.py::TestDeriveChatSessionId::test_launch_profile_prefix_keeps_the_unprefixed_id
2026-09-28 04:34:30 -07:00
teknium1
24b1c32bd2 fix(terminal,api-server): trim profile qualification of session-derived sandbox keys (#123989, salvage #124057 + #123991)
Fold the two salvaged fixes to the class and the salvage bar:

- tools/terminal_tool.py: `_qualify_task_key` reuses `_routed_home_task_key`
  instead of a duplicate qualifier; branch 2 (session-isolated sandboxes) AND
  branch 3's `session:<key>` (SSH-style backends) carry the routed profile, so a
  session id two profiles share (header-less API fingerprint, a DM chat id served
  by two bots) never resolves to one `_active_environments` slot. Cwd records
  keep the symmetric qualification. No routed home → raw key, unchanged.
- gateway/platforms/api_server.py: `_derive_chat_session_id` namespaces only
  named routed profiles, so default/standalone ids are byte-identical and live
  Open WebUI conversations survive the upgrade.
- Tests trimmed to two invariants: A→B→A across two temp profile homes under
  multiplex (distinct sandbox keys, alias follows the qualified parent, cwd
  isolated per profile), and derived-id namespacing without moving default.
  Both fail on 5912ed81ed, pass here.
2026-09-28 04:34:30 -07:00