Commit Graph

7451 Commits

Author SHA1 Message Date
kshitijk4poor
24b33d42e0 fix(honcho): every honcho.json writer holds _refresh_lock and reads strictly
The CLI's _write_config took only the best-effort file lock, so an in-process
refresh thread could still interleave with a command's read-modify-write. The
dashboard's Honcho save (_write_provider_honcho) still seeded its whole-file
rewrite from a tolerant reader, so a honcho.json that exists but does not parse
was replaced by the active host's block alone from the UI - the same bug class
this PR closes on the CLI and refresh paths. `hermes profile create --clone`
swallowed the new ConfigWriteRefused as "plugin not installed".

One parametrized test covers the web writer for the corrupt and parseable cases.
2026-09-13 19:05:28 +05:30
teknium1
0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
teknium1
0c0875b746 chore: delete orphaned bench data, datagen examples and stale one-off docs
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.

- mcp-research-data/: 224K of July tool-search bench result rows. The
  harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
  gitignored dir; the rows were committed by hand once and the headline
  numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
  WebResearchEnv that no longer exists; the yaml paths point at a
  configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
  for a resolved bug, two RFCs whose work shipped, an unimplemented
  profile-builder proposal, the kanban dialog mock HTML and the kanban v1
  spec PDF. profile-routing.md duplicated the profile_routes section of
  website/docs/user-guide/multi-profile-gateways.md.

Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
2026-09-13 06:06:46 -07:00
teknium1
e73f94fa83 refactor(platforms): every plugin setup wizard uses the shared declines_reconfigure gate
Fourteen platform plugins hand-rolled the "already configured? Reconfigure? [y/N]"
gate at the top of interactive_setup (env check + info line + prompt_yes_no(..., False)),
with drifting wording ("X: already configured" vs "X is already configured." vs
"already enabled") and, for LINE and SimpleX, raw input() loops with their own
EOF/KeyboardInterrupt handling and no gate at all. Fixes to the gate (non-interactive
handling, wording, default) therefore reached only the core Telegram/BlueBubbles/webhook
wizards.

- hermes_cli/setup_platforms.py: `_declines_reconfigure` becomes the public
  `declines_reconfigure(label, question, *env_vars)` (any-of env check, so Matrix's
  token-or-password gate fits); `_save_prompted` becomes `save_prompted` alongside it.
  No alias kept; the three core callers are updated.
- buzz, dingtalk, discord, feishu, google_chat, irc, matrix, mattermost, raft, slack,
  teams, wecom: the hand-rolled gate is replaced by one `declines_reconfigure(...)` call;
  post-decline extras (Discord allowlist nudge, Slack manifest refresh, Raft "Keeping"
  line) stay local and unchanged.
- line, simplex: the raw input() loops move onto hermes_cli.cli_output.prompt (masked
  for secrets, "" on Ctrl-C/EOF) and gain the shared gate on their primary env var.

Behavior change: the gate's info line is now uniformly "<Label>: already configured"
(DingTalk/Feishu/WeCom lose the trailing period + inline ID; Buzz/IRC/Google Chat/Raft/
Teams no longer echo the current value in that line). Feishu and WeCom now gate on the
app/bot ID alone instead of ID AND secret. LINE and SimpleX gain a "Reconfigure?" [y/N]
prompt when already configured; their prompts now honour HERMES_NONINTERACTIVE and print
via the CLI helpers instead of bare print(). Prompt defaults (No) are unchanged everywhere.

Not touched: WhatsApp's gate keys on WHATSAPP_ENABLED being truthy (a "false" value must
not count as configured), which the shared any-set gate cannot express — left hand-rolled.

Test: tests/plugins/platforms/test_interactive_setup_reconfigure_gate.py parametrized over
the 14 wizards — with the primary env var set and the user declining, each wizard must have
called declines_reconfigure with that var and returned without prompting or saving.
Sabotage: reverting mattermost's gate fails that row.
2026-09-13 05:32:38 -07:00
teknium1
7944175e62 fix(model): custom-to-custom switch drops the previous endpoint's inline api_key
model_selection_config_updates cleared model.api_key/api only when the target
was not a custom provider, so custom:legacy-box -> custom:other-box (or the
same 'custom' provider on a different base_url) carried endpoint A's inline
secret to endpoint B. The old dashboard writer cleared on any provider change.
The inline key now survives only a same-route re-pick (same provider AND same
normalized base_url); an explicitly submitted key is re-added by the dashboard
after this shape is applied, as before.
2026-09-13 05:21:02 -07:00
teknium1
08dfa32da7 fix(dashboard): profile-create validates the model in the dashboard's home, reports rejections
_write_profile_model validated under the NEW profile's HERMES_HOME. A just-created
profile has no providers:, no .env and no auth, so every non-env provider the
create dialog offered (anthropic, ollama, providers:-keyed custom) came back
'Unknown provider' / 'Could not resolve credentials' and create returned a
silent model_set: false. The picker read THIS dashboard's catalog, so validation
now runs scoped to the process home (validate_in) while the write still lands in
the new profile. PUT /api/profiles/{name}/model keeps validating in the target
profile (it is an existing, credentialed home). A real rejection now surfaces as
model_error in the create response instead of a log line.
2026-09-13 05:21:02 -07:00
teknium1
466ccac5ab fix(windows-update): a user-launched hermes serve/dashboard is never a Desktop backend
Canonicalising _is_backend_argv onto _hermes_holder_subcommand dropped the
'-m hermes_cli.main' entry-shape discriminator the old predicate had. That
widened _orphaned_desktop_backend_pids to tree-kill a standalone hermes.exe
serve / hermes dashboard whose console parent died. The Desktop's only spawn
shape is -m hermes_cli.main (apps/desktop/electron/main.ts), so _is_backend_argv
now forwards to _looks_like_desktop_control_plane — the same canonical-subcommand
AND entry-shape predicate — instead of a second copy. Test matrix gains a
'Desktop backend?' column with plain hermes serve / hermes.exe dashboard rows.
2026-09-13 05:21:02 -07:00
teknium1
0885e19a75 fix(dashboard): a rejected model switch is a 400, never a flattened model: block
_denormalize_config_from_web ran the new switch_model validation inside the
pre-existing 'except Exception: pass' disk-read fallback. Any rejection
(offline models.dev, no OpenRouter key, unlisted model) left config['model'] a
flat string, and PUT /api/config's deep-merge wrote that string over the on-disk
model: dict, destroying provider/base_url/api_mode/context_length/model_slots.

Only the load_config() read keeps its fallback; validation propagates as the
HTTPException(400) the caller's http_failure passes through. Invariant test PUTs
a rejected model against a real config.yaml and asserts 400 + byte-identical file.
2026-09-13 05:21:02 -07:00
teknium1
11576390fe refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.

Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.

Sites -> canonical:
  hermes_cli/cli_model_switch_mixin.py::_persist_global_switch          -> deleted; _commit_model_switch calls persist_model_selection
  hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
  gateway/slash_commands_model.py::_persist_model_switch_to_config       -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
  tui_gateway/model_switch.py::_persist_model_switch                     -> deleted; _apply_model_switch calls persist_model_selection
  hermes_cli/web_server_config.py::_apply_main_model_assignment          -> apply_model_selection(result) (+ explicit custom api_key)
  hermes_cli/web_server_config.py::_validated_main_model_selection       -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
  hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
  acp_adapter/server.py::_resolve_model_selection                        -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError

Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.

Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.

Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
2026-09-13 05:21:02 -07:00
teknium1
3e066dfedd fix(computer_use): approval goes through the shared gate; no callback now fails closed
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.

_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.

Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
  headless -q, plain library use): destructive actions are now REFUSED
  with a BLOCKED error and never reach the backend. Previously they
  silently ran. cron honors approvals.cron_mode, unattended platforms
  approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
  approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
  command_allowlist entry (`cua:click:background`) and is scoped to that
  action+mode — the old blanket "always_approve unlocks everything for
  the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
  "BLOCKED: Action timed out ..."); the error JSON keeps `action`.

Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
2026-09-13 05:21:02 -07:00
teknium1
f537c0b4c3 refactor(status): CLI, gateway and TUI /status render the same field set from hermes_cli/status_report.py
The three /status renderers (hermes_cli/cli_session_mixin.py::_show_session_status,
gateway/slash_commands_status.py::_handle_status_command, tui_gateway/methods_session.py
session.status) each hand-built Session ID / Path / Title / Model (provider) / Created /
Last Activity / Tokens / Agent Running with their own getattr(agent, "model") fallback
chain, their own updated_at/last_updated_at/last_activity_at scan and their own timestamp
format. A fix to one (a new last-activity column, a placeholder change) silently missed the
other two.

hermes_cli/status_report.py::build_status_fields now derives the common facts once and
returns them as structured, display-ready data; status_lines() renders the English
"Label: value" form for the CLI and TUI. The gateway keeps translating through its
existing t("gateway.status.*") catalog keys (no locale change); the CLI keeps reasoning /
approvals / context, the gateway keeps free-tier / context / queue depth / Matrix scope,
the TUI keeps its Project line. tui_gateway/methods_session.py::_status_dt and the
CLI's inline updated_at loop are gone; cli_session_mixin._timestamp_or stays for its
remaining history-timestamp caller.

Behavior change: none intended for populated sessions. Unified edge cases: a
SessionDB row with an unparseable started_at now falls back to now() on the TUI as it
already did on the CLI, and the TUI's fallback on a bad updated_at is the created stamp on
both surfaces.

Test: tests/hermes_cli/test_status_report_contract.py drives the three real renderers with
one session (distinctive model, provider, title, stamps, token count) and asserts each
output carries every common value. Sabotage-verified: builder dropping tokens -> red;
TUI hand-formatting the model line -> red; restored -> green.
2026-09-13 05:21:02 -07:00
teknium1
991b23ad8d refactor(sessions): one session-id minter; QQ update-prompt key from build_session_key
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.

Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".

gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.

Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
2026-09-13 05:21:02 -07:00
teknium1
c1e58f4cb2 fix(process-identity): desktop reap, Windows updater and profile liveness use the canonical matchers
Two kill/relaunch predicates decided identity by argv substring, the bug class root AGENTS.md
forbids: hermes_cli/dashboard_procs.py::_is_desktop_local_serve_cmdline (`"serve" not in cmd`,
on the orphan-reap KILL path) and hermes_cli/update_cmd_windows.py::_is_backend_argv
(`" serve" in argv_low`, in the very file that defines _hermes_holder_subcommand). Both now ask
the canonical token classifier; host/port are read as flag values, not substrings.

hermes_cli/profiles.py::_check_gateway_running open-coded rungs 1/3 of
gateway.status.resolve_gateway_liveness and skipped the multiplexer rung; it is now that ladder
scoped to the profile dir (pid probe keeps cleanup_stale=False so a probe for another profile
never unlinks its PID file). The gateway/status.py ladder itself is untouched.

Behavior change: `hermes kanban --preserve-cache --host 127.0.0.1 --port 0` and
`-m dashboard serve`-style argv are no longer classified as serve backends (never killed /
relaunched as one); a named profile served by the live default multiplexer now reads as
running from _check_gateway_running (previously only via the separate
_served_by_running_multiplexer OR at some call sites).
2026-09-13 05:21:02 -07:00
teknium1
c0be0e0826 refactor(display): one think-tag list; ACP tool titles derive from agent/display previews
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.

acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
2026-09-13 05:20:26 -07:00
teknium1
c3de99bbf0 refactor(state): one carrier-aware user-turn rewind behind CLI /undo, /retry, gateway and TUI
CLI `_rewind_persisted_user_turn`, TUI `_rewind_active_session_history` and gateway
`rewind_session` each re-ran get_active_message_ids -> get_messages_as_conversation ->
split_user_originated_turn -> rewind_to_message with their own warm/durable comparison
helpers and three different out-of-range contracts (RuntimeError / ValueError / None).

The durable transcript is the authority for a rewind, so the implementation now lives
with the data: `SessionDB.rewind_user_turn` (hermes_state_rewind.py) with one typed
out-of-range error (`RewindTargetUnavailableError`). Surfaces keep only lock, eviction
and rendering glue and map that error to their own message.
2026-09-13 05:20:26 -07:00
teknium1
d364473620 refactor(model-providers): fetch_models stubs drop to the ABC; commandcode and AI Gateway catalog GETs go through open_credentialed_url
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
2026-09-13 05:19:48 -07:00
teknium1
1512edfb85 refactor(tools): terminal, execute_code, MCP and the bounded collector truncate through one head/tail helper
Four copies of the 40/60 head/tail algorithm with a near-identical notice
(terminal_tool_result, mcp_tool_content, code_execution_tool,
environments/base_output) collapse into tools/tool_output_truncate.py, so the
ratio and the `... [<LABEL> TRUNCATED - N <unit> omitted out of T total] ...`
marker are defined once. execute_code keeps byte mode + spill path and only
shares the notice/split. Visible change: the terminal notice now uses
thousands separators like the other three (`9,000 chars` not `9000 chars`).

kanban_specify._truncate: comment claimed escape stripping the body never did;
comment now says what the plain clamp is for.
2026-09-13 05:09:43 -07:00
teknium1
398234748f refactor(agent): one Retry-After parser and one reset-grammar table feed every retry wait
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.

The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
2026-09-13 05:09:43 -07:00
teknium1
5aa1a50c57 fix(config): good-config backup only for the active home; fixture-based effective-config contract
- load_user_config_effective wrote backups/config/*.good.* into ANY home it
  read, so doctor and the TUI cwd lookup created backup dirs inside other
  profiles. The copy is now taken only when the path is the active home's
  config (the only home load_config ever backed up).
- The effective-config test re-composed the implementation's own primitives
  (a mirror); it now pins a literal expected dict for a fixture of user file +
  managed overlay + env, and fail_closed asserts yaml.YAMLError.
- Four repointed gateway tests carried duplicate _load_gateway_config setattr
  lines (one silently overriding the other); deduped to the intended dict.
2026-09-13 05:09:06 -07:00
teknium1
fd5693bc92 fix(config): kanban decompose and local-models status keep their fail-open config reads
Repointing the two _load_config copies at load_config_readonly() dropped the
except-Exception guards the old copies had. load_config_readonly runs
ensure_hermes_home(), which can raise FileNotFoundError / HomeInitializationError,
so decompose_task (promises ok=False) and /api/local-models/status (garnish
that must render degraded) would raise / 500 instead. The guard is restored
at both sites with the same breadth the old code had.
2026-09-13 05:09:06 -07:00
teknium1
469a87f7a5 refactor(config): collapse thin _load_config copies onto the canonical readers
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.

tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
2026-09-13 05:09:06 -07:00
teknium1
93d0dba281 fix(profiles): HERMES_HOME-blind path sites resolve through hermes_constants
worktree_gc archived untracked files under ~/.hermes regardless of the active
profile or HERMES_HOME (and the Windows LOCALAPPDATA default); dashboard_procs'
remote-lock dir, both bot_mode `_default_home` copies, methods_bot_relay's
`_relay_root` (a hand copy of get_default_hermes_root's profiles/ strip) and
load_hermes_dotenv's default home all re-derived env-or-`~/.hermes` by hand and
so diverged from the platform default. Each now calls the canonical getter
with the intent it already had: get_hermes_home() where the active profile
matters (archive), get_process_hermes_home() where the process asset must stay
visible under a routed-profile override (locks, bot mode, startup .env),
get_default_hermes_root() for install-wide relay state.
2026-09-13 05:09:06 -07:00
teknium1
b91088d768 refactor(config): one effective-user-config loader replaces 9 hand-rolled raw→overlay→expand pipelines
Every defaults-free config reader (gateway runtime, TUI gateway, cron
scheduler + job snapshot, `hermes send` env bridge, doctor memory section,
hermes_cli/main early parse, hermes_time, hermes_logging, the gateway
fallback-chain refresh) re-implemented "read config.yaml + managed overlay +
${VAR} expansion" by hand, in three different orders, and none of them
replayed the model-key canonicalization or the last-known-good recovery that
load_config() gained. An admin-pinned `${VAR}` expanded on one surface and was
bridged literally on another; `model: {name: x}` resolved to an empty model
everywhere except the gateway.

hermes_cli/config_effective.py::load_user_config_effective is the one
primitive: user file → ${VAR} → managed overlay → _normalize_root_model_keys,
no DEFAULT_CONFIG merge, sharing read_raw_config's parse cache and serving the
last good parse (in-process, then backups/config/*.good.*) on torn YAML;
`fail_closed=True` raises for the one caller that keeps its own last-good
state (the fallback-chain refresh). gateway/run.py::_load_gateway_runtime_config
is deleted — it was _load_gateway_config plus expansion, and _load_gateway_config
now expands.

Behavior change: _load_bridge_config, send_cmd._load_hermes_env and
doctor_state._doctor_memory_config expanded BEFORE the overlay; they now match
load_config (managed `${VAR}` expands against the process env only). All nine
sites gain model-key canonicalization and last-good recovery.
send_cmd._load_hermes_env now routes its .env read through
env_loader._load_dotenv_with_fallback so the credential sanitizer runs.
2026-09-13 05:09:06 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1
6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
teknium1
dd1baee0e4 refactor(secrets): drop scope-aware env shims; runtime_provider and the voice/xai tools read the canonical getters
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.

tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.

Behavior change: none.
2026-09-13 05:07:50 -07:00
teknium1
c849bc383a refactor(env): agent.secret_scope.load_env_file is the only .env tokenizer; six hand parsers collapse onto it
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.

Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.

Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.

Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
2026-09-13 05:07:50 -07:00
teknium1
226df89f74 refactor(redact): one secret-pattern source; a2a, gateway chat and monitoring egress scrub through redact_for_egress
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.

Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.

Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).

Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
2026-09-13 05:07:50 -07:00
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
714013c493 fix(utils): new non-secret atomic writes follow the process umask again
Every hand-rolled writer this PR folded into utils._atomic_write created a
NEW file with write_text()/open("w"), i.e. at 0o666 masked by the umask
(0644 under 022). The canonical helper publishes through mkstemp, whose
temp is 0600, and with no explicit mode and no existing target to copy bits
from it left that 0600 in place - so debug, model_catalog, profiles,
breadcrumbs, worktree_ops, web_result_cache, plugin_compat, write_approval,
rich_sent_store, active_sessions and the google_meet state files were
silently tightened to owner-only, the volume-mount hazard
_restore_file_metadata's own docstring warns about. Undeclared in the PR.

Fix at the canonical: when mode is None and the target does not exist,
apply default_new_file_mode() (0o666 masked by the umask, read via the
umask two-call trick with a transient 0o077 so a racing thread can only get
a tighter file). The helper is hermes_cli/backup._default_new_file_mode
moved into utils and reused. Secret writers (mode=0o600) are 0600 before,
during and after as before; an existing target keeps its bits; on non-POSIX
the helper returns None so nothing is chmod'd.
2026-09-13 05:07:11 -07:00
teknium1
30657d197d refactor(update): psutil Android installer extracts through the shared safe-tar guard
hermes_cli/psutil_android carried its own tar path-traversal / link-member
guard (a 0.87 copy of archive_safe.safe_extract_targz). One guard for every
tar.gz we extract; the installer keeps raising PsutilAndroidInstallError so
its callers' except clauses are unchanged.

Behavior change: the psutil path now also rejects Windows-absolute and
backslash-smuggled member names (archive_safe.normalize_archive_parts), and
chmod failures on extracted files are suppressed identically.
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
teknium1
2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
Franci Penov
f361971eed feat(gateway): fire agent_loop_stopped plugin hook on interrupt
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.

_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.

Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.

Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
2026-09-12 22:19:44 -07:00
nyx573
0007a4c2f9 feat(auth): OpenRouter OAuth PKCE login via hermes auth add openrouter --type oauth
Browser login against openrouter.ai/auth (S256 PKCE, bind-first OS-assigned loopback port,
POST /api/v1/auth/keys code exchange) that stores the minted key as a plain api_key pool entry
with source manual:openrouter_pkce, so it rotates and resolves exactly like a pasted key.

Salvaged from #102639 (nyx573) onto the facade+siblings layout: the flow lives in the new
auth_openrouter sibling and rides the shared loopback helpers instead of appending to the
auth.py facade; auth_commands gains a table row rather than a provider branch. OpenRouter echoes
no `state`, so the CSRF nonce rides in the callback path (a redirect that guesses the port but not
the nonce is a 404 and never reaches the exchange); remote/SSH sessions use OpenRouter's
documented headless paste-the-code mode instead of an unreachable loopback listener. OpenRouter
keeps its API-key default when --type is omitted so the documented `--api-key` form is unchanged.
2026-09-12 22:07:41 -07:00
Teknium
3310298a37 refactor(auth): share the loopback PKCE listener between OAuth flows
Move the S256 verifier/challenge pair, the loopback callback handler, the bind-first
listener and the serve-until-redirect loop out of auth_spotify into auth_device_flow so a
second loopback PKCE provider does not copy 80 lines of HTTP-server plumbing. Spotify's
behaviour and error codes are unchanged; only its private copies are deleted.
2026-09-12 22:07:41 -07:00
teknium1
08a2e7dbcc refactor(cli): vim mode is a config key only — drop the /vim slash command
Maintainer ruling: no new slash command for this. `display.vim_mode: true`
in config.yaml enables vi keybindings in the composer at startup; the
NORMAL/INSERT/REPLACE status-bar label stays. Removes the CommandDef, the
handler, its dispatch-table entry and slash-command docs; documents the key
under Display Settings.
2026-09-12 22:00:02 -07:00
Sam Foreman
750b2fc5e2 feat(cli): wire vi editing mode into the prompt_toolkit Application
Implements vim mode for the classic CLI input composer:

- Pass editing_mode=EditingMode.VI to Application when display.vim_mode
  is set, else EditingMode.EMACS (prompt_toolkit's own default), so the
  change is inert for users who have not opted in
- Add _handle_vim_command() to toggle at runtime and persist the choice
  to display.vim_mode; the live Application is updated in place so the
  toggle takes effect without a restart
- Add _vim_mode_label() and surface NORMAL/INSERT/REPLACE in the status
  bar while vim mode is active, as requested in the issue

Credit to #4325 by @SHL0MS, which first identified the EditingMode
wiring and the /vim toggle; that PR has gone stale against main. This
revives the approach, adds the missing config default and the status
indicator, and covers it with tests.

Closes #4254.
2026-09-12 22:00:02 -07:00
Sam Foreman
2a5373da2c feat(cli): add display.vim_mode config and register /vim command
Introduces the opt-in surface for vi editing in the input composer:

- display.vim_mode config default (False, so existing users are
  unaffected and prompt_toolkit keeps its standard emacs bindings)
- /vim command registered with on|off|status subcommands, matching
  the established /battery and /timestamps pattern

Part of #4254.
2026-09-12 22:00:02 -07:00
Teknium
7b037f0efa feat(cli): opt-in git_branch status-bar field (⎇ current branch)
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.

Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.

Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
2026-09-12 21:47:53 -07:00
Teknium
053f8b1b17 fix(kanban): make archive-time worker termination race-safe and audited
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:

- Snapshot status/pid/claim INSIDE the archive txn so the kill only
  happens when this caller wins the archive transition; a losing
  concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
  skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
  hold the SQLite write lock. Safe because archived is terminal — no
  dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
  event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
  -> no signal, no event), live E2E: worker survived archive on main,
  terminated (<0.3s, clean SIGTERM) with the fix.

Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
2026-09-12 21:45:15 -07:00
PINKIIILQWQ
907f409dc4 fix(kanban): archive_task now kills the worker process
archive_task() was a pure DB operation — it cleared worker_pid,
claim_lock, and status from the tasks row but never sent SIGTERM
to the actual OS process. A running worker stayed alive until it
next called kanban_complete/kanban_block and discovered it was
archived, burning API quota and compute resources.

Fix: snapshot pid+claim_lock before the write_txn clears them,
then call _terminate_reclaimed_worker() — same function reclaim_task
uses — which sends SIGTERM, waits 5s, then SIGKILL if still alive.
The termination metadata is included in the 'archived' event so
operators can see what happened.

Order matches reclaim_task: terminate first, then DB update.
For non-running / non-local tasks, _terminate_reclaimed_worker
returns immediately as a no-op.

Closes #33774 reprise: the scratch-workspace side was fixed in
fc8afd500, but the orphaned-process side was never addressed.
2026-09-12 21:45:15 -07:00
Teknium
77f0c83ec3 fix(sessions): honor CLAUDE_CONFIG_DIR and CODEX_HOME in foreign session discovery
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.

_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.

Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
2026-09-12 21:39:10 -07:00
salch-cred
ad0398eed8 fix(kanban): preserve sticky block on tasks created with initial_status=blocked (#107398) 2026-09-12 21:27:21 -07:00
Teknium
0f3199bd65 fix(dashboard): OAuth start routes resolve pollers late so test mocks intercept the spawned thread
The oauth router imported _nous_poller/_minimax_poller/_xai_device_poller
from web_server_oauth at module level, so tests patching the owning module
("hermes_cli.web_server_oauth._minimax_poller") patched a binding the
router never read. The REAL poller then ran on the leaked daemon thread,
called the live MiniMax token endpoint from CI, and the in-flight
getaddrinfo segfaulted the interpreter during a later test's fixture setup
(CI run 34323790818, tests/hermes_cli/test_web_oauth_dispatch.py flake).

Route the three pollers through the existing late() seam (web_deps), the
same mechanism every other monkeypatch-sensitive symbol in this router
already uses, so the patch wins at thread-spawn time. Regression test
proves the mock intercepts and the real poller body never runs; it fails
on the old module-level import (sabotage-verified).
2026-09-12 21:00:55 -07:00
Søren L. Hansen
0037a4b17a feat(cron): add resnap action to adopt the current global inference default
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.

Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
2026-09-12 20:57:21 -07:00
686f6c61
2bc9ed9f23 fix(mcp): fully redact credential headers in MCP probe errors and test display (salvage #97466)
`hermes mcp test` resolved Authorization headers and printed first4***last4
— still a reusable credential fragment — and probe exceptions that echoed
`Authorization: Bearer <value>` reached the CLI error line and the dashboard
`POST /api/mcp/servers/{name}/test` response verbatim.

Redact once at the `_probe_single_server` raise seam so every consumer
(`mcp add`, `mcp test`, `mcp login`, `mcp configure`, the dashboard probe,
`hermes doctor`, catalog probes) prints already-safe text. Recognized
credential header fields (Authorization/Proxy-Authorization plus
agent.redact._SECRET_HEADER_NAMES) have their complete value replaced with
***; bare Bearer/Basic/Token/Digest spans are covered; the generic redactor
runs force=True as a second pass. CLI header display fails closed: only pure
${ENV} template values print.

Salvaged from PR #97466 by @686f6c61 (base predated the mcp_config/web_routers
decomposition; re-applied onto current main, test seams repointed to the
defining modules tools.mcp_tool_loop / tools.mcp_tool_lifecycle).

Inspired by Claude Code 2.1.268: "Fixed /mcp and /plugin server details,
claude mcp list/get, and MCP login errors showing secrets resolved from
${VAR} placeholders in MCP configs."

Fixes #97460

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-12 20:49:18 -07:00
Hermes fleet-fix
1916cb249d security: make state databases and snapshots owner-only 2026-09-12 20:43:42 -07:00
teknium1
205645ee42 fix(gateway): register gateway.multiplex_profiles; explicit migrate --multiplex flips it with no standalone secondary
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.

`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
2026-09-12 18:35:21 -07:00
teknium1
acbecf588a fix(profiles): --clone leaves messaging channels behind; --clone-channels opts in
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).

Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.

`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.

The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
2026-09-12 18:35:21 -07:00