Commit Graph

1384 Commits

Author SHA1 Message Date
kshitijk4poor
52d203d041 fix(agent): read the trim flag with a default — bare test stubs reach _execute_tool_calls without _set_defaults 2026-09-20 18:09:19 +05:30
kshitijk4poor
e74a84f0b6 refactor(agent): declare the trim flag in _CONTROL_STATE; trim only on a completed batch
Gate Lows: an in-flight exception pins the executor frames through its traceback, so the
finally-placed trim could not release the result on that path and would burn the cooldown;
the flag now survives to the next completed batch. The default lives beside _executing_tools.
2026-09-20 18:09:19 +05:30
kshitijk4poor
05e09aee68 fix(agent): trim after the tool batch unwinds, not inside the commit that still holds the result
Gate review (three lenses converged): inside _commit_tool_result the >=1 MB string is still
referenced by the publish frames (the returned tuple, managed.result / batch.results, the
tool.completed callback), so gc.collect + malloc_trim could not release it and merely spent
the 60 s cooldown. The commit now only sets agent._trim_after_tool_batch; the finally of
AIAgent._execute_tool_calls consumes the flag once every executor frame is gone, coalescing
N large results in one batch into one trim. The import is lazy like every sibling call site
(keeps ctypes out of the agent import chain).
2026-09-20 18:09:19 +05:30
kshitijk4poor
ddc55f217b fix(agent): Responses-upgrade predicate declines every external-process profile, not three names
``_provider_model_requires_responses_api`` is the one predicate both primary routing
(``agent_init._finalize_routing``) and GPT-5 fallback activation
(``chat_completion_helpers._fallback_api_mode_resolved``) consult; the fallback path had no
ACP exclusion at all, so activating ``copilot-acp / gpt-5.x`` as a fallback selected
codex_responses and the next request died with ``'CopilotACPClient' object has no attribute
'responses'`` (#65842). Key the exclusion on the profile's external_process auth_type so the
bundled facade and out-of-tree ACP plugins are covered by the same line.
2026-09-20 12:36:34 +05:30
Scott Glover
9bb0c4a40f fix: keep Copilot ACP fallbacks on chat completions
CopilotACPClient exposes the chat.completions facade, not the Responses API. Prevent generic GPT-5 routing from switching copilot-acp fallbacks or either supported provider alias to codex_responses.

Add parameterized regressions for all three provider IDs and validate with focused suites plus live forced Terra-to-Luna fallback requests.

(cherry picked from commit 687699b6e97df6b91b9eb9c90fe4840c802c0a27)
2026-09-20 12:36:34 +05:30
teknium1
3934fe551c fix(codex): sweep the retired gpt-5.3-codex slug off every Codex-OAuth surface
What: the contributor's pick removes the slug from DEFAULT_CODEX_MODELS and the
forward-compat templates. This follow-up finishes the class on the sibling
surfaces that still taught users the retired slug:

- hermes_cli/cli_model_switch_mixin.py: `/model` -> openai-codex replaced an
  untouched default with the literal "gpt-5.3-codex" when live discovery failed;
  it now uses DEFAULT_CODEX_MODELS[0] so the fallback can never drift from the
  curated list again.
- run_agent.py: the Codex silent-hang hint no longer recommends the retired slug.
- website/docs (+ zh-Hans): the Codex OAuth vision/fallback examples used the
  retired slug and claimed a default that no longer exists (openai-codex has no
  implicit auxiliary model); examples now use gpt-5.4 and say to set `model`.
- tests: one invariant test pins the slug out of the curated list and every
  template tuple (#52492); hint/watchdog tests assert the retired slug is absent;
  catalog fixtures switched to live slugs.

Why: the ChatGPT Codex backend returns HTTP 400 "not supported when using Codex
with a ChatGPT account" for gpt-5.3-codex (#52492; second field report incl. the
official CLI on the #81558 thread, 2026-08-10). Same shape as e8955f222c, which
dropped gpt-5.2-codex / gpt-5.1-codex-max / gpt-5.1-codex-mini. Live discovery
still surfaces the slug if the backend re-enables it.
2026-09-19 10:22:33 -07:00
teknium1
7395931545 fix(agent): release_clients() closes the Codex app-server session too
Gateway cache eviction pops the AIAgent and soft-releases it through
release_clients(), which kept `_codex_session` alive on the assumption
that it was session tool state. It is an LLM client: the evicted
instance is discarded and the rebuilt agent spawns its own app-server
child, so every evicted Codex-route session leaked one `codex
app-server` process (plus its MCP children) for the gateway's lifetime.
close() already owned this since e623432b892; the soft path is the
remaining hole (#72548, field report in #66671).

Test: red on base, green here; idempotent across repeated release/close
and exception-safe (reference cleared before close()).
2026-09-19 09:43:20 -07:00
teknium1
ae1b5d79b2 fix(sessions): one-shot runs get a distinct oneshot source that pickers hide
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).

- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
  inherited UI transport label without an explicit --source) resolves to `oneshot`; the
  platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
  (HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
  three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
  `sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
  recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
  documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
  search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
  workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
  instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
  session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
  `--source` flag reference (explicit flag always stored as given).
2026-09-17 09:06:59 -07:00
teknium1
0010273f1e fix(sessions): an explicit --source tui|desktop survives the one-shot label drop
_session_source_for_agent now drops an inherited tui/desktop
HERMES_SESSION_SOURCE for finite chats, but main.py exports the
documented --source flag through the very same variable, so
'hermes chat -q --source tui' was silently relabelled 'cli' with no way
to opt out. main.py marks an explicit flag with
HERMES_SESSION_SOURCE_EXPLICIT=1 and the resolver keeps the label when
that marker is set; no internal launcher passes --source tui/desktop, so
inherited transport labels are still dropped.
2026-09-16 17:50:21 -07:00
teknium1
c5f6336fa5 fix(sessions): one-shot runs stop inheriting the tui/desktop session source
A finite `hermes chat -q` (or `hermes -z` one-shot) launched from inside a TUI
or Desktop session inherits HERMES_SESSION_SOURCE=tui/desktop — the terminal
tool bridges the session env into child processes — and
`_session_source_for_agent` honoured that inherited transport label over the
child's own platform. The row was persisted as a `tui` chat, so it appeared in
the TUI/WebUI session pickers as a resumable conversation and `hermes -c` in
the TUI could continue it.

The resolver now drops an inherited UI-transport source (`tui`, `desktop`) when
the finite-chat marker HERMES_SINGLE_QUERY_SESSION=1 is set, falling back to the
child's platform (`cli`). Automation sources (kanban, tool, cron, a2a, ...) are
still inherited on purpose — kanban dispatch relies on it. `run_oneshot` sets
the same marker as `hermes chat -q` already does.

Deliberately NOT a new `oneshot` source and NOT a forced `tool` source: `tool`
is the documented opt-in (`--source tool`) for integrations that must stay out
of session lists, `hermes -c` resolves the latest session by source=cli, and
every picker (TUI gateway deny list, dashboard automation category,
session_search, `hermes sessions list`) already filters on the existing
sources. Whether plain one-shot runs should be hidden by default is a product
call left to the maintainer (see PR body).

Part of #112550
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-16 17:50:21 -07:00
teknium1
ebd106f3ec fix(codex): floor every effort >= high and keep the run-budget cap over it
Two gaps in the high-effort silence floor (#112909):

- The floor matched only {'high','xhigh'}. 'max' is a real Codex wire level
  (CODEX_GPT56_EFFORTS / CODEX_ASTRA_EFFORTS) and 'ultra' is the Hermes-internal
  Codex tier, so the tiers that think LONGEST still got the 12s/120s/90s
  small-prompt fuses. Match by EFFORT_LADDER position >= 'high' instead.

- The stale floor was applied in _resolve_nonstream_watchdogs AFTER
  AIAgent._compute_non_stream_stale_timeout, undoing its run-budget cap: a
  100s cron/kanban budget capped stale at 60s but the watchdog got 300s, so one
  hung call could outlive the whole run. The floor now lives inside
  _compute_non_stream_stale_timeout before the cap; the helper keeps flooring
  only the TTFB/idle implicit defaults.

Tests: max/ultra params + a run-budget invariant, both red without this commit.
2026-09-16 17:45:01 -07:00
fangliquan
20ade400ab fix(auxiliary): propagate OpenCode session header 2026-09-16 17:22:17 -07:00
teknium1
0023b904e4 fix(compression): re-arm the drifted-prompt INFO on session reset so it is once per session
The INFO-once flag `_compaction_prompt_drift_logged` lived on the AIAgent and
was never cleared, so on a CLI agent that /new, /resume or /branch-rebinds
sessions the line was INFO once per agent lifetime, not once per session as
the commit and PR text claim. reset_session_state is the per-session reset
point for agent-held counters (`_user_turn_count` lives there); clear the
flag alongside it.
2026-09-16 17:12:33 -07:00
outpoints
5237cab756 fix(honcho): thread logical cwd through agent construction
(cherry picked from commit b1d7207c45311be658592c6ad34ee84634fed0ee)
2026-09-15 22:30:11 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
Siddharth Balyan
e0ef0eb9c3 manage_connections covers local MCP servers; setup_mcp leaves the schema (NS-867, PR1) (#109517)
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema

One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.

MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.

Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.

`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.

Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.

`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.

The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.

* wip(desktop): connection.request store, resume restore, card routing for MCP targets

Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.

* fix(config): hermes update turns on the connections toolset for saved toolset lists

`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.

Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.

`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.

* refactor: anti-slop pass on the desktop slice; shorten added comments

Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.

* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed

The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.

session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.

lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.

vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.

* chore: drop __pycache__ files swept in by an over-broad git add

* fix(desktop): correlate the connection.request row with the model's tool call by reason

The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.

* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds

* fix(connections): settle reason derives from target state, never from the renderer

A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.

* fix(desktop): a pending connection card re-arms on resume and activate

The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.

* style: literal wording in added comments, docstrings and docs

* fix: shared gateway-event contract and config-schema category for the connection events

connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.

* style: import order (perfectionist) in the desktop and shared files this PR touches

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-15 00:41:13 +05:30
kshitijk4poor
32a0d2d2e1 refactor(agent): one home for the stream parse-error markers
The SDK's malformed-frame ValueError substrings were hand-copied into the
retry classifier and the error summary; a third SDK message would need
both remembered. The mid-tool retry branch no longer clears
partial_tool_names itself — _start_stream_attempt resets it for every
attempt. buffer_anthropic_tool_input's docstring states the trade: the
knob stays on the turn's kwargs and the changed tools block costs one
prompt-cache miss.
2026-09-14 20:01:19 +05:30
joaomarcos
dfb4caf4b7 fix(anthropic): recover malformed streamed tool JSON 2026-09-14 20:01:19 +05:30
joaomarcos
6bc0e9e6df fix(agent): same-model review fork keeps the parent's affinity header and Portal conversation root (#109964)
Trimmed salvage of #110045 (deltas 1 + 2 only), stacked on the #110009 scope inheritance:

- `declared_conversation_scope` treats an inherited value as a DECLARED scope only when it
  carries the `gwk_` prefix. A rotated CLI parent publishes no affinity scope (None → sticky
  key falls back to the conversation root); the fork now publishes exactly the same instead
  of an explicit physical lineage root. `resolve_prompt_cache_scope` honors any inherited
  value directly, so the body `prompt_cache_key` still matches.
- `build_cache_parity_fork` snapshots `parent._conversation_root_id()` as
  `_cached_conversation_root`; with `_session_db=None` the fork's own walk fell back to the
  parent's PHYSICAL id, so after a compression rotation the review's Portal
  `conversation=` tag fragmented usage attribution across one logical conversation.

Dropped from the original: copying `_gateway_session_key` onto the persistence-detached
fork (no cache-identity consumer reads it there; the compression-boundary hooks were
deliberately severed by `_detach_fork_compression`), and the defensive
hasattr/callable/try wrapper around `_conversation_root_id()`.
2026-09-14 06:55:54 -07:00
Teknium
f3f5c4f7c7 Port from RooCodeInc/Roomote#1796: per-task image forwarding on delegate_task
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.

Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.

Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
2026-09-13 21:05:42 -07:00
Teknium
1fec70ea48 fix(guardrails): catch repeating multi-call cycles in the stall guard
Port from can1357/oh-my-pi#10521: their loop guard only hashed single-call
turns, so a model replaying the same multi-call batch every iteration was
never counted; they widened the hash to the whole batch. Hermes has the
same blind spot in a different shape: observe_call tracks a CONSECUTIVE
identical-call streak, so an A,B,A,B,... cycle of identical (args, result)
pairs resets the streak on every alternation and runs to the iteration
budget unflagged (live-reproduced: 60 calls in a 2-cycle, zero notices,
no hard stop).

Add a period-2..4 cycle detector over a bounded per-turn call history:
notice on the STALL_GUARD_IDENTICAL_CALL_THRESHOLD-th identical lap,
hard stop at no_progress_block_after laps under hard_stop_enabled — the
same thresholds the period-1 streak uses. Cycles whose results change
between laps never fire (real progress); cycles made only of poller-exempt
tools are exempt (legitimate waiting), matching single-call semantics.
Widen the streak-stop propagation seam in run_agent.py to carry the new
decision code.
2026-09-12 21:07:46 -07:00
Justin Bennington
8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Justin Bennington
d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Erosika
55b3ea0b11 fix(gateway): count each bot message once in the loop guard and consume the author variable
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it
again, and the busy path asks a third time. Each call counted one loop-guard event, so a
Telegram bot tripped the budget after a third of the configured messages. The verdict now only
refuses a chat that is cooling down. The ingress gate counts an admitted bot message once.

`parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag,
and returns None for an author with neither id nor name. Names keep format characters and
non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR
before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive
number. Issue numbers move out of code comments.
2026-09-10 10:27:07 -07:00
Erosika
8969511209 feat(memory): carry the turn's author to sync_turn
on_turn_start already received the author trio. sync_turn did not, so a
provider that wanted to write the turn under its author had to stash state
between the two hooks. sync_turn now takes turn_author as a keyword-only
argument, and MemoryManager sends it only to providers whose signature
accepts it, so existing providers keep working unchanged.

build_turn_context resets the author on the agent at the start of every turn
so a cached gateway agent never carries a bot author into the next human
turn. agent/turn_author.py holds the parsing and the HERMES_TURN_AUTHOR
carrier.

MemoryProvider.identity_signature() is a new optional hook: the identity
values a provider writes under, declared by the provider itself, for the
gateway's agent cache to key on.
2026-09-10 10:27:07 -07:00
liuhao1024
c0714575c3 fix(review): distinguish explicit /refine from unattended reviews and surface staged consolidations
Review follow-up on #105944 (#105921):

- explicit /refine forks now run under the refine_review write origin
  (explicit flows from the CLI/gateway handlers through
  _spawn_background_review_now and spawn_background_review_thread down
  to build_cache_parity_fork), so a user-requested review keeps the
  full memory operation set; only automatic reviews stay behind the
  unattended delete gate.
- the unattended delete gate now stages the denied replace/remove (or
  whole batch) into the pending store instead of dropping it: the
  fork's own review summary is never published, so a plain denial lost
  the consolidation request with no surfacing path. The staged proposal
  carries a proposal_staged marker that summarize surfaces as an action
  line, and a staging failure still fails closed to a plain denial.
- regression tests: explicit-path origin pass-through, refine_review
  keeping replace working, near-limit denial end to end (add rejected
  by budget -> replace staged -> proposal surfaces, store unchanged).
2026-09-09 12:19:13 +05:30
Teknium
dca7a90cf8 fix(agent): reclaim background processes by execution owner
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.

Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
2026-09-07 04:38:59 -07:00
kshitijk4poor
311b980bb6 fix(prefix-cache): drop the workspace pin at session boundaries; bind session cwd for /context
The pin from the previous commit lives in _SESSION_STATE but nothing cleared it, so a CLI
/new, /resume or /branch (same AIAgent, reset_session_state + _invalidate_system_prompt)
replayed the previous session's git snapshot into the new session's prompt. Clear it in
reset_session_state next to the other session anchors; one invariant test (red without it).

The TUI/Desktop session.context_breakdown RPC ran the prompt builder on the RPC thread
with no session cwd bound, so it re-probed against the backend's cwd and overwrote the
session's pin — one /context between compactions restored the divergence this fix removes.
Bind the session context around the build like the live rebuild in server.py does.

Also: trim _coding_parts' docstring to the WHY, drop the isinstance/len guard on a value only
this function writes, and remove the tests' assertions on the private pin shape.
2026-09-06 22:45:31 +05:30
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
4fcd9ede33 simplify(compat): repoint the last stragglers (run_agent computer_use import, 8 test files patching old facade paths) 2026-09-03 15:15:53 -07:00
Teknium
912a497ce2 simplify(compat): tools/browser_tool — repoint run_agent cleanup_browser + chat_completion_helpers _is_headed_mode to the defining browser_tool_* siblings 2026-09-03 14:17:11 -07:00
Teknium
7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium
53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium
fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium
2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
Teknium
eb8a30cc2f simplify(compat): skills_sync/skill_manager/kanban/computer_use — drop 37 re-exports/aliases, repoint 8 callers + 9 test files 2026-09-03 13:27:01 -07:00
Teknium
8e1a5b9f46 simplify(compat): tools/wake_word — drop 5 re-exports, repoint 1 test 2026-09-03 13:24:36 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
fb14bc4e11 review-fix(whitespace): strip trailing whitespace and EOF blank lines introduced by this PR
Trailing-whitespace-only edits so 'git diff --check BASE HEAD' is clean
(16 diagnostics across 11 files). No code changes.
2026-09-03 09:31:54 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
561b053f79 perf(agents): run per-child timers on one shared scheduler thread
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog).  A
profiled session with ~130 children was carrying ~1000 threads.  All
of these timers now run on a single process-wide daemon thread.

- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
  one Condition-driven daemon thread.  schedule(fn, interval) -> handle;
  handle.cancel(wait=) blocks for an in-flight run like the old join.
  A callback returning False stops itself; a raising callback is logged
  at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
  scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
  idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
  stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
  _lease_refresh_interval; lease-lost / refresh-error interrupt paths
  and the stop-event fencing are unchanged; the join(timeout=1.0) is now
  cancel(wait=1.0) so the interrupt clear still runs after any in-flight
  tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
  schedule(); the poll body is _tick(), same sampling state machine.

Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132.  At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
2026-09-03 02:44:24 -07:00
Teknium
c96568f66c perf(delegation): finished delegate children no longer pin their transcripts in the parent heap
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:

1. bind_subagent_parent() stored the agent strongly in the
   `hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
   for its own turn, and every asyncio Handle/Future scheduled during
   that turn (LSP reader loops, kernel pipe transports) snapshots the
   Context — 56 live Contexts held 14 finished children after the bench.
   The ContextVar now holds a weakref (non-weakrefable doubles fall back
   to a closure); get_active_subagent_parent() dereferences it.

2. AIAgent.close() cleared _session_messages but not the
   _db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
   every successful DB flush) nor _streamed_assistant_text_parts, so the
   agent — kept alive by (1) — retained every message dict. close() now
   drops both.

The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.

Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
2026-09-03 02:35:37 -07:00
kshitijk4poor
914d8a0bd6 fix(recovery): name the real state.db in the copy-pasteable recovery banners
The gateway broadcast and turn-failure explanation printed a literal
~/.hermes/state.db; now that the line is a command the operator is meant to
run as-is, interpolate _default_db_path() so profile / HERMES_HOME installs are
pointed at the store that actually failed. Also: split a comment that a merge
fused onto the logger line in session_lost_and_found.py, and fix an inverted
test docstring.
2026-09-03 11:28:21 +05:30
sal
a15f96450b fix(recovery): make the printed salvage command satisfy the real CLI contract
Review blocker on e62940d: every state-db guidance site printed

  hermes sessions recover --source <db>

but cmd_sessions rejects that shape with exit 2 ("--output is required
unless --inspect-only is used") before any snapshot is taken — the user
follows the instruction during a corruption incident and gets nothing.

All five state-db sites now print the established two-stage operator
contract (the same shape `sessions repair` failure output and
docs/state-db-recovery.md already use):

  hermes sessions recover --source <db> --inspect-only
  hermes sessions recover --source <db> --output recovered-state.db

with the stop-the-gateway precondition stated for the gateway/turn
banners, and --inspect-only leading in the hermes_state refusal strings
(inspection before writing anything).

New TestEmittedCommandsSatisfyCliContract dispatches the exact emitted
flag shapes through the real cmd_sessions and asserts they pass the
contract gate (rc != 2) on a scratch DB, plus a premise test pinning
that the v1 no-flag shape is still rejected with rc 2 — so a guidance
string can never again pass a source-substring test while the command
it prints deterministically fails.

Noted for merge order: #101423 and #101168 also touch
hermes_cli/session_recovery.py. They are complementary recovery-integrity
work, not duplicates of this guidance/gate fix; whichever lands second
should rebase and rerun the lost_and_found + session-recovery suites.

(cherry picked from commit 34dc59a284509e76a0342c36d03a2a437aa8a3b9)
2026-09-03 11:28:21 +05:30
sal
5d9a2110ba fix(recovery): stop pointing sqlite3 .recover guidance at the live state.db
Refs #100368. The forensics thread established that a sqlite3 CLI with
the WAL-reset opener bug (fixed 3.51.3+ / backports 3.50.7 / 3.44.6;
Debian/Ubuntu system shells 3.45.1/3.46.1 are in the vulnerable band)
unlinks the live -wal/-shm pair when pointed at a live state.db whose
writer's DMS lock has been cancelled, splitting the store into two
concurrent generations whose acknowledged writes vanish while both
report integrity_check ok. Hermes' own corruption banners instructed
exactly that command.

- gateway corruption broadcast, run_agent corrupt-cause explanation,
  hermes_state repair-budget and forensic-backup refusals, and the
  kanban manual-recovery hint now route operators to
  `hermes sessions recover --source <db>` (which snapshots the damaged
  bundle before any shell touches it) and warn against a raw sqlite3
  shell on the live file
- find_sqlite3_cli() now refuses a WAL-reset-vulnerable shell for the
  page-level salvage lane even on the snapshot, reusing the canonical
  gate from hermes_cli.sqlite_runtime so the embedded runtime and the
  salvage shell can never disagree
- find_sqlite3_cli_refusal() records why a shell was refused so the
  lost_and_found lane can tell the operator exactly what to install
  instead of a generic "not found"
- regression tests cover the version gate (vulnerable/fixed matrix, the
  mirror check), every refusal reason, and each guidance site

Test plan:
- scripts/run_tests.sh tests/hermes_cli/test_sqlite3_cli_salvage_gate.py
  tests/test_state_db_repair_loop_cap.py
  tests/run_agent/test_corruption_recovery_guidance.py
  tests/hermes_cli/test_session_recovery_lost_and_found.py
  tests/hermes_cli/test_session_recovery.py tests/test_sqlite_wal_reset_gate.py
  tests/hermes_cli/test_sqlite_runtime.py - 91 passed, 1 skipped locally

(cherry picked from commit e62940d1021e80e9b7d6423ced1cbdfe7dd0c37d)
2026-09-03 11:28:21 +05:30
Teknium
d68a4c01a7 refactor(run_agent): restore main() docstring verbatim (fire renders it as --help output) 2026-09-02 22:51:39 -07:00
Teknium
55385a8111 refactor(run_agent): collapse defensive layers in close/todo-hydration/listing paths; -31 LOC 2026-09-02 21:53:57 -07:00
Teknium
116ca1db1e fix: sibling Nous 401 recovery adopts a peer's refresh instead of rotating again
The stampede fix added a `stale_access_token` hint to
resolve_nous_runtime_credentials() so a process whose bearer just 401'd
adopts a token a sibling already rotated instead of re-POSTing the shared
grant — but only the credential-pool caller passed it. The main agent's
401 path (run_agent._try_refresh_nous_client_credentials), the auxiliary
client rebuild, and the proxy adapter all called force_refresh=True with
no hint, so `_already_rotated_by_peer` could never fire: N subagents
hitting hourly expiry still issued N serialized refreshes, each one
invalidating the token a sibling had just adopted.

Live 12-process A/B against a fake Portal: 12 refresh POSTs / 9 distinct
final tokens before, 1 POST / 1 token after.
2026-09-02 21:22:07 -07:00
Teknium
a208c541f1 refactor(run_agent): fold codex/sanitization forwarders into staticmethod aliases, pack multi-line signatures/calls; -198 LOC 2026-09-02 21:04:11 -07:00