prompt() dispatches slash commands on a worker thread before claiming the
turn, so /reset and /compress could clear or rebind state.history (and null
agent._session_db) underneath a live run_conversation. /model and
session/set_model could likewise swap state.agent mid-turn, which also made
_finish_turn emit a spurious compression-rotation update.
Mutating commands now hold a per-session command_op flag for their whole
run: they are rejected while a turn or another op is in flight, prompts
arriving mid-op queue instead of claiming the turn, and the queue drains
through the same helper _finish_turn uses. Gateway parity: these commands
are idle-only there.
backfill_acp_session_cwd had no production caller, so rows minted before
the column was written stayed unassigned in Desktop until someone ran it by
hand. The manager now runs the idempotent UPDATE once per process on first
DB use (injected or acquired), with a test through create_session().
Also maps the contributor email for attribution.
ACP sessions stored their workspace only inside the model_config JSON blob
(a correct choice when the cwd column did not exist yet), so Desktop, the
Projects sidebar and hermes sessions list showed every editor session as
unassigned. create_session now passes cwd, an existing row (the live path,
since the agent flushes the transcript incrementally) gets the column
promoted via update_session_cwd, update_cwd() moves it on reopen, and git
branch/root are probed off the interactive path under the same generation
contract tui_gateway/session_workdir.py uses. backfill_acp_session_cwd
promotes model_config.cwd for rows minted before this change.
Squashed from the five commits of #115707; the accidentally committed
Windows cache files under %SystemDrive% are dropped.
Multi-file V4A proposals set EditProposal.path to a comma-joined display
string, and should_auto_approve_edit evaluated it as one path: the
sensitive-name check saw only the last segment and the workspace check
resolved the joined string under the session cwd. A patch touching .env
plus a normal file auto-approved under 'session', and one carrying an
absolute path outside the workspace auto-approved under
'workspace_session' — both without the interactive prompt.
EditProposal now carries a paths tuple with every real target; the
sensitive check runs any() and the workspace check all() across them,
falling back to (path,) for single-file proposals. The joined string
remains for display only.
Fixes#115213
_finish_turn persisted, emitted provenance/final text, then released
is_running/current_prompt_text with a bare block before draining queued
prompts. Any exception mid-tail skipped the release, wedging the session
running and stranding the queue. Indent the tail into try; the finally
releases is_running/current_prompt_text first, then drains the queue.
Regression test raises on the final-text session_update and asserts the
session goes idle and the queued prompt still drains. (#115588)
Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.
Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
resolve_runtime_provider() selects a provider-scoped credential pool and returns
it as runtime["credential_pool"]; oneshot and the gateway hand it to AIAgent, but
acp_adapter/session.py::_make_agent dropped it, so a long-lived ACP process had
_credential_pool=None and could not refresh/rotate on HTTP 401 after OAuth
token expiry — the only recovery was restarting the ACP process (#70292).
Forward the pool by identity like the other surfaces. The pool is already
provider-scoped and its selected entry matches the agent's initial api_key, so
the existing account-isolation guards are preserved rather than bypassed.
Salvaged from PR #70293 (the cherry-pick claimed in #77029 never reached
acp_adapter/session.py); regression test asserts the pool object is retained.
Fixes#70292
`SessionManager._make_agent` and the Feishu doc-comment agent built their
`AIAgent` without `reasoning_config`, so `agent.reasoning_effort: none`
never reached those sessions: the transport applied its default effort,
which non-reasoning models such as gpt-4o-mini reject with HTTP 400 and
which silently re-enables thinking everywhere else. Both surfaces now go
through `hermes_constants.resolve_reasoning_config`, the same chokepoint
the CLI, gateway, TUI, cron and `hermes -p` already use, resolved against
the model the session actually runs so per-model overrides apply.
Ported from PR #85164 by @Chinmayrawat15 (oneshot hunk already on main).
Fixes#85153
The warm-up imported only the provider module. holographic / mnemosyne import
numpy at module top, but hindsight defers the ML stack to is_available() ->
_check_local_runtime() (importlib of hindsight / sentence_transformers), which
ran later on a to_thread worker racing acp-mcp-discovery — the reported hindsight
stack was still reachable. The deadlock partner is numpy's lazy _core init in
every reporter's dump, and a plain `import numpy` up front was every reporter's
workaround, so import_memory_provider_module now also imports numpy (best-effort)
once the provider module is in.
Also: import_memory_provider_module() defaults to the configured memory.provider,
so entry.py drops its duplicate config resolver and outer try; the "ONLY thread"
comment is reworded — hermes_cli's plugin-discovery thread is already running
when hermes acp dispatches.
prompt / cancel / set_session_model / set_session_mode / set_config_option still
called session_manager.get_session inline. For an id not in memory that runs
_restore -> _make_agent (config, memory-provider import, SessionDB) on the loop —
the hang class session/new just left — and, since restores are single-flight, it
also parks the loop on _restore_lock while an off-loop session/load is in flight.
Route the five sites through asyncio.to_thread like new/load/resume/fork.
Test: a parametrized invariant over the five handlers with a slow DB restore;
ticks=0 on the previous head, green now.
Every faulthandler dump in the thread shows session/new stuck in numpy's
create_module on the main thread while another thread (MCP discovery / ACP
stdin reader) sits in the same lazy import chain — a first-time native
extension import racing another thread deadlocks on Windows (holographic,
mnemosyne and hindsight all reproduce it; a sitecustomize `import numpy`
before any thread exists resolves it every time).
hermes acp now imports the configured memory.provider's module on the main
thread before the MCP-discovery thread and asyncio.run() start (Windows only —
the deadlock is Windows-specific and the import is paid once either way).
plugins.memory.import_memory_provider_module imports the module without
constructing a provider or running register(); the agent build later finds it
in sys.modules.
Trimmed from #91775 (@tigercraft4): same placement and gating; reuses the
existing plugin loader instead of a second module-import routine.
Co-authored-by: tigercraft4 <tigercraft4@tigercraft4.com>
session/new, session/load, session/resume and session/fork constructed a full
AIAgent (config load, memory-provider import, SessionDB) inline in the request
coroutine, freezing the loop that serves every JSON-RPC request — a host saw an
agent that answered initialize and then nothing, with no error anywhere.
Run the construction through asyncio.to_thread. Because restores can now
overlap, SessionManager.get_session serializes the DB-restore path under a
lock and re-checks the in-memory map, so two session/load for one id share one
agent build.
Slimmer redo of #85001 (@SHL0MS): same direction, without the async wrapper
layer and future map — the lock + re-check gives the same single-flight.
Co-authored-by: SHL0MS <SHL0MS@users.noreply.github.com>
The ACP permission and edit-approval bridges self-denied after a fixed 60 s
while the editor's approval card was still waiting (raised on #73403 by an ACP
host maintainer). Default the bridges' timeout to the existing approvals.timeout
config knob (300 s, same resolver as CLI/gateway prompts), read per request.
The `except ValueError` in set_session_model wrapped both the switch_model rejection and
the _make_agent rebuild, so a rebuild ValueError (provider disabled in config, context
window below the floor) was reported as -32602 by accident. _switch_model now raises a
dedicated ModelRejected(ValueError) at the rejection site and set_session_model catches
only that; rebuild ValueErrors keep the -32603 internal-error path. The slash /model path
still sees the rejection text via str(exc).
set_session_model already validates through hermes_cli.model_switch.switch_model
(11576390fe), so an unadvertised modelId is refused before the session mutates
(#72439's main atom). The rejection surfaced as JSON-RPC -32603 "Internal error"
though, which clients attribute to the agent rather than to the request; it is now
RequestError.invalid_params (-32602) carrying the switch_model reason. _switch_model
also assigned state.model before the rebuild, so an agent-build failure left the
session persisted on a model the live agent did not run; the assignment now follows
the successful build.
_make_agent swallowed a resolve_runtime_provider failure at debug and built a bare
AIAgent, which dies with the first-run "No LLM provider configured. Run `hermes
setup`" text on a configured machine (#91090's residual ask). The fallback stays, but
when the bare build fails the swallowed resolution error (revoked OAuth, disabled
provider, ...) is raised instead, chained to the fallback failure.
Direction credited to @z0zero (#72579: -32602 + atomic session state) and
@webtecnica (#91100: do not swallow the resolution failure).
Co-authored-by: z0zero <z0zero@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
The closer in await_permission() only recognised an allow when the response
outcome was the SDK AllowedOutcome class, while the edit-approval requester
duck-types (outcome == "selected"). A client answering with a plain selected
outcome therefore had its edit applied but the edit-approval-N bubble closed
"failed". Duck-type the closer on the wire discriminator so its terminal
status always matches the decision. Live-pass side-effect on this PR.
request_permission carries a synthetic perm-check-N / edit-approval-N
ToolCallUpdate in status pending; clients materialise it as its own
bubble and nothing ever moved it on, so it spun forever. await_permission
now takes an optional send_update and, once the outcome is known, closes
that id: completed when an allow option was selected, failed on deny,
timeout or a failed request. The server wires it for both the command
approval callback and the edit approval requester.
tool.completed carries is_error, but the ACP bridge dropped it and
re-derived status from the result text alone, so a tool cancelled by a
user interrupt (plain-text result) or one returning an error dict closed
as completed. Pass the flag through close_tool_call and OR it into
build_tool_complete's failed predicate; the text heuristic stays as the
fallback for the step-closer path.
The turn-end sweep ran on the loop thread, where _send_update's
future.result(timeout=5) stalls the loop for 5s per open call and the
terminal update only lands after the response. Flush in the executor
body's finally instead, which also covers the executor-exception path
with one call site. Test drives prompt() with an open tool.started on
both paths: red on origin/main (no close) and on the previous head
(loop stalled), green now.
Follow-up to the salvaged #114442 (@hteo1337), which closes each ACP tool call
from its own ``tool.completed`` (the mechanism PR #50741 by @liuhao1024 filed in
June, and PR #27854 by @godlin-gh filed first in May via ``tool_complete_callback``) and fails whatever is still open at turn end, draining
``tool_call_ids`` and ``tool_call_meta`` together.
Live repro of the fallback path exposed a third mechanism the issue did not
name: ``make_step_cb`` passed ``prev_tools[i]["arguments"]`` — the wire JSON
*string* — as ``function_args`` into ``build_tool_complete``, whose content
builders index it as a dict. The ``.get`` on a str raised inside the swallowed
step callback, so a ``write_file`` close never reached the client even when a
next step existed. Coerce with ``coerce_tool_args`` (meta args when absent).
Tests: the contributor's six tests folded into two invariants — a call is
closed exactly once from ``tool.completed`` with the step closer standing down;
the step fallback survives JSON-string arguments and the turn-end flush fails
the remaining call while emptying BOTH per-turn dicts.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: godlin <ganlinbupt@gmail.com>
The step callback only fires on the next step, so a turn's last tool calls stayed in_progress forever, and a blocked or permission-denied call projects no tool.completed at all. Close each call from its own tool.completed event, stand the step-callback fallback down once completions arrive, and fail whatever is still open when the turn ends.
Terminal-failure paths (HTTP-200 content-policy refusal, ``_Trunc.end_turn``, retry
exhaustion, interrupt before any assistant text) persist the accepted user row and return
before ``finalize_turn``, so ``user`` stays the durable conversation tail. The next prompt
appends a second user row, ``repair_message_sequence`` merges the pair, and the provider
is asked to act on the failed request again. The gateway compensates with
``_hmwa_close_failed_turn`` (#108033); standalone ACP, the CLI and the TUI/Desktop hand
``result["messages"]`` straight back as history and had no closer.
Close it once, at ``agent/conversation_loop.py::run_conversation`` — the seam every
envelope leaves through — with a Hermes-authored assistant boundary
(``agent/turn_failure_copy.py::FAILED_TURN_NOTICE`` / ``PARTIAL_FAILED_TURN_NOTICE``,
which the gateway now aliases instead of keeping its own copy). Idempotence is keyed on
``SessionDB.latest_conversation_role`` (durable state, not content), so a redelivery or a
tail another writer already closed is a no-op and the gateway's closer no-ops in turn.
The context-pressure classes (``compression_exhausted``, ``compression_deferred``,
``failure_reason == "context_overflow"``) are excluded: appending to an oversized session
is the #1630 growth loop; their repair is rotation.
Adjacent defect from the same report: ``acp_adapter/server.py::_finish_turn`` called
``final_response.startswith`` on ``None`` for an interrupted turn — the same one-line fix
PR #64471 by @israellot filed first (its wider prompt()-restructure is superseded by the
current ``_finish_turn`` shape).
Slimmer redo of #114168 by @kendrickkester (same seam and invariants; the +1023-line
PR carried a new copy module, an accepted-turn re-anchoring scan and an 859-line suite).
Two invariant tests: the real ACP path (loopback provider, refusal then a new prompt) and
the durable-tail idempotence / overflow exclusion.
Co-authored-by: Kendrick Kester <kendrick.kester@gmail.com>
Co-authored-by: Israel Lot <israel.lot@gmail.com>
Review follow-up: asyncio.to_thread already runs its callable inside
contextvars.copy_context(), so wrapping _register_pinned in a second copy was a
no-op. The cwd pin set inside the worker still does not leak back to the caller
(tests/acp_adapter/test_server.py pin assertion unchanged, still green).
The default stdio child cwd now reads agent.runtime_cwd.resolve_context_cwd()
at spawn time (previous commit), but ACP new_session/load_session register the
client's MCP servers outside the per-turn cwd pin, so a hosted ACP session with
logical cwd /workspace/a — the reporter's exact scenario — still spawned its
stdio servers in the hermes-acp process directory. Run register_mcp_servers in
a copied context with set_session_cwd(state.cwd); run_coroutine_threadsafe
carries that context onto the MCP loop task and the long-lived server task
copies it, so reconnects respawn in the same directory.
Also: trim the salvaged tests to the two invariants (session pin becomes the
default; explicit config cwd wins) — the TERMINAL_CWD fallback and the
missing-directory→None cases are runtime_cwd's own contract, pinned in
tests/agent/test_runtime_cwd.py, and the native default (no anchor → None) is
already pinned by test_start_preserves_native_default_cwd. Document the `cwd`
key and its default in the MCP docs.
Shared-process caveat recorded in the PR body: MCP connections are per
process/registry scope, not per chat session, so the anchor is read once at
connect time (profile-level terminal.cwd for gateway/cron; the owning session
for ACP-provided servers).
Closes the remaining atoms of #112600.
A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
its siblings still resolved credentials against config's `default`: the CLI
auth-fallback rung, `--resume` credential re-resolution, the gateway
provider-override helper (channel overrides, persisted /model switches,
API-server provider refresh), the gateway fallback chain, the TUI /model
switch-from runtime and ACP agent construction. With a `*-free` default the
OpenCode free-tier rung fired first and a Go-only model was built against
the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
optional `target_model` and the two test stubs of it accept the kwarg.
B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
provider matched by opencode_provider_family, including custom providers
merely named after a family (`opencode-go-bridge`, #85589) whose relay the
user declared explicitly in `providers:`. The family heal now applies to the
built-in canonical providers only; custom prefix-named providers keep their
per-model api_mode routing and /v1 handling. Documented in the providers
guide.
C) Same function: the official-host check uses parsed.hostname (a port no
longer defeats the heal) and only the path is edited, so query/fragment
round-trip instead of being dropped.
Fixes#112600
Salvage of #93452 (@outpoints). Keep one invariant per fix:
resolver-level "automatic title never remaps a strategy session",
integration "provider routes by logical workspace, not process cwd",
agent-level "title provenance + cwd reach the provider", deferred
Desktop/TUI build threads the session cwd, seeded branch titles are
derived, and the workspace-move E2E. Drop the plumbing/legacy-shape
tests that re-assert the same contract.
acp_adapter: AIAgent(cwd=...) now stamps session_cwd itself, so the
direct assignment after construction was a duplicate.
build_model_state dropped every inventory row whose slug name matched a
named `providers:` entry and relabelled the current provider `custom:<key>`
unconditionally. When a `providers:` key shadows a canonical provider name
(providers.openrouter: -> proxy with a models: list), a session genuinely
running on openrouter.ai lost all canonical rows and its current id became
custom:openrouter:..., which resolves to the proxy base_url — picking the
"current" row silently re-routed the session to a different endpoint.
Only rows flagged is_user_defined are replaced by the named catalogs now,
and the current provider is promoted to custom:<key> only when the session
base_url matches that entry's api_url or no canonical row for the raw key
exists (the reported relay/MixedCase class keeps its custom:<key> id).
Review finding: named entry shadowing a canonical provider dropped canonical rows and re-routed the current session to the proxy.
Gateway, dashboard, ACP server and the CLI already hold one registry-shared
SessionDB per state.db path, yet several in-process call paths still opened
a bare SessionDB() beside it. Each one is a full writer: schema init, write
lock, token-writer thread and a close-time WAL checkpoint. On a dashboard
serving overlapping requests that stacked up to the "5 live SessionDB
handles" precursor within seconds; per #110544 they are now harmless to each
other's WAL generation, but the leak itself remained.
Pure readers attach read_only=True (no writer connection, no write lock):
- plugins/hermes-achievements/dashboard/plugin_api.py::scan_sessions
(dashboard, per background scan and per /rescan; highest-frequency site)
- hermes_cli/console_engine.py::_session_db (dashboard console; list,
stats and export are reads; rename/optimize opt in to a writer)
- tools/process_registry_results.py::_owns_result (gateway, per retained
result load)
- hermes_cli/main.py::_session_db (last-session / title / cwd lookups)
- hermes_cli/terminal_breadcrumbs.py, hermes_cli/status.py,
hermes_cli/main_tui_launch.py (lookup one-shots)
Writers share the process's registry handle (hermes_state_registry.acquire;
close()/release_or_close release one refcount):
- acp_adapter/session.py::SessionManager._get_db — the AIAgent it builds
acquires the same path, so the ACP server held two writers per process
- hermes_cli/kanban_db_dispatch.py::_retag_legacy_worker_sessions
(gateway dispatcher tick)
- hermes_cli/main.py::_create_titled_session, hermes_cli/oneshot.py,
hermes_cli/foreign_sessions.py — the CLI acquires the same handle a
moment later
The "N live SessionDB handles" warning now counts only writable members:
read-only attaches are the sanctioned per-request shape for dashboard
routers and CLI lookups, and counting them turned a healthy topology into
an operator alarm (#100896 field reports of restart loops keyed on it).
Live repro (one registry writer + 4 overlapping dashboard/gateway paths in
one process): before 5 writable opens, 5 live handles, warning fired;
after 1 writable open (the registry handle), 0 from the request paths,
no warning.
Refs #100896#103339
The ACP schema (agent-client-protocol 0.9.0, ContentChunk.messageId) says
"Both clients and agents MUST use UUID format for message IDs". The
ported allocator emitted hermes-assistant-N strings, which a strict
client may reject or fail to group. A fresh uuid4 per message keeps the
grouping semantics and can never collide with an earlier turn's id, so
the counter/prefix state is gone.
Tests trimmed to three invariants: chunks share one UUID until the None
flush sentinel (empty deltas ignored), thought + text share an id, and
the no-allocator shape stays unchanged.
ACP clients group streamed agent_message_chunk / agent_thought_chunk
updates into one assistant reply by messageId, and use a new id to
start the next reply (root-reply replacement semantics). Hermes' ACP
adapter sent every chunk without a messageId, so clients that replace
'the current assistant message' per chunk collapsed separate
autonomous turns into one bubble.
- AssistantMessageIdAllocator (per ACP session, monotonic across
turns): a contiguous run of reasoning + text deltas shares one
hermes-assistant-N id; the None flush sentinel Hermes core emits
before tool execution / at end of stream closes it.
- make_message_cb / make_thinking_cb stamp update.message_id when an
allocator is provided; legacy no-allocator shape unchanged.
- Unstreamed final responses open their own id; plugin-transformed
responses reuse the streamed message's id (replacement).
- Tests: grouping until flush, thought+text sharing, monotonic ids,
empty-string vs None sentinel, legacy shape.
set_session_model and the /model slash command called switch_model synchronously
on the loop thread; it does ~10 s of network I/O on a cold cache (models.dev,
custom-endpoint probes), stalling every ACP session in the process. Same
asyncio.to_thread the gateway uses for the same call.
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.
acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.
`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
`_finish_turn` gated the history update on truthiness, so a turn that
legitimately returned an empty transcript left the previous history
in place and the client kept rendering stale state. Same falsey-vs-absent
class as the step-callback fix (#10845); gate on key presence.
Refs #10844
make_step_cb() used 'or' to fall back from 'result' to 'output' key:
result = tool_info.get('result') or tool_info.get('output')
This collapsed valid falsey results (empty string, 0, False) to None,
because Python's 'or' operator treats all falsey values as missing.
Use explicit key presence check instead:
result = tool_info.get('result') if 'result' in tool_info else tool_info.get('output')
This preserves the actual value when the 'result' key exists (even if
falsey), and only falls back to 'output' when 'result' is truly absent.
Adds regression tests for empty string, zero, and output-key fallback.
Fixes#10845.
`_persist()` prepared `session_meta` (cwd + provider/base_url/api_mode) but
the create path wrote only `{"cwd": ...}`; the snapshot only landed on a
later update. A restart before that update restored the session with
provider/base_url = None.
Use the same `session_meta` for create and update.
Cherry-pick of PR #9883 by @Ruzzgar (release.py mapping hunk dropped —
already mapped on main).
Fixes#9812
session/set_model with "anthropic:claude-sonnet-5" while already on anthropic
parsed to the same provider, so the bare-name fallback detect_provider_for_model
ran and could hand the session to OpenRouter because the bare name exists in
its catalog. Detection now runs only when parse_model_input consumed no prefix.
Diagnosed by AlexFucuson9 in #59294; same fix, expressed as "input unchanged
by parsing" so it needs no provider-name table.
Keep two invariants covering empty create/cwd/save/fork, genuine content and existing-row metadata. Preserve original authorship and avoid source-only legacy pruning: an empty ACP row does not prove its owner is dead. Native ACP wire plus a local streaming model fixture verifies the first turn and nonempty fork remain durable.
Fixes#104724
The ACP wire contract has session/new as a separate round trip from
session/prompt precisely so a client can open a session before it knows a
prompt is coming, and at least one shipping client opens sessions it will
never prompt: bb discovers a model catalog by spawning a throwaway
`hermes acp`, sending initialize + session/new, reading
NewSessionResponse.models, and killing the process. Its discovery cache
TTL is 60s, so an open editor re-probes continuously.
_persist() created the state.db row the moment a session was created, so
every such probe left a permanent message_count=0 shell. Measured on one
workstation: 116 accumulated, appearing on a time cadence rather than a
per-conversation one (26 real ACP user turns vs 13 shells in a day; exactly
120.0-minute spacing overnight with zero user activity).
The shells are indistinguishable from real chats in the session list and
`hermes sessions prune` cannot remove them: the rows are never ended, so
ended_at stays NULL and prune skips them by design, leaving
`hermes sessions delete <id>` one id at a time as the only remedy.
Defer row creation until the session has history. Nothing is lost for a
genuine conversation: AIAgent._ensure_db_session() creates the row on the
first turn and the post-prompt save_session() lands the ACP metadata on
top, which makes the create-time write redundant for every session except
the one case that should not be recorded at all.
Gate on state.history rather than message_count so fork_session, which
deep-copies a non-empty history into a fresh id, still persists at once.
Verified by differential execution of an identical probe against both
trees: on 693641aa8b an unprompted session/new adds a row, with this
change it adds none, and a session that does receive a prompt persists
in both. tests/acp: 138 passed. Reverting only the source change leaves
the two new assertions failing, confirming they exercise the fix.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
tools/approval.py no longer re-exports sibling names (approval_context/prompt/floors/detection/
human_wait/smart/gateway_wait); it imports only what it uses. Siblings reference sibling-defined
names directly (module-attribute reads on tools.approval_context so patching the defining module
still works); only facade-owned state (_lock, _gateway_queues, _permanent_approved, _denied,
_denial_breaker_addendum, _gateway_notify_cb) is still read back through tools.approval.
approval_detection calls its own _command_detection_variants instead of late-binding through the facade.