Commit Graph

5626 Commits

Author SHA1 Message Date
John Paul Soliva
3591bf3bb8 fix(compression): turn-start in-place compaction no longer hides the newest summarized turn (#120187)
archive_and_compact takes the newest tail_count durable rows as the carried tail's superseded
originals (active=0, compacted=0: hidden from display and session_search). The CLI and the gateway
persist a turn's user row only after the turn-start preflight, so at that compaction the row rides
in the carried tail with no durable original of its own. The positional rewind then reached one
row past the tail and flagged the newest summarized message as a superseded duplicate. Each
turn-start compaction hid one more (A5, then A16 in a two-compaction run), on the default config.
The TUI/Desktop are immune because they persist the user row at submit.

tail_count now leaves out this turn's rows that never reached state.db: no persisted marker and
no _row_id, counted within the carried tail window only. That includes the user row and any
unflushed scaffolding. It does so only while the turn holds the session turn lease. Between turns
(a manual /compress, the gateway's pre-turn hygiene) the turn anchor is left over from the last
turn, and a gateway transcript reload is durable but unmarked. Counting those rows would leave
carried originals at compacted=1, the duplicate recall #86366 fixed.
2026-09-23 10:05:30 -04:00
beardthelion
0cbe552888 fix(scope): failed profile-scoped secret reads must not borrow os.environ
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.

Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
2026-09-23 06:46:02 -07:00
beardthelion
2c3a75beaa fix(secret-scope): scoped misses fail closed under a foreign-home scope
get_secret returned os.environ on a scoped miss whenever multiplex was
off, but non-multiplex hosts serve foreign homes too (dashboard/desktop
backend, per-profile cron, MCP owner scopes, kanban spawn-env builds),
where os.environ is the launch profile's. Bound scopes now carry the
home they were built for; serves_routed_profile detects a foreign scope
even when the binder deliberately skips the HERMES_HOME override, and
the miss returns the caller's default. Every production binder stamps
its home; own-home scopes keep the deliberate env overlay.
2026-09-23 06:46:02 -07:00
c0d1ngHUB
857e651b5f fix(relay): the logical LLM close must not pop through a concurrent turn's scope
The logical LLM scope close in ``relay_llm._complete_logical`` popped its handle with
the unguarded ``pop_relay_scope``. Two consecutive calls in one session share a physical
scope stack, so when the sibling turn's live scope sat above the handle the native
binding raised::

    RuntimeError: invalid argument: scope handle is not at the top of the stack

``_complete_logical`` catches that, logs "logical LLM finalization failed" with a
traceback and returns early, so the early-return path also skips the
``turn.logical_llm_calls`` cleanup until the handle is retried. Observed in production as
one traceback per overlapping turn (~14/day on a busy local profile).

``pop_relay_scope_if_top`` already exists for exactly this case and is used by the
shared-metrics task close (PR #116685, #115471); the logical-LLM seam was missed. Guarding
``pop_relay_scope`` itself is NOT an option — ``_pop_with_drain`` relies on that raise to
detect a stacked sibling and drain it.

The skipped scope is reclaimed by the existing session-close drain
(``RelayRuntime._close_scope_handle``), so the handle still leaves
``logical_llm_calls`` and the session unwinds without an orphan.

Test: ``test_logical_close_skips_pop_under_concurrent_turn_scope`` pushes a sibling scope
in the session context, completes the logical call, and asserts the sibling (not ours) is
still on top and that the sibling's scope survives. It fails on upstream ``main`` with the
exact ``RuntimeError`` above and passes with the guard.

Fixes #115471 (the remaining call site).
2026-09-23 06:31:43 -07:00
teknium1
eb8960fead fix(api_server): keep one memory provider per session across requests (#120116)
The api_server platform builds a fresh AIAgent per request (per-request
callbacks, model route, ephemeral prompt), so the memory provider was
re-initialised on every request. External providers deliver recall as the
PREVIOUS turn's background prefetch held on the provider instance, so a
continued session (X-Hermes-Session-Id, previous_response_id, declared
session key) never received automatic recall, and for hindsight
local_embedded each init also restarted the embedded daemon, killing the
retain still in flight. Pre-existing: the same probe fails on main before
the hindsight catalog migration (526d135a96, bundled provider).

ApiServerMemorySessions parks the session's initialised MemoryManager
between requests (exclusive check-out/check-in, keyed by profile home +
session id, LRU/idle eviction under the owning profile's scope) and
AIAgent(memory_manager=...) adopts it instead of loading and initialising
the provider again. /v1/chat/completions, /v1/responses, session chat and
/v1/runs all go through the same two seams (_create_agent, turn finally).
2026-09-23 05:41:31 -07:00
teknium1
d9ec0bcf56 fix(relay): read segments config through the gateway's own reader, keep -z in its launch dir (#95577)
Follow-up to the salvaged #88281 commits: read gateway.telemetry.session_segments
through hermes_cli.config_effective.load_user_config_effective(), the exact
primitive gateway.run._load_gateway_config() now delegates to, instead of
read_raw_config() + managed overlay. The raw read dropped ${VAR} expansion
(max_turns: ${SEG_TURNS} parsed as 0) and model-key canonicalization.

The same import also rebound TERMINAL_CWD to the home dir in `hermes -z`
(which does not pre-set it like cli.py): the launch dir's AGENTS.md was
never loaded and terminal commands ran in $HOME (#95577). The subprocess
guard now asserts TERMINAL_CWD / HERMES_QUIET / _HERMES_GATEWAY stay unset
too, and the fixtures patch the effective reader, which both the old and
new call paths go through.

Found by the core entrypoint parity E2E matrix (oneshot row, context_file).
2026-09-23 05:39:26 -07:00
Mr.XYS
853312c1fc fix(relay): stop importing gateway.run from non-gateway hosts (#87183)
gateway.run executes os.environ["_HERMES_GATEWAY"] / HERMES_QUIET /
HERMES_EXEC_ASK at module top level. _segments_config() late-imported it,
so any CLI/TUI/desktop/cron process hitting the relay turn lifecycle had
HERMES_EXEC_ASK=1 written into its own environ, routing dangerous-command
approvals onto the gateway path (tools/approval.py:3344) where the CLI has
no notify_cb — the approval panel never renders and commands hang forever
in pending_approval.

Read the gateway telemetry config via read_raw_config() + managed overlay
instead (no import side effects), and point the test fixture at the new
read path.
2026-09-23 05:39:26 -07:00
teknium1
2ad268b314 fix(agent): re-anchor current_turn_user_idx before every request after mid-turn compaction
Mid-turn compaction rebuilds `messages` without always handing back a new
current_turn_user_idx: the pre-API pressure gate (run_preflight_compression) returns
"continue" with the pre-compaction index, and any future compaction site would do the
same. The request builder splits the replay prefix at that index and canonicalizes
messages[:idx]. With a stale index the split lands inside the current turn's tool rows:
canonicalization drops the assistant tool_call whose result fell past the split, the
sanitizer then drops the orphaned result, and the model silently loses the turn's
earlier tool output even though state.db (and the in-memory history) still hold it.
persisted != sent, and a resumed session sends a different history than the live one.

prepare_iteration now validates the index before every request (it must land on this
turn's user row, verbatim or via its user-originated view) and re-anchors otherwise,
so every mid-turn compaction path is covered in one place. The post-tool gate's own
re-anchor (previous commit) keeps _persist_user_message_idx aligned immediately; this
check is a no-op there.

Found by the compaction E2E suite (tests/e2e/core/compaction, persisted-prefix
invariant): a seeded tool-heavy session that crosses the trigger after a tool round.
2026-09-23 05:36:34 -07:00
Andrey
f27417fc40 fix(agent): re-anchor after post-tool compression 2026-09-23 05:36:34 -07:00
teknium1
24f03c41c1 fix(agent): hosted providers no longer get the local-server "wait and /retry" context rejection
The unexplained-rejection gate from #114644 ended the turn with "another request on the same
server was probably holding its capacity ... wait and /retry" whenever a server said "context
exceeded" without a count while the local estimate sat under half the known window. That cause
only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any
public endpoint) the same rejection means the route's real window is smaller than the one Hermes
assumes, so /retry failed identically every turn and the conversation was never compressed:
Discord bots on claude-opus-5-5 were stuck repeating the message.

The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container
DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ
entry says which endpoints get the message.
2026-09-23 04:38:27 -07:00
teknium1
4b17b284ce test: keep cross-cutting architecture lints; drop stale OS-fake baseline entries
Restore the cheap repo-wide guards the per-file triage classed as source reads
but that protect recurring bug classes (<2s total):
- subprocess env scrubbing near spawn sites (credential leakage)
- gateway UTF-8 encoding= on file I/O (Windows mojibake)
- no raw yaml.safe_load of config.yaml (lost ${ENV} expansion)
- CLI subprocess.run timeouts (hung CLI)
- no locked readers on the shared state.db connection (#99349 segfault)
- CI classifier outputs / live-comment watch list match real workflows
- relay imports no platform crypto (relay trust boundary)
- Desktop relay deliver budget mirrors the Python deadlines (#93911)
- no native title= on Desktop buttons (DESIGN.md rule)

Drop _BASELINE entries in check_os_marker_fakes.py for files that no longer
fake macOS (the checker fails on stale entries), and remove doc/comment
pointers to deleted tests.
2026-09-23 03:15:26 -07:00
teknium1
b3624cc3af docs(imessage): hints match the stripped-on-link send path
Photon messages with a link now arrive stripped rather than raw, and BlueBubbles
keeps link URLs, so the hints stop describing the old leaks.
2026-09-23 03:10:15 -07:00
teknium1
6fa6df75f2 fix(imessage): platform hints steer replies to plain texting style
The Photon hint told the model "Markdown is rendered (bold, italics, lists,
code)", so replies came back with headers, code fences and backticks. That
is wrong on the real delivery path: the sidecar sends any message containing
a URL as raw text (every *, #, ``` and | shows literally), and even on the
markdown path spectrum-ts flattens headings to bold, tables to "a | b"
rows, and turns code into Unicode math-monospace glyphs that break when
copied. The hint also claimed attachments are metadata-only, which stopped
being true when native media send landed.

Both iMessage hints (Photon plugin, BlueBubbles built-in) now ask for a
texting register: short, answer first, no headers/tables/fences/backticks,
commands on their own plain line so they copy, bare URLs. BlueBubbles also
notes that strip_markdown drops the URL of [text](url) links.
2026-09-23 02:36:34 -07:00
John Paul Soliva
526d135a96 fix(compression): keep the /compress here N tail in state.db under in-place compaction
compress_now() hands only the head to _compress_context() and rejoins the
kept exchanges in memory afterwards. With compression.in_place (the
default), the commit's archive_and_compact() archives every active row at
or below the lease watermark, which includes the kept tail's rows, and
inserts only the compacted head. Its tail_count rewind then lands on the
newest rows under the watermark, which are the kept tail itself, so those
rows end up active=0, compacted=0 with no live copy. No surface writes them
back: the CLI re-flushes only after a rotation, the gateway skips in-place
on purpose, and the TUI only swaps its in-memory history. The exchanges the
user asked to keep verbatim are gone on resume, and on the gateway from the
very next message.

compress_now() now passes copies of the tail as verbatim_tail. The in-place
commit stores head + tail in the same archive_and_compact() transaction,
joined with the same seam rejoin_compressed_head_and_tail() builds in
memory, and adds the tail to tail_count. The rewind flags then land on the
kept tail's originals and on compress()'s own carried rows, which were
left compacted=1 and shown twice in the resumed display history. The
copies are stamped as persisted and compress_now() returns the stored list
instead of rejoining the tail a second time. Rotation, no-op and
rolled-back commits are unchanged: the copies stay unstamped and the
caller's tail is rejoined as before.
2026-09-23 01:11:58 -07:00
ehz0ah
d57c2a3254 fix(memory): forward committed entry identity to providers 2026-09-23 01:06:32 -07:00
teknium1
38c289c014 fix(models): drop gpt-6-terra, a tier OpenAI never published
#119410 registered gpt-6-terra alongside sol and luna from the request text, and the Codex
forward-compat synthesis put it (and -900k) in the live /model picker on every surface.
Nothing serves it: not the Codex account catalog (astra, sol, luna), not OpenRouter (same
three, plus -pro), and OpenAI's model page 404s. Forward-compat is for published tiers the
account catalog has not listed yet, not for guessed names. Removed from the Codex fallback
list and template chain, the context/900k tables, the effort ladder prefixes, the aux-client
family list, and the Nous/OpenRouter static catalogs; catalog JSON regenerated. A contract
test pins the synthesized GPT-6 set to the published tiers.
2026-09-23 00:44:33 -07:00
Brian Fernstrom
91df54184d fix(state): preserve summarized tool rows in micro-compaction
Micro-compaction carries a non-contiguous prefix and suffix around its summary marker. Using tail_count=len(result)-1 incorrectly marked summarized assistant/tool rows as rewind-only active=0, compacted=0.

Pass the exact unchanged carried messages instead and resolve their durable originals transactionally by row id or unique identity+timestamp, preserving summarized rows as compacted history.

Fixes #118481
2026-09-23 00:29:30 -07:00
alt-glitch
31e59b9441 feat: setup agent can search the catalog and install plugins and skills through the approval card
The setup profile's `setup` toolset was empty. It now carries one tool, manage_catalog:

- search: catalog plugins (the Plugins tab's live catalog resolver) and hub skills, with
  whether each is already installed in the default profile. Read-only.
- install: opens the same connection operation manage_connections opens, with rows of kind
  plugin / skill. Nothing installs until the user approves a row. An approved row installs
  into `default` (or the profile the Advanced modal named) through dashboard_install_plugin /
  the hub's headless install, so the catalog pin, kill list, security scan and live
  activation (#119644) are the host's. The row settles with the live MCP tool names and the
  plugin's skill.

The model sends catalog ids and an action only; every other key is refused before anything
runs. An unknown id or a plugin this OS cannot run is drawn failed with the installer's own
text. Anywhere a catalog card cannot be drawn (TUI, CLI, messaging, registry dispatch) the
result is the `hermes plugins install` / `hermes skills install` pointer.

- contract: plugin/skill targets follow the MCP transitions.
- run.apply_answer / reissue route a card answer to the module that owns the operation.
- tool_search: `setup` joins the direct-surface toolsets, so the guide's one tool is never
  deferred behind tool_search.
- docs: tools reference, toolsets reference, plugin catalog page.

Linear NS-964.
2026-09-23 08:04:33 +05:30
Siddharth Balyan
3e00a356a4 fix(mcp): one reader for mcp_servers.<name>.enabled (#119567)
The `enabled` key had four parsers. The MCP client (`_parse_boolish`) read
`enabled: 0` as on; the toolset resolver and editor (`_parse_enabled_flag`)
read it as off. The server list (`summarize_server`, `/api/mcp/servers`) read
any non-`False` value as on, so `enabled: "false"` showed on while the agent
skipped it. The catalog and `hermes mcp list` accepted only true/1/yes, so
`enabled: on` showed off while the server ran.

`tools/mcp_tool_common.py::mcp_server_enabled` is now the only reader, and
every surface calls it. `_parse_boolish` treats YAML numbers by truthiness
(0 off, other numbers on). Everything else keeps the client's semantics:
the off words are off, absent / null / junk stay on, with the existing
warning for junk.

The desktop MCP page mirrors the rule in `serverEnabled`
(`apps/desktop/src/lib/mcp-servers.ts`). One case table
(`mcp-enabled-cases.json`) drives the Python invariant test and the vitest
test, so the page and the runtime cannot drift apart again.
2026-09-22 22:00:24 +00:00
teknium1
1d10cef836 fix(codex): ask the models catalog as the newest client so GPT-6 Sol/Luna show up (#119412)
The ChatGPT Codex models endpoint gates entries on client_version against each
model's minimal_client_version. Hermes sent the "0.0.0" sentinel, which used to
return the whole account catalog; since the GPT-6 Sol/Luna rollout it returns a
frozen legacy list (astra + the 5.6 trio) while any version at or above the newest
minimal_client_version (0.155.0, 1.0.0, 99.0.0 alike, live 2026-09-22) returns
everything the account is entitled to. So the picker and the context-window probe
hid gpt-6-sol / gpt-6-luna that the vendor's own client showed for the same account.

Both request sites now go through fetch_codex_catalog_entries(): try the
newest-client URL first and fall back to the "0.0.0" sentinel only when that
answer is non-200 or empty (the backend used to reject out-of-sequence versions
that way). No local ~/.codex/models_cache.json is consulted: a missing or stale cache (this
host records 0.147.0, which hides gpt-6-sol at 0.155.0) must not decide what the
account can see.

Report by @Stone441 (#119412); supersedes #119420 by @KoNit-K, whose fix read the
compatibility version from the local `~/.codex` cache.
2026-09-22 13:07:45 -07:00
kshitijk4poor
71a2fe399b refactor(compression): startup banner uses _effective_threshold_cap; trim comments and tests
`_cap_binds` was the last inline copy of the cap-clamped-to-window predicate that #119406
centralised; the banner still prints the raw configured cap. Plugin engines without the
helper report "not binding", as before when no cap was set. The threshold_tokens comment
and the new test module stop narrating the incident; the parametrize takes the whole agent
config so the `{}` failure path is spelled out rather than derived from a ternary.
2026-09-23 00:25:29 +05:30
teknium1
79ec1f2a34 feat(models): GPT-6 Sol/Terra/Luna replace the 5.6 tiers in Nous/OpenRouter catalogs, with Codex -900k variants
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.

Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.

/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
2026-09-22 11:50:39 -07:00
kshitijk4poor
94e8a255df refactor(compression): one effective-cap helper; drop dead _config_threshold_percent fallback
`min(threshold_tokens_cap, context_length)` behind the same `cap is not None and cap > 0`
guard was written in both `_derive_trigger` and `_apply_threshold_tokens_cap`, contradicting
`_derive_trigger`'s "one place the trigger math lives". Both now call
`_effective_threshold_cap(context_length)`. The `getattr(self, "_config_threshold_percent",
...)` fallback was dead: `__init__` sets the attribute unconditionally and no bare-`__new__`
test instance reaches `_derive_trigger`. Behaviour-preserving; the cap tests still fail when
the helper is mutated to return None.
2026-09-23 00:10:27 +05:30
kshitijk4poor
ed9e227257 fix(agent): threshold_tokens inline default matches DEFAULT_CONFIG like every other compression key
`_parse_compression_config` documents "Defaults here MUST match DEFAULT_CONFIG" because the
config-load failure path hands it `{}`. Every key had an inline fallback except the cap
added in #115986, so a broken config.yaml kept `threshold=0.50` but silently dropped the
256K cap — the pre-#115986 500K trigger on 1M-window models. An explicit null still
means ratio-only.
2026-09-23 00:10:27 +05:30
Siddharth Balyan
4094ab610d hermes_platform.resolver: locate/inspect/probe tiers, AppResolver, gh lookup migrated (NS-921) (#118065)
* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates

Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.

* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources

Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.

* refactor(copilot): gh candidates through locate_command and the Homebrew table

First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.

* feat(platform): availability() over an application declaration

locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.

* feat(platform): application declarations parsed into AppDef per OS

The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.

* feat(mcp): check_fn honours a registered application declaration

_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).

* feat(skills): requires_apps gate through registered declarations

Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.

* docs: application declarations page

The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.

* test(platform): declaration parser, availability, gates

The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
2026-09-22 17:08:30 +05:30
carlotestor
74183e1793 perf(stream): resolve plugins.stream_reasoning_deltas once per stream, not per token
_fire_reasoning_delta called stream_reasoning_deltas_enabled() on every reasoning
token; each call took _CONFIG_LOCK and paid load_config()'s full deepcopy on a
cache hit, so every concurrently streaming thread in a WebUI/gateway process
serialized behind one lock on the token path.

Cache the opt-in on the agent for the life of one stream (reset in
_reset_stream_delivery_tracking, so the next request re-reads config) and use
load_config_readonly() for the lookup itself.

(cherry picked from commit 20fab007e5c3e399e3923bb6d66649918292825c)
2026-09-22 16:15:34 +05:30
kshitijk4poor
28bd8cc08c fix(memory): never dedupe indented recall lines
Only column-0 bullets participate in `_drop_repeated_recall_lines`.
An indented line is a continuation of the bullet above it (nested
child, provenance, wrapped prose) and is kept verbatim, so identical
children under two different parents both survive. The docstring
already promised this; the code now matches.

PROOF: with the old code, `- Project A\n  - status: active\n- Project B\n
  - status: active\n` lost B's child (probe s4_eff_nested.py) and the
new assertion in test_a_bullet_with_continuation_lines_is_never_touched
fails (AssertionError at :49); after the change the probe returns the
input unchanged and tests/agent/test_memory_context_dedupe.py passes.
2026-09-22 15:55:09 +05:30
kshitijk4poor
8f618d8409 refactor(memory): match the recall bullet regex once per line
`_RECALL_BULLET_RE.match(stripped)` was evaluated twice per line and the inline
comment restated the docstring's per-section scoping paragraph. Compute
`is_bullet` once and keep a one-line comment. Behaviour-neutral.

PROOF: mutation `is_bullet = True` → tests/agent/test_memory_context_dedupe.py
1 failed (repeated-bullet), restored → all 44 tests across the PR's four test
files pass; ruff clean; footguns clean; real import OK.
2026-09-22 15:55:09 +05:30
kshitijk4poor
af0305c95a fix(memory): any column-0 non-bullet line opens a new dedupe section
The per-section reset only fired on `#…`, `**…` and `---` lines. The
in-tree RetainDB provider writes prose headings (`Profile:`,
`Relevant memories:`, `Instructions:`), so its second section still
collapsed into the first: `Profile:\n- None\nRelevant memories:\n- None`
lost the second `- None` and left a heading claiming nothing — the
exact failure 931d240e14 set out to fix, for markdown headings only.

Now any non-empty, non-bullet line at column 0 (a heading of any style,
a `---`/`***`/`___` rule, prose) clears `seen`; bullets and indented
continuation lines stay inside the current section.

PROOF: tests/agent/test_memory_context_dedupe.py::
test_identical_placeholder_bullets_under_different_headings_both_survive
(RetainDB prose-heading shape + markdown shape). Red with
agent/memory_manager.py at the previous commit (prose case drops the
second `- None`), green after; the continuation-line carve-out test
stays green.
2026-09-22 15:55:09 +05:30
kshitijk4poor
e1ea7e74f0 fix(memory): repeated-bullet dedupe is scoped per section
The `seen` set from #117399 was block-global, so identical placeholder
bullets under different headings ("- (none recorded)" under two stores)
collapsed into the first section and left the second heading claiming
nothing. Reset the set at every heading (`## …` / `**…**`, the PR's own
heading detection) and at `---`, so only a repeat inside the same
section is dropped.
2026-09-22 15:55:09 +05:30
John Paul Soliva
83a1a9a2bb fix(memory): only a self-contained bullet is deduped, and a bold heading is not a bullet
Review follow-up on two real defects in the first cut, both reproduced before fixing.

A dropped duplicate left its continuation lines behind, and they re-parented under the surviving
bullet: `- prefers draft PRs` / `  (logged 12 Jan, supermemory)` followed by the same headline
logged elsewhere collapsed into ONE entry carrying BOTH provenance lines — inventing a record
neither provider reported, then stamping it into `api_content` to be replayed every turn. That is
worse than the duplicate the change set out to remove. A bullet that carries continuation lines is
now never dropped and never suppresses a later one, so two entries that share a headline and differ
underneath it both survive.

And `stripped[:1] in ("-", "*")` treated `**Preferences**` as a list item, so a repeated bold
section heading was silently deleted — contradicting the docstring's promise that headings survive.
The marker test now requires a real bullet: marker, whitespace, content (`[-*+]\s+\S`), which also
keeps `*emphasis*` out. Numbered items stay out of scope, and the docstring says so rather than
claiming "list items" generally.

Measured again on the same 166 real blocks: still 219,260 of 1,302,220 body bytes — 16.8%,
unchanged. Every duplicate in that sample is a self-contained bullet, so correctness cost nothing.

Tests: a bullet with continuation lines is untouched; a nested bullet is a child, not a repeat;
same-depth indented duplicates still collapse; bold headings, emphasis and numbered items are left
as written. Removing the continuation guard fails two. The tautological assertion the review flagged
is replaced by an exact whole-body comparison.

(cherry picked from commit 37a36f23658b8ca1e112f46ff60a705a66f80a5b)
2026-09-22 15:55:09 +05:30
John Paul Soliva
cbe23de61a perf(memory): a recalled line is stated once per memory-context block
Every user turn injects a `<memory-context>` block built from the providers' prefetch, and nothing
dedupes it: a provider merges several stores and `prefetch_all` merges several providers, so one
prefetch routinely surfaces the same fact two or three times.

It is not paid once. The composed block is stamped into the user row's `api_content` sidecar and
replayed verbatim on every later request for as long as that row is in context — deliberately, so
the prompt-cache prefix stays byte-stable. Compaction cannot reclaim it either: `_demote_tool_result_at`
bails on anything that is not `role == "tool"`, so a user row is never shrunk and drags its memory
dump along while counting against the tail budget. A duplicated bullet therefore costs its tokens
once per turn, for the life of the row, while telling the model nothing that same block has not
already said.

Measured on a real operator state.db (166 rows carrying a block, 46 sessions): 219,260 of 1,302,220
body bytes — 16.8%, ~54,815 tokens on first send alone — are list items byte-identical to an earlier
line of the SAME block. Median block ~2,034 tokens per user turn.

Only list items are considered, and only when identical after stripping: headings, prose, blank
lines and `---` rules are left exactly as the provider wrote them, so sections still read as
written and a line repeated deliberately as prose is untouched. Within one block only — deduping
against earlier turns would save far more (73% of bullet bytes in that sample were already sent
earlier in the session) but changes what the model sees on a turn, so it is not done here.

The "provider returned pre-wrapped context" warning stays keyed on sanitization alone: a deduped
bullet is routine, not a provider fault.

Fixes #117397

(cherry picked from commit e3b08ca9b1708a7cddc2f1a2191b51d60d194624)
2026-09-22 15:55:09 +05:30
Siddharth Balyan
70f5dc5f46 feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours

The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.

- `tools/connectors/portal/`: a client for the portal's tool-list route and a
  JSON cache under the Hermes home, one file per portal origin and connector.
  An entry is fresh for 24 hours. After that the read revalidates with the
  stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
  serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
  chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
  `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
  live in `tui_gateway/methods_connectors_account.py`.

The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.

* feat(connectors): catalog, accounts and member tool rules by RPC

The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.

- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
  the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
  first. The body is a union on `mode`, so a reader can name who turned a
  tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
  connector, or one connector on or off), with the revision the user saw. A
  stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
  write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
  page can show one card per app.

* feat(connectors): connect an app without a chat session

Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.

- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
  `connectors.operation.wake` and `connection.respond` take `owner`, a union
  on `type`: `session` (today's behaviour and authorization) or `account`
  (routed by `profile`, authorized by the live transport like `mcp.*`).
  `session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
  thread, under the profile's scope, so the watcher reads the account and
  settles the operation. A second connect for an app that is already
  connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
  address, so its updates go out on the session-less broadcast path.

* feat(mcp-catalog): eighteen more bundled entries name their hosted connector

A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.

* refactor(connectors): the account handlers share one gate, one params model and one write table

The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.

The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.

The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.

An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.

Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.

* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"

The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.

Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.

* fix(connectors): a connect from the page returns to the app after sign-in

The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.

Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.

* test(connectors): defer the new connector RPC coverage

The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.

Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.

Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.

* fix(cli): the connection panel hands the tool thread back at once

The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.

For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.

The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.

Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.

* fix(connectors): "run it again" lives in the library, so the classic CLI can use it

Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.

`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.

Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.

* feat(connectors): the account list and disconnect go through the portal

`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.

The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.

`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.

* fix(connectors): the account RPCs answer what the portal really sends

Checked against the portal source and against the staging and production
services.

- Errors are read from the upstream error code, not the HTTP status. A rule
  write answered 409 for a stale revision and for a user with no organisation;
  both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
  403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
  sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
  the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
  portal's own result for this user, with its stamp and without provider or
  subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
  and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
  answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
  invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
  so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.

Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.

* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link

Found by two adversarial reviews of the RPC layer and its types.

- `connectors.connect` from a chat session with no open operation is refused
  (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
  registry with no card: it made a link nobody watched, returned a reply
  without the required `settled` field, and named an operation that was never
  registered. There is one way into an operation: the agent's call, or the
  account owner's `connectors.connect`. "Run it again" inside an open
  operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
  profiles with the same key no longer cross-deliver a sign-in link. The event
  payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
  OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
  `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
  The last one is display data: the gateway enforces the rules, the backend
  only passes the list on. The phantom `name` and `description` are gone, and
  the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
  produces it. The contract generator now fails when a contract enum and its
  domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
  as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
  instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
  on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
  `optional-mcps/` are untouched by this PR again.

anti-slop: no net-new findings (15 touched files).

* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect

The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.

- The agent turn now declares how a link can reach the user
  (`tools/connectors/turn.py`): CARD when the agent was built with a
  connection callback, SIDE for a subagent or a background turn, LINK for a
  headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
  tool batch in the agent loop and read by the connector dispatch path, which
  never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
  link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
  CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
  tool on any path that derives the tool list, and a connector call on an
  unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
  does, so no `connection.update` is emitted for an operation no client asked
  for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
  `respondToConnectionRequest` share one guard and send nothing for a settled
  or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
  repeated it to the user.

Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.

* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients

A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.

- `tools/tool_labels.py` is the one place that turns a bridged call into a
  label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
  → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
  keeps its own emoji, verb and primary-argument preview. A batch gets exactly
  one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
  failure text on the row of the call that failed. With friendly labels off
  it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
  carry a typed `labels` field. It does not depend on the classic CLI's
  display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
  labels, one row per call. The labels reach the row under a key no tool
  argument can use. The connect card it drew under a failed tool result is
  gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
  `manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
  "Reading tool details · N tools".

Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.

* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly

- A failed hosted search or describe used to return nothing, by design, so the
  model saw only local tools and told the user that a connected app was
  missing. The local results are unchanged; when the hosted leg failed, the
  `tool_search` and `tool_describe` results carry
  `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
  and one hint line. A rejected token is `sign_in_expired`; an entitlement
  refusal or a shut gate adds nothing. `tool_describe` no longer lists those
  names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
  is a hosted connector account; `mcp: true` only when the user asks for an MCP
  server, a local server or an install, or when the name exists only in the
  catalog; connect and reconnect are hosted verbs, install, enable and
  authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
  gateway does not know the connector (confirmed on that failure path) and the
  name is a catalog entry does the target fail with "X is a local MCP server.
  Call manage_connections with action install ...". It is a per-target
  outcome: other targets of the same call keep their links and their card. A
  vendor failure on a name both sides know stays an ordinary failed row. The
  MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
  USER asks for that app again; the description and the settled-result notes
  say so. A builder saw the model refuse a direct user request.

Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.

* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled

Reproduced on the real Ink TUI with the rig, then fixed:

- The keyboard was dead during the sign-in wait: the card kept a `submitting`
  flag that the normal OAuth path never cleared, and Esc went through the same
  guard. The in-flight state now belongs to the answered row and clears when
  that row moves, when any later frame of the operation arrives, or after
  five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
  (the input handler had no branch for this overlay); Shift+arrows scroll the
  transcript and the card ignores them; arrow keys no longer move the text
  cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
  operation stayed in the store, and a resume dropped the pending card. The
  flag survives idle, a resume shows the pending card again, a session switch
  clears it.
- States with no branch: `not_connected` and a row with no link fell into the
  credential form; `expired` vanished with no note. The title and the row text
  now name the action (connect, reconnect, install, enable, authorize); a
  failed or expired row with no fields offers Try again / Skip; a failed row
  WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
  per app states the outcome. A settled or dismissed operation id is
  remembered, so no replay or resume can reopen its card. Esc in the last
  "Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
  the card in one sentence.

Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.

* chore(connectors): remove the comments and docstrings this branch added

Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.

Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.

* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures

Found by the end-to-end runs on the pushed head.

- Desktop: after a window reload, Continue on the restored card sent nothing.
  The answer looked up the backend that holds the session with the runtime
  session id, the lookup wants the stored id, and a failed lookup returned
  silently. When the lookup gives no owner the answer now goes out on the
  window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
  a refused scope, a non-member and a missing organisation alike: the handler
  runs with the gateway's globals and did not import the reason enum, so its
  own error mapping raised. `connectors.accounts.remove` caught auth failures
  in its generic branch. `org_required` was mapped on `policy.set` only. All
  six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
  `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
2026-09-22 13:57:51 +05:30
teknium1
966d091d6a fix(aux-hooks): tolerate test seams that stub relay metadata without task keys
test_chat_sdk_transform_bypass stubs _relay_auxiliary_metadata with an empty
metadata dict; read aux_task/api_mode with .get so the seam keeps working.
2026-09-22 01:19:12 -07:00
teknium1
0e5809566f feat(plugins): fire pre/post_auxiliary_call events on every auxiliary LLM call (#79733)
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.

- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
  post_api_request payload shape plus `aux_task`, `api_request_id`
  (`aux-...`, shared by every attempt of one logical call), `retry_count`,
  `streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
  turn is in flight; fail-open (a raising/hung subscriber is logged and
  the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
  attempt shares (_relay_sync_completion / _relay_async_completion /
  _relay_sync_stream) run under the hook pair — retries and fallbacks
  included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
  sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
  tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
  aux_task and no api_request events; raising subscriber never breaks
  the call). First is red on origin/main.

Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.

Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-22 01:19:12 -07:00
kshitijk4poor
e4f76b8277 refactor(agent): inline the single-use _finish closure in _sample_summary_records
`_finish` (agent/context_compressor.py:3586-3588) only wrapped
`_coverage` and re-derived `shown` from the final `selected`; it was
called once, at the return. With `_merged` hoisted above the slice loop
the classmethod now has two closures (`_merged`, `_render`) plus
`_coverage` instead of four. Same expressions, same evaluation order
(`_render` first, then the coverage counters), so the output and
coverage dict are byte-identical — verified on the 47-shape capture and
probe_C2 (0 violations, base identity x4 True).
2026-09-22 13:40:13 +05:30
kshitijk4poor
8e12cd92d7 refactor(agent): integer head/tail split in _bound_oversized_record
`_bound_oversized_record` (agent/context_compressor.py:3485-3488) split
`remaining` with float arithmetic (`int(remaining * 0.5)`) and guarded
the tail slice with `if tail_len else ""`. `remaining = limit -
marker_reserve` is >= 1 because the line above already returned when
`limit <= marker_reserve`, so `head_len = remaining // 2 <= remaining - 1`
and `tail_len = remaining - head_len >= 1` always: the guard can never
fire. Use `remaining // 2` and drop the dead guard.

Byte-identical: `int(r * 0.5) == r // 2` holds for every non-negative
int (checked exhaustively for r in 0..10**6), and the guard was dead.
Capture of 800 bounded-record cases (4 record shapes x remaining 1..200)
plus the 47-shape sampling capture is identical before/after.
2026-09-22 13:40:13 +05:30
kshitijk4poor
e17d49174d refactor(agent): drop unreachable trailing-gap branch from lean-sampling _render
The `if cursor < len(records)` block at the end of `_render`
(agent/context_compressor.py:3580-3583) duplicated the gap-marker
arithmetic of the in-loop branch but had no path to it.

Invariant: `_render` is only reached with non-empty `records`
(empty returns early), and n >= 1, so the slice-build loop always runs
its last iteration, which anchors `end = len(records)` and
`start = end - 1 >= 0`, hence `end > start` and the final slice ends at
`len(records)`. `_merged` keeps `max(end)` for the last interval, and the
extension pass only grows the newest slice backward (`(s - 1, e)`) or
older slices forward but never past the next slice's start, so the last
slice's end stays `len(records)`. Therefore `cursor == len(records)`
after the loop for every input.

Byte-identical across the 47-shape capture (40 probe_C2 shapes + 7 edge
shapes incl. single-record and empty); probe_C2 -> 0 violations;
under-cap base identity x4 True.
2026-09-22 13:40:13 +05:30
kshitijk4poor
56b63ecf82 refactor(agent): merge lean-sampling slices once via _merged instead of inline
The slice-build loop in _sample_summary_records hand-rolled the same
"overlapping/touching slice -> extend previous" fold that the `_merged`
closure 25 lines below implements (agent/context_compressor.py:3563-3568
vs :3590-3597). Hoist `_merged` above the loop, append raw (start, end)
pairs, and merge once before the extension pass; the first `_render`
no longer wraps `selected` in a redundant `_merged` since it is already
merged. One interval-merge implementation instead of two.

Output is byte-identical: the merge is applied once to the same ordered
slice list, so the extension pass starts from the same `selected`.
Verified with a 47-shape capture (40 probe_C2 shapes + 7 edge shapes)
diffed before/after: identical sampled text and coverage; probe_C2
40 shapes -> 0 violations; under-cap base identity x4 True.
2026-09-22 13:40:13 +05:30
kshitijk4poor
f3642a27d8 docs(agent): say why the extension pass re-renders after the cap pre-check
The comment claimed "marker widths shrink as records leave a gap" — but
shrinkage can only make the render smaller, which the pre-check already
tolerates. The exact re-render exists for the opposite case: a gap's
first index can gain a digit or thousands separator (999 -> 1,000, +2
chars) when the added record moves a marker, and that growth can push a
render sitting at cap over it. Comment only; no behaviour change.

Finding: $D/regate/C.md suggestion 1 (agent/context_compressor.py:3619).
2026-09-22 13:40:13 +05:30
kshitijk4poor
be63d5bcd3 fix(agent): hand out lean-sampling headroom round-robin across slices
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.

Before -> after (records per slice, fill unchanged):
  10000x100  [1,1,1,1,1,1,1,8]      -> [1,2,2,2,2,2,2,2]      fill 0.944
   8000x100  [2,2,2,2,2,2,2,5]      -> [2,2,2,2,2,3,3,3]      fill 0.957
   4000x300  [4,4,4,4,4,4,4,11]     -> [4,5,5,5,5,5,5,5]      fill 0.987
   2000x600  [9,9,9,9,9,9,9,15]     -> [9,9,10,10,10,10,10,10] fill 0.995
    500x2000 [37,...,37,39]         -> [37,...,37,38,38]      fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.

Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).

Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
2026-09-22 13:40:13 +05:30
kshitijk4poor
c94fe7a259 docs(agent): say sampled_chars counts display chars in _sample_summary_records
The coverage docstring (agent/context_compressor.py:3510-3511) said the
counters count "record content only", but `sampled_chars` sums the
*display* records (post `_bound_oversized_record` truncation) while
`input_chars` sums the raw records. Spell that out so telemetry
consumers do not compute `omitted` two different ways (gate 2c
suggestion, L3512-3513/3524-3525).
2026-09-22 13:40:13 +05:30
kshitijk4poor
eb05bde6df refactor(agent): pre-declare summary_input_* telemetry keys
`_begin_compression_telemetry` (agent/context_compressor.py:1947-1960)
did not seed the six `summary_input_*` counters that
`_record_summary_input_coverage` adds on the lean path, so the emitted
`compression_attempt` line (conversation_compression.py:1451-1470) had a
different key set for lean and legacy attempts. Seed them as None so the
schema is stable; probe: lean and legacy attempts now emit identical key
sets (gate 2c suggestion, L1947-1960).
2026-09-22 13:40:13 +05:30
kshitijk4poor
949d1707f9 fix(agent): spend leftover lean-sampling budget on whole neighbouring records
The whole-record greedy fill in `_sample_summary_records`
(agent/context_compressor.py:3541,3555-3563) stops a slice as soon as the
next record would cross `target` (~20K), so each slice can be short by up
to one record. With mid-sized records that wastes a large share of the
160K cap: probe worst case (10K records) filled 90,660/160,000 (57%),
8K records 129,021 (81%); main filled 159,988 by splitting records.

Add a bounded extension pass after `selected` is built: newest slice
grows backward, older slices grow forward, one whole record at a time,
only while the exact rendered length stays <= cap and the slice does not
run into its neighbour (adjacent slices are merged so no separator is
lost). Records are never split and the cap is never exceeded.

After: worst random shape 150,581 (94%, remaining headroom < one 10K
record), 8K records 153,117 (96%), 1551-record region 159,996 with
968 whole records, 0 partial, 8 buckets within +-10%, telemetry
input_chars == sampled_chars + omitted_chars (probe: $D/fold/probe_C_fold.py).
The existing keeps_record_boundaries fixture (1.2K records) fills 99.1%
without this pass, so no fill assert was added there — it would be toothless.
2026-09-22 13:40:13 +05:30
kshitijk4poor
f94e296ae6 refactor(agent): remove unreachable overflow-trim loop in _sample_summary_records
Delete the post-render overflow-trim loop (agent/context_compressor.py
:3591-3607) and state the invariant in its place.

Why it is unreachable: every display record is pre-bounded to `target`
via _bound_oversized_record; each slice's greedy fill stops before
`size + sep + next > target`, so a slice renders to <= target chars and
the n slices to <= n*target (a merge of overlapping/adjacent slices is
their union: overlap only removes chars, adjacency adds one 2-char
separator but drops one marker of >= 2 chars). There are at most n-1
markers, each <= marker_len because marker_len is formatted with
first=last=len(records) and elided=total_len, the maxima. Since
target = (cap - (n-1)*marker_len) // n, the rendered length is
<= n*target + (n-1)*marker_len <= cap. Gate mutation M5 (loop disabled)
left all 7 sampler tests green; 30 random shapes never entered it.
2026-09-22 13:40:13 +05:30
kshitijk4poor
5aeef2889f refactor(agent): drop unreachable re-trim in _bound_oversized_record
Delete the `if len(res) > limit` re-trim at the end of
`_bound_oversized_record` (agent/context_compressor.py:3488-3489).

Why it can never fire: `marker_reserve` is the marker width formatted
with `elided=len(record)`, the largest value `elided` can take, so the
real marker is <= marker_reserve; `head` is `record[:head_len].rstrip`
(<= head_len) and `tail` is `record[-tail_len:].lstrip` (<= tail_len);
head_len + tail_len == limit - marker_reserve. Hence
len(head + marker + tail) <= limit always. The branch was untested and
unreachable (gate 2c W2).
2026-09-22 13:40:13 +05:30
kshitijk4poor
33a1d46dfc feat(agent): record summary-input coverage telemetry (from #118390)
Lean sampling is the documented budget mechanism, but nothing reported
how much of a region actually reached the summarizer, which is the gap
the reporter narrowed #118362 to. _sample_summary_records now returns
record-level coverage counters alongside the bounded transcript and
_generate_summary writes them into the active compression telemetry
(summary_input_chars / sampled_chars / omitted_chars / record_count /
sampled_record_count / elided_record_count), so the existing
compression_attempt JSON log line carries them with no transcript
content. Char counters cover record content only, not separators or
markers. Elision markers also name the one-based record range they
stand for, so a reader can tell how many turns a gap hides, not only
how many characters.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-22 13:40:13 +05:30
kshitijk4poor
a3a03d48d4 refactor(agent): sampler takes serialized records only
_sample_summary_input accepted either a flat string or a record
sequence and, for strings, rebuilt records with content.split("\n\n").
The PR's own third commit established that "\n\n" is not a record
delimiter (message bodies contain blank lines), so that fallback would
re-introduce split records for any caller that reached it. The only
production caller already passes records to _sample_summary_records;
drop the dual-signature wrapper and move the test callers to records.
2026-09-22 13:40:13 +05:30
kshitijk4poor
e28157be98 fix(agent): keep _serialize_for_summary byte-identical to main
The record split added .rstrip("\n") to every serialized record. That
changes the bytes _serialize_for_summary returns for legacy tail_mode
and for agent/micro_compaction.py's _micro_summarize_one, which never
sample and never needed it: a message with trailing newlines was no
longer reproduced verbatim in the summarizer prompt. Sampling receives
records structurally, so an empty pseudo-record can no longer become the
tail anchor and the strip has no remaining purpose.
2026-09-22 13:40:13 +05:30
joaomarcos
04fe735c57 fix: preserve structural record framing in lean summary sampling
Preserve record framing structurally so messages containing internal blank lines are not split into pseudo-records.

- agent/context_compressor.py:
  * Extract _serialize_records_for_summary() to retain turn boundaries as a sequence of records.
  * Add _sample_summary_records() to sample across serialized turn records directly without lossy double-newline splitting.
  * Correct separator accounting in omission markers at gap boundaries so separator counts match full serialized text.
  * Update _sample_summary_input() to support structural records while preserving flat-string fallback with trailing fragment pruning.
  * Wire _generate_summary() to use _serialize_records_for_summary() and _sample_summary_records() in lean mode.
- tests/agent/test_context_compressor.py:
  * Test multi-paragraph serialized messages maintain role boundaries and keep newest message paragraphs together.
  * Test newest message ending in blank lines does not produce empty tail anchor.
  * Test exact omission marker separator accounting.
  * Test end-to-end _generate_summary prompt assembly with multi-paragraph messages.

(cherry picked from commit f68f1bfcf037dc5ddecf6cc4032023cd88c496c8)
2026-09-22 13:40:13 +05:30