Builds on tancou's #119129 (cherry-picked above): the pin now lives in
get_routing_process_hermes_home() and only the four routed-profile DECISIONS read it.
get_process_hermes_home()/get_hermes_home() keep following HERMES_HOME, so an env-only
home switch in a multiplexed process resolves as before.
- set_multiplex_active(True) pins the launch home only when no host pin exists, and
set_multiplex_active(False) releases only the pin it created itself. A transient toggle
(gateway_migrate._multiplex_read_mode, cron external-worker restore) no longer drops an
embedding host's explicit pin_process_hermes_home(launch).
- profiles._cleanup_gateway_service binds set_hermes_home_override(profile_dir) beside the
env write. Under the previous head, DELETE /api/profiles/<x> from a multi-profile dashboard
resolved get_service_name() against the pinned launch home -> bare `hermes-gateway`, and
disabled/stopped/unlinked the HOST multiplexer's unit. Same path serves rename_profile.
Tests (red on the previous head): explicit pin survives True->False; env readers follow the
env while pinned; two-home delete removes hermes-gateway-victim and leaves hermes-gateway.
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.
Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.
Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.
Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.
Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.
Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`test_language_is_per_profile_under_multiplex` switched profiles by rewriting
`os.environ["HERMES_HOME"]` after `set_multiplex_active(True)`. That is the T1
standalone contract (environ IS the profile); a multiplexed turn binds
`set_hermes_home_override`, and since the launch home is pinned at activation
(#119242) a later env mutation is deliberately ignored. The two assertions are
unchanged; only the profile-switch mechanism now matches production.
Every "does this task serve a ROUTED home" decision (serves_routed_profile,
_is_process_home, _is_routed_home, env_loader._process_hermes_home) compares the
home override with get_process_hermes_home(), which read os.environ["HERMES_HOME"]
live. A host that mirrors the served profile into that env var per turn (hermes-webui)
made every served profile look like the launch one: MCP registry scope None, bare
cross-profile connection names, launch residue kept in served child envs, the launch
GATEWAY_ALLOW_ALL_USERS grant seeded into the served scope.
set_multiplex_active(True) now pins the launch home (hermes_constants.
pin_process_hermes_home; first pin wins, an embedding host may pin explicitly) and
get_process_hermes_home() returns the frozen value while multiplex is active.
Standalone hermes -p x gateway run (multiplex inactive) keeps following the env.
No os.environ fallthrough is added anywhere.
Closes#119242
archive_and_compact takes the newest tail_count durable rows as the carried tail's superseded
originals (active=0, compacted=0: hidden from display and session_search). The CLI and the gateway
persist a turn's user row only after the turn-start preflight, so at that compaction the row rides
in the carried tail with no durable original of its own. The positional rewind then reached one
row past the tail and flagged the newest summarized message as a superseded duplicate. Each
turn-start compaction hid one more (A5, then A16 in a two-compaction run), on the default config.
The TUI/Desktop are immune because they persist the user row at submit.
tail_count now leaves out this turn's rows that never reached state.db: no persisted marker and
no _row_id, counted within the carried tail window only. That includes the user row and any
unflushed scaffolding. It does so only while the turn holds the session turn lease. Between turns
(a manual /compress, the gateway's pre-turn hygiene) the turn anchor is left over from the last
turn, and a gateway transcript reload is durable but unmarked. Counting those rows would leave
carried originals at compacted=1, the duplicate recall #86366 fixed.
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.
Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
Several profile-scope binders called set_hermes_home_override and
set_secret_scope before the try/finally that releases them. A raise in
scope setup (a corrupt or removed profile home) propagated with the
foreign override still bound to the caller's context, silently re-homing
every later read — and in _reregister_orphaned_adopters it also skipped
every remaining adopter. The set calls now run inside the try with
None-guarded resets; the same shape is fixed in the routed-turn scope,
the cron external worker (which also leaked the multiplex flag), the
kanban worker scope, the MCP OAuth paths, the launch-profile policy,
and model_switch, which releases partially-bound scopes on raise.
get_secret returned os.environ on a scoped miss whenever multiplex was
off, but non-multiplex hosts serve foreign homes too (dashboard/desktop
backend, per-profile cron, MCP owner scopes, kanban spawn-env builds),
where os.environ is the launch profile's. Bound scopes now carry the
home they were built for; serves_routed_profile detects a foreign scope
even when the binder deliberately skips the HERMES_HOME override, and
the miss returns the caller's default. Every production binder stamps
its home; own-home scopes keep the deliberate env overlay.
The logical LLM scope close in ``relay_llm._complete_logical`` popped its handle with
the unguarded ``pop_relay_scope``. Two consecutive calls in one session share a physical
scope stack, so when the sibling turn's live scope sat above the handle the native
binding raised::
RuntimeError: invalid argument: scope handle is not at the top of the stack
``_complete_logical`` catches that, logs "logical LLM finalization failed" with a
traceback and returns early, so the early-return path also skips the
``turn.logical_llm_calls`` cleanup until the handle is retried. Observed in production as
one traceback per overlapping turn (~14/day on a busy local profile).
``pop_relay_scope_if_top`` already exists for exactly this case and is used by the
shared-metrics task close (PR #116685, #115471); the logical-LLM seam was missed. Guarding
``pop_relay_scope`` itself is NOT an option — ``_pop_with_drain`` relies on that raise to
detect a stacked sibling and drain it.
The skipped scope is reclaimed by the existing session-close drain
(``RelayRuntime._close_scope_handle``), so the handle still leaves
``logical_llm_calls`` and the session unwinds without an orphan.
Test: ``test_logical_close_skips_pop_under_concurrent_turn_scope`` pushes a sibling scope
in the session context, completes the logical call, and asserts the sibling (not ours) is
still on top and that the sibling's scope survives. It fails on upstream ``main`` with the
exact ``RuntimeError`` above and passes with the guard.
Fixes#115471 (the remaining call site).
The api_server platform builds a fresh AIAgent per request (per-request
callbacks, model route, ephemeral prompt), so the memory provider was
re-initialised on every request. External providers deliver recall as the
PREVIOUS turn's background prefetch held on the provider instance, so a
continued session (X-Hermes-Session-Id, previous_response_id, declared
session key) never received automatic recall, and for hindsight
local_embedded each init also restarted the embedded daemon, killing the
retain still in flight. Pre-existing: the same probe fails on main before
the hindsight catalog migration (526d135a96, bundled provider).
ApiServerMemorySessions parks the session's initialised MemoryManager
between requests (exclusive check-out/check-in, keyed by profile home +
session id, LRU/idle eviction under the owning profile's scope) and
AIAgent(memory_manager=...) adopts it instead of loading and initialising
the provider again. /v1/chat/completions, /v1/responses, session chat and
/v1/runs all go through the same two seams (_create_agent, turn finally).
Follow-up to the salvaged #88281 commits: read gateway.telemetry.session_segments
through hermes_cli.config_effective.load_user_config_effective(), the exact
primitive gateway.run._load_gateway_config() now delegates to, instead of
read_raw_config() + managed overlay. The raw read dropped ${VAR} expansion
(max_turns: ${SEG_TURNS} parsed as 0) and model-key canonicalization.
The same import also rebound TERMINAL_CWD to the home dir in `hermes -z`
(which does not pre-set it like cli.py): the launch dir's AGENTS.md was
never loaded and terminal commands ran in $HOME (#95577). The subprocess
guard now asserts TERMINAL_CWD / HERMES_QUIET / _HERMES_GATEWAY stay unset
too, and the fixtures patch the effective reader, which both the old and
new call paths go through.
Found by the core entrypoint parity E2E matrix (oneshot row, context_file).
Follow-up to salvaged PR #87187:
1. The comment in _segments_config() claimed gateway.run's module top
level sets HERMES_EXEC_ASK, but #86043 moved it to start_gateway().
Only _HERMES_GATEWAY and HERMES_QUIET remain at module level. Updated
the comment to reflect this, noting #86043 for historical context.
2. Added timeout=30 to the subprocess.run in _run_isolated test helper
to prevent CI hangs if the subprocess deadlocks.
Address AI-review feedback on #87187:
- Drop raising=False from the read_raw_config monkeypatches so a typo in
the target path fails loudly instead of silently no-op'ing (which would
read the real user config during tests).
- Add a subprocess regression guard proving relay_runtime._segments_config()
never drags gateway.run into sys.modules. Its module top level sets
_HERMES_GATEWAY / HERMES_QUIET / HERMES_EXEC_ASK, flipping a non-gateway
host onto the gateway approval path and hanging dangerous commands in
pending_approval (#87183). A monkeypatch of read_raw_config can't catch a
future refactor that reintroduces the import; this runs the real path in a
fresh process.
Leftover monkeypatch on gateway.run._load_gateway_config would still
import gateway.run during the test run, setting _HERMES_GATEWAY /
HERMES_QUIET in the test process. Mirror the _set_segments fix.
gateway.run executes os.environ["_HERMES_GATEWAY"] / HERMES_QUIET /
HERMES_EXEC_ASK at module top level. _segments_config() late-imported it,
so any CLI/TUI/desktop/cron process hitting the relay turn lifecycle had
HERMES_EXEC_ASK=1 written into its own environ, routing dangerous-command
approvals onto the gateway path (tools/approval.py:3344) where the CLI has
no notify_cb — the approval panel never renders and commands hang forever
in pending_approval.
Read the gateway telemetry config via read_raw_config() + managed overlay
instead (no import side effects), and point the test fixture at the new
read path.
Drives the real AIAgent loop: a tool round, then the pre-API gate compacts two
historical rows away before the next request. That request must still end with this
turn's user row, the assistant tool_call and its result. Red on origin/main and with
only the post-tool re-anchor; green with the prepare_iteration check. Folds the
post-tool re-anchor invariant into the same file (renamed to cover both gates).
Passes the now-required current_turn_user_idx to the existing post-tool prune-wiring
test call (signature change from the salvaged commit; no assertion changed).
The unexplained-rejection gate from #114644 ended the turn with "another request on the same
server was probably holding its capacity ... wait and /retry" whenever a server said "context
exceeded" without a count while the local estimate sat under half the known window. That cause
only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any
public endpoint) the same rejection means the route's real window is smaller than the one Hermes
assumes, so /retry failed identically every turn and the conversation was never compressed:
Discord bots on claude-opus-5-5 were stuck repeating the message.
The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container
DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ
entry says which endpoints get the message.
Restore the cheap repo-wide guards the per-file triage classed as source reads
but that protect recurring bug classes (<2s total):
- subprocess env scrubbing near spawn sites (credential leakage)
- gateway UTF-8 encoding= on file I/O (Windows mojibake)
- no raw yaml.safe_load of config.yaml (lost ${ENV} expansion)
- CLI subprocess.run timeouts (hung CLI)
- no locked readers on the shared state.db connection (#99349 segfault)
- CI classifier outputs / live-comment watch list match real workflows
- relay imports no platform crypto (relay trust boundary)
- Desktop relay deliver budget mirrors the Python deadlines (#93911)
- no native title= on Desktop buttons (DESIGN.md rule)
Drop _BASELINE entries in check_os_marker_fakes.py for files that no longer
fake macOS (the checker fails on stale entries), and remove doc/comment
pointers to deleted tests.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Deleted: the seven files whose subject was plugins/memory/hindsight itself
(provider, config schema, env perms, local-runtime hint, templates, health
grace timeout, root guard) and the hindsight-only cases inside the multiplex
identity-scope, provider-thread, BOM-tolerance and session-switch suites.
Retargeted: the dashboard memory-provider config-surface tests used hindsight
as the only bundled provider with a flat `<home>/<name>/config.json` declared
schema; they now install a synthetic user plugin (`flatprov`) into the isolated
HERMES_HOME so the generic declared/live PUT + GET paths stay covered. The
lazy_deps plugin-owned-range test keeps mem0ai as its subject.
compress_now() hands only the head to _compress_context() and rejoins the
kept exchanges in memory afterwards. With compression.in_place (the
default), the commit's archive_and_compact() archives every active row at
or below the lease watermark, which includes the kept tail's rows, and
inserts only the compacted head. Its tail_count rewind then lands on the
newest rows under the watermark, which are the kept tail itself, so those
rows end up active=0, compacted=0 with no live copy. No surface writes them
back: the CLI re-flushes only after a rotation, the gateway skips in-place
on purpose, and the TUI only swaps its in-memory history. The exchanges the
user asked to keep verbatim are gone on resume, and on the gateway from the
very next message.
compress_now() now passes copies of the tail as verbatim_tail. The in-place
commit stores head + tail in the same archive_and_compact() transaction,
joined with the same seam rejoin_compressed_head_and_tail() builds in
memory, and adds the tail to tail_count. The rewind flags then land on the
kept tail's originals and on compress()'s own carried rows, which were
left compacted=1 and shown twice in the resumed display history. The
copies are stamped as persisted and compress_now() returns the stored list
instead of rejoining the tail a second time. Rotation, no-op and
rolled-back commits are unchanged: the copies stay unstamped and the
caller's tail is rejoined as before.
Keep the two tests that are red on unchanged main: previous_content comes
from the locked native-store result (single + batch replace/remove, batch
ordering) and never from caller arguments or build_metadata. Drop the
uncommitted-batch cases (they pass on main -- the no-notification gate on a
failed write predates this change) and the memory/user target axis (same
store code path).
Salvage of #118903 by @ehz0ah.
Micro-compaction carries a non-contiguous prefix and suffix around its summary marker. Using tail_count=len(result)-1 incorrectly marked summarized assistant/tool rows as rewind-only active=0, compacted=0.
Pass the exact unchanged carried messages instead and resolve their durable originals transactionally by row id or unique identity+timestamp, preserving summarized rows as compacted history.
Fixes#118481
tools-reference.md listed manage_catalog under the connections section, which the
reference-docs contract resolves against the connections toolset. The agent-runtime
post-hook ownership test enumerates every tool in AGENT_RUNTIME_POST_HOOK_TOOL_NAMES;
manage_catalog now runs through both executor paths there.
An empty catalog now legitimately triggers the 0.0.0 sentinel fallback (two requests per
site), so the fake answers one model and the one-request-per-site count still holds; every
request still carries the residency headers.
`_cap_binds` was the last inline copy of the cap-clamped-to-window predicate that #119406
centralised; the banner still prints the raw configured cap. Plugin engines without the
helper report "not binding", as before when no cap was set. The threshold_tokens comment
and the new test module stop narrating the incident; the parametrize takes the whole agent
config so the `{}` failure path is spelled out rather than derived from a ternary.
`_parse_compression_config` documents "Defaults here MUST match DEFAULT_CONFIG" because the
config-load failure path hands it `{}`. Every key had an inline fallback except the cap
added in #115986, so a broken config.yaml kept `threshold=0.50` but silently dropped the
256K cap — the pre-#115986 500K trigger on 1M-window models. An explicit null still
means ratio-only.
* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates
Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.
* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources
Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.
* refactor(copilot): gh candidates through locate_command and the Homebrew table
First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.
* feat(platform): availability() over an application declaration
locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.
* feat(platform): application declarations parsed into AppDef per OS
The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.
* feat(mcp): check_fn honours a registered application declaration
_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).
* feat(skills): requires_apps gate through registered declarations
Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.
* docs: application declarations page
The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.
* test(platform): declaration parser, availability, gates
The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
_fire_reasoning_delta called stream_reasoning_deltas_enabled() on every reasoning
token; each call took _CONFIG_LOCK and paid load_config()'s full deepcopy on a
cache hit, so every concurrently streaming thread in a WebUI/gateway process
serialized behind one lock on the token path.
Cache the opt-in on the agent for the life of one stream (reset in
_reset_stream_delivery_tracking, so the next request re-reads config) and use
load_config_readonly() for the lookup itself.
(cherry picked from commit 20fab007e5c3e399e3923bb6d66649918292825c)
Only column-0 bullets participate in `_drop_repeated_recall_lines`.
An indented line is a continuation of the bullet above it (nested
child, provenance, wrapped prose) and is kept verbatim, so identical
children under two different parents both survive. The docstring
already promised this; the code now matches.
PROOF: with the old code, `- Project A\n - status: active\n- Project B\n
- status: active\n` lost B's child (probe s4_eff_nested.py) and the
new assertion in test_a_bullet_with_continuation_lines_is_never_touched
fails (AssertionError at :49); after the change the probe returns the
input unchanged and tests/agent/test_memory_context_dedupe.py passes.
The memory pick carries at most two invariant tests. The section-scoping
assert added by the follow-up (identical placeholder bullets under prose
AND markdown headings both survive) is the second half of the same
invariant the repeated-bullet test states — once per section — so it
lives there now, and test_identical_placeholder_bullets_under_different_headings_both_survive
is gone.
PROOF: with the `seen.clear()` section reset at agent/memory_manager.py:303
replaced by `pass`, tests/agent/test_memory_context_dedupe.py fails
1/2 (test_a_repeated_bullet_is_kept_once_in_first_position, prose-shape
assert); reverted, 2/2 pass. ruff clean, footguns clean, real import OK.
The per-section reset only fired on `#…`, `**…` and `---` lines. The
in-tree RetainDB provider writes prose headings (`Profile:`,
`Relevant memories:`, `Instructions:`), so its second section still
collapsed into the first: `Profile:\n- None\nRelevant memories:\n- None`
lost the second `- None` and left a heading claiming nothing — the
exact failure 931d240e14 set out to fix, for markdown headings only.
Now any non-empty, non-bullet line at column 0 (a heading of any style,
a `---`/`***`/`___` rule, prose) clears `seen`; bullets and indented
continuation lines stay inside the current section.
PROOF: tests/agent/test_memory_context_dedupe.py::
test_identical_placeholder_bullets_under_different_headings_both_survive
(RetainDB prose-heading shape + markdown shape). Red with
agent/memory_manager.py at the previous commit (prose case drops the
second `- None`), green after; the continuation-line carve-out test
stays green.
Each pick keeps the two tests that pin its invariant; the rest restated
the same behaviour from other angles.
- #117399: keep test_a_repeated_bullet_is_kept_once_in_first_position and
test_a_bullet_with_continuation_lines_is_never_touched (9 removed).
- #117789: keep test_archiving_a_skill_does_not_move_the_epoch and
test_installing_a_real_skill_still_moves_the_epoch (4 removed).
- #117951 (rewritten): one test, test_recording_a_reply_does_not_open_a_
second_connection — counts sqlite3.connect per record_obligation, red on
main by construction (main opens 2).
Review follow-up on two real defects in the first cut, both reproduced before fixing.
A dropped duplicate left its continuation lines behind, and they re-parented under the surviving
bullet: `- prefers draft PRs` / ` (logged 12 Jan, supermemory)` followed by the same headline
logged elsewhere collapsed into ONE entry carrying BOTH provenance lines — inventing a record
neither provider reported, then stamping it into `api_content` to be replayed every turn. That is
worse than the duplicate the change set out to remove. A bullet that carries continuation lines is
now never dropped and never suppresses a later one, so two entries that share a headline and differ
underneath it both survive.
And `stripped[:1] in ("-", "*")` treated `**Preferences**` as a list item, so a repeated bold
section heading was silently deleted — contradicting the docstring's promise that headings survive.
The marker test now requires a real bullet: marker, whitespace, content (`[-*+]\s+\S`), which also
keeps `*emphasis*` out. Numbered items stay out of scope, and the docstring says so rather than
claiming "list items" generally.
Measured again on the same 166 real blocks: still 219,260 of 1,302,220 body bytes — 16.8%,
unchanged. Every duplicate in that sample is a self-contained bullet, so correctness cost nothing.
Tests: a bullet with continuation lines is untouched; a nested bullet is a child, not a repeat;
same-depth indented duplicates still collapse; bold headings, emphasis and numbered items are left
as written. Removing the continuation guard fails two. The tautological assertion the review flagged
is replaced by an exact whole-body comparison.
(cherry picked from commit 37a36f23658b8ca1e112f46ff60a705a66f80a5b)
Every user turn injects a `<memory-context>` block built from the providers' prefetch, and nothing
dedupes it: a provider merges several stores and `prefetch_all` merges several providers, so one
prefetch routinely surfaces the same fact two or three times.
It is not paid once. The composed block is stamped into the user row's `api_content` sidecar and
replayed verbatim on every later request for as long as that row is in context — deliberately, so
the prompt-cache prefix stays byte-stable. Compaction cannot reclaim it either: `_demote_tool_result_at`
bails on anything that is not `role == "tool"`, so a user row is never shrunk and drags its memory
dump along while counting against the tail budget. A duplicated bullet therefore costs its tokens
once per turn, for the life of the row, while telling the model nothing that same block has not
already said.
Measured on a real operator state.db (166 rows carrying a block, 46 sessions): 219,260 of 1,302,220
body bytes — 16.8%, ~54,815 tokens on first send alone — are list items byte-identical to an earlier
line of the SAME block. Median block ~2,034 tokens per user turn.
Only list items are considered, and only when identical after stripping: headings, prose, blank
lines and `---` rules are left exactly as the provider wrote them, so sections still read as
written and a line repeated deliberately as prose is untouched. Within one block only — deduping
against earlier turns would save far more (73% of bullet bytes in that sample were already sent
earlier in the session) but changes what the model sees on a turn, so it is not done here.
The "provider returned pre-wrapped context" warning stays keyed on sanitization alone: a deduped
bullet is routine, not a provider fault.
Fixes#117397
(cherry picked from commit e3b08ca9b1708a7cddc2f1a2191b51d60d194624)
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
tests/agent/test_context_compressor.py:3397 asserted
`sampled.count("chars elided") < 7`, hardcoding the n-1 gap count that
the 8-slice constant implies and that the comment explained in prose.
Use `ContextCompressor._SAMPLED_INPUT_SLICES - 1` so the invariant
("the extension pass closed at least one initial gap") survives a
change to the slice count instead of silently becoming a change
detector. `_SAMPLED_INPUT_SLICES - 1 == 7` today, so the assertion
value is unchanged.
No fixture exercised the final `selected = _merged(selected)` in
_sample_summary_records: the re-gate mutant that dropped it survived
all 8 bounding tests and 40 random probe shapes. Add 20 x 8.5K records
(170K > 160K cap, headroom > one 8.5K gap) so the round-robin extension
closes the gaps between the newer slices. Assert consecutive shown
records are joined by exactly "\n\n" (an unmerged adjacent pair renders
as "...xxx[USER]: record-..." with no separator), remaining gaps carry a
correctly numbered marker, fewer than 7 markers survive, and
sampled_record_count equals the whole records shown.
RED with the final `_merged` call removed: AssertionError (14, 15);
GREEN at head.
Finding: $D/regate/C.md suggestion 2 (M-c, agent/context_compressor.py:3623).
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.
Before -> after (records per slice, fill unchanged):
10000x100 [1,1,1,1,1,1,1,8] -> [1,2,2,2,2,2,2,2] fill 0.944
8000x100 [2,2,2,2,2,2,2,5] -> [2,2,2,2,2,3,3,3] fill 0.957
4000x300 [4,4,4,4,4,4,4,11] -> [4,5,5,5,5,5,5,5] fill 0.987
2000x600 [9,9,9,9,9,9,9,15] -> [9,9,10,10,10,10,10,10] fill 0.995
500x2000 [37,...,37,39] -> [37,...,37,38,38] fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.
Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).
Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
tests/agent/test_lean_single_aux_call.py:163 accepted either `<seg09>`
or a `content[-500:]` substring. The last slice is anchored to the
newest record, so the fallback can never be what passes; assert
`"<seg09>" in out` outright and drop the now-unused `content`.
Red when the tail anchor is defeated (end = len(records) - 1): 1 failed;
green at head (gate 2c suggestion).
Lean sampling is the documented budget mechanism, but nothing reported
how much of a region actually reached the summarizer, which is the gap
the reporter narrowed #118362 to. _sample_summary_records now returns
record-level coverage counters alongside the bounded transcript and
_generate_summary writes them into the active compression telemetry
(summary_input_chars / sampled_chars / omitted_chars / record_count /
sampled_record_count / elided_record_count), so the existing
compression_attempt JSON log line carries them with no transcript
content. Char counters cover record content only, not separators or
markers. Elision markers also name the one-based record range they
stand for, so a reader can tell how many turns a gap hides, not only
how many characters.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Replace the eight sampler tests with two invariants. The boundary test
now drives _generate_summary so it fails on main for the actual defect
(a retained section starting mid-record) instead of on a helper name,
and folds in the multi-paragraph case since blank lines inside a message
were what broke the string-split approach. Dropped: covers_multiple_
regions and generate_summary_lean_samples_structural_records (green on
main, no teeth), omission_markers_account_separators_exactly (asserted
only that a marker exists), newest_message_ending_in_blank_lines (moot
once records are structural), and preserves_oversized_last_record (a
record above the whole cap is unreachable through _serialize_for_summary
because _CONTENT_MAX bounds message bodies). The remaining oversized
test asserts behaviour — tail kept, oversized record bounded — not the
marker wording.
_sample_summary_input accepted either a flat string or a record
sequence and, for strings, rebuilt records with content.split("\n\n").
The PR's own third commit established that "\n\n" is not a record
delimiter (message bodies contain blank lines), so that fallback would
re-introduce split records for any caller that reached it. The only
production caller already passes records to _sample_summary_records;
drop the dual-signature wrapper and move the test callers to records.
Preserve record framing structurally so messages containing internal blank lines are not split into pseudo-records.
- agent/context_compressor.py:
* Extract _serialize_records_for_summary() to retain turn boundaries as a sequence of records.
* Add _sample_summary_records() to sample across serialized turn records directly without lossy double-newline splitting.
* Correct separator accounting in omission markers at gap boundaries so separator counts match full serialized text.
* Update _sample_summary_input() to support structural records while preserving flat-string fallback with trailing fragment pruning.
* Wire _generate_summary() to use _serialize_records_for_summary() and _sample_summary_records() in lean mode.
- tests/agent/test_context_compressor.py:
* Test multi-paragraph serialized messages maintain role boundaries and keep newest message paragraphs together.
* Test newest message ending in blank lines does not produce empty tail anchor.
* Test exact omission marker separator accounting.
* Test end-to-end _generate_summary prompt assembly with multi-paragraph messages.
(cherry picked from commit f68f1bfcf037dc5ddecf6cc4032023cd88c496c8)