Commit Graph

40773 Commits

Author SHA1 Message Date
kshitijk4poor
b3a1900e72 refactor(gateway): reuse utils.file_signature for the pairing cache key
Also collapse the artifact-store lock lookup to a single setdefault.
2026-09-23 21:47:15 +05:30
SilverNine
709aaf9764 fix(slack): bound nested attachment block text across the whole unfurl array
a74e0155b6 made attachments[].blocks[] reach the agent through
_append_link_unfurls, but rendered each attachment's blocks with no ceiling.
Slack allows 20 attachments per message, so one alert could project 20x what
a single attachment does (measured: 3,247 chars for 1 -> 64,855 for 20 with
8x400-char rich_text sections each), while the top-level blocks path caps
once at 6000.

Share one budget (_SLACK_UNFURL_BLOCKS_MAX_CHARS, the same 6000 the top-level
path uses) across the array: the first attachment keeps its body, later ones
are truncated against the remainder, and a spent budget still leaves every
header visible. After: 3,247 -> 6,659 chars at 20 attachments.

(cherry picked from commit 3efc1e4c532c38a220d4c6ea53745d59d334aa3e)
2026-09-23 21:47:15 +05:30
Michael Versluis (Berry)
aca906df2c fix(gateway): keep status caches bounded across awaits (#87479)
Telegram's per-(chat_id, status_key) status-message cache grew without
bound; give it the same _STATUS_MESSAGE_IDS_MAX=2000 FIFO half-trim the
Slack adapter already has. In both adapters, guard the post-await write-back
after a successful edit with a compare-before-write (only re-store the id if
the cached entry is still the one we edited) so an eviction or replacement
that happened during the await is not undone.

Partial salvage of #87480: kept the Telegram bound and both compare-before-write
guards (re-applied by hand, 17586 behind), defined the max as a class attr
like Slack instead of an instance attr, dropped the 4 new tests.

(cherry picked from commit 43ae95e98e)
2026-09-23 21:47:15 +05:30
devorun
b331e5d282 fix(buzz): persist channel cursors off the event loop
_handle_events called _save_cursors inline whenever a batch moved the
cursor. It ends in atomic_json_write (mkstemp + fsync + os.replace), and
on the WebSocket transport _handle_events runs once per inbound EVENT
frame, so every message paid an fsync on the gateway's event loop,
stalling every other adapter and in-flight turn for its duration.

Split the snapshot from the write: the payload is still built on the
loop (_channel_state is loop-owned), and the write goes through
asyncio.to_thread. The snapshot is taken under an asyncio.Lock so a
slower, older write can never land after a newer one and regress the
durable cursor. connect() keeps the synchronous _save_cursors.

(cherry picked from commit 299929741429262f497b86a67940ffb37faa1694)
2026-09-23 21:47:15 +05:30
spfcraze
5a035357f5 perf(gateway): make no-override session scan lazy behind debug
The "No session model override" debug line eagerly built a list over every
session's conversation.model_override as a logger.debug argument, 2-4 times
per inbound message, even when DEBUG logging was off. Gate the whole call
on logger.isEnabledFor(logging.DEBUG) so the scan costs nothing in normal
operation.

Re-applied by hand from #77810 (targeted gateway/run.py; the code now lives
in gateway/run_turn.py, logic unchanged). Dropped the PR's routing test.

(cherry picked from commit c3c2bf3801)
2026-09-23 21:47:15 +05:30
kshitijk4poor
9d8ec4c99c test(gateway): cover single-flight off-loop artifact store construction
One invariant test replacing the dropped 392-line file: five racing first
requests build the store once, off the loop thread, and cache hits skip
the lock.
2026-09-23 21:47:15 +05:30
Kyzcreig
f276ff3f6f fix(gateway): move the one-shot artifact transport off the event loop
Partial salvage of #119405: kept the four api_server.py hunks (per-profile
single-flight asyncio.Lock dict, _artifact_store_for_async, awaited
to_thread for store.store and store.load), dropped the 392-line test file
because it is mostly change-detector/timing assertions.

(cherry picked from commit cdfcc39181)
2026-09-23 21:47:15 +05:30
devorun
73dc249cca fix(weixin): persist the long-poll cursor off the event loop, only when it moves
_poll_loop called _save_sync_buf inline after every getUpdates response.
It ends in atomic_json_write (mkstemp + fsync + os.replace), so each poll
cycle blocked the gateway's event loop for the duration of an fsync, and
it did so even when nothing changed: an empty long-poll, and the timeout
sentinel from _get_updates, echo the current buffer back, so the same
value was rewritten every cycle.

Write only when the buffer differs from the one in memory, and dispatch
the write with asyncio.to_thread. The loop awaits it before the next
poll, so writes stay serialized.

(cherry picked from commit 316b775941ffcc7e1bdd17c044e879cbca1eb14f)
2026-09-23 21:47:15 +05:30
Adolanium
d644764c9d perf(gateway): cache pairing approved list by file mtime instead of re-reading per message
The authz gate calls PairingStore.is_approved on every inbound message,
and each call re-read and re-parsed {platform}-approved.json from disk
plus rebuilt per-user alias sets. Approvals change only through
explicit pairing writes (this process or the pairing CLI), which always
bump the file mtime, so a (mtime_ns, size)-keyed cache is an exact
invalidator. A stat per message replaces a read + parse per message.

The loud PermissionError warning in _load_json still runs on every
real read, and same-process writes through _save_json invalidate
automatically via the mtime bump.

(cherry picked from commit 0a0d2296bdf36ae58ca4443c2d26b8fc577c88dc)
2026-09-23 21:47:15 +05:30
kshitijk4poor
7bcaf554b0 test(gateway): trim macOS proxy-cache tests to the two invariants
Keep hit-caching and TTL expiry; drop reset/non-darwin/failure-path
tests that duplicate the same monkeypatched fork counter.
2026-09-23 21:47:15 +05:30
John Paul Soliva
16b18ce602 perf(gateway): cache the macOS system-proxy probe instead of forking per chunk
_detect_macos_system_proxy() forks `scutil --proxy` with no cache, and
resolve_proxy_url() calls it on the SEND path — inside the per-chunk sender in
_send_chunks and again per media attachment, not once per adapter. It is also
reached whenever no explicit proxy is configured, which is the common case.

Measured on macOS: a bare `scutil --proxy` fork is 11.05 ms, and
resolve_proxy_url("DISCORD_PROXY") is 11.209 ms — essentially all fork+exec, to
read an OS setting that changes when someone edits Network Settings or joins a
VPN. A three-chunk reply paid ~34 ms of it; a message with two attachments
paid more again. macOS-only, so it never shows up in CI.

Memoise the answer for 60 s. The TTL is the staleness a proxy change can
suffer; a send that goes out on a stale answer fails and is retried, the same
outcome as any transient proxy error. A failing scutil is cached too, so a
broken or slow one cannot re-fork per chunk either. reset_macos_proxy_cache()
forces a re-read.

resolve_proxy_url: 11.209 -> 0.742 ms (-93%), same value returned.

(cherry picked from commit 635b06c1007e5806bf83e9d1aef2db0be1a486c2)
2026-09-23 21:47:15 +05:30
kshitijk4poor
a2d5bb07c5 chore: map contributor emails for the perf salvage stack 2026-09-23 21:47:15 +05:30
brooklyn!
bb4130f6f5 test(desktop): cover Unicode profile initials and default glyph behavior 2026-09-23 11:14:55 -05:00
brooklyn!
2b18ca7867 fix(desktop): preserve Unicode graphemes in profile initials
Use one shared grapheme-safe initial without widening the existing rail or changing profile identity and color behavior. Adapt the single-initial direction from the competing proposals.

Co-authored-by: Moisés Valero <moisesvs84@gmail.com>
2026-09-23 11:14:55 -05:00
brooklyn!
3f06838e04 test(desktop): keep archive instructions aligned with row gestures 2026-09-23 11:14:28 -05:00
mrchen0709
afea7ff5fc fix(desktop): teach the real archive gesture in the archived-sessions copy
The Archived sessions intro told users to Ctrl/⌘-click a chat in the sidebar to
archive it. That gesture opens the chat in a new tab: resolveSessionRowClick maps
⌘/⌃-click to 'newTab' and only ⌥+⇧-click to 'archive'. Anyone following the copy
never archives anything, and that line is the only in-app hint the gesture exists.

Swap the modifier in the six locales carrying the string, keeping each
translation's wording. The resolver and the row's own comment already document
⌥+⇧ — only the user-facing copy was stale.

Fixes #107999
2026-09-23 11:14:28 -05:00
brooklyn!
af63503a0f test(desktop): preserve changelog semantics across locale labels 2026-09-23 11:11:29 -05:00
brooklyn!
c312d39931 fix(desktop): localize update groups and fallback release copy
Adapt David Metcalfe’s proposal to the current update overlay, preserving commit identity, ordering, counts, and the pure formatter’s English defaults.

Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
2026-09-23 11:11:29 -05:00
Xipong
fe9a5ccafa fix: offload dashboard log reads from event loop
(cherry picked from commit 6a41c1a3bb66b409625df9204362d64258bee36d)
2026-09-23 21:30:08 +05:30
kshitijk4poor
e13f985844 fix(web): move the remaining dashboard file reads off the event loop
Widen #117979 to the sibling routes in the same router: /api/files/read
(inline read_bytes+b64encode), /api/fs/read-text (bounded prefix read) and
/api/fs/read-data-url (_fs_read_bytes+b64encode) all still ran their file
I/O on the loop thread. Reuse the same helper / asyncio.to_thread pattern so
a slow disk or large bounded payload no longer stalls every other request.
2026-09-23 21:30:08 +05:30
wangtao
3505e4ce77 fix(web): offload dashboard media encoding
(cherry picked from commit 351b77daab262f710ab1f380e39005ff7e45d4a7)
2026-09-23 21:30:08 +05:30
kshitijk4poor
61202bf2fb perf(codex): categorise each distinct non-ASCII character once in the format-control scan
No ASCII code point is a Unicode format control (Cf), so ASCII text returns on the
O(1) isascii() flag and other text categorises set(text) minus ASCII instead of every
character. Output is unchanged (same unicodedata predicate).

Python 3.12, per call: CJK 9.2k chars 3.60 -> 1.57 ms, mixed 3.49 -> 0.30 ms,
ASCII 3.16 -> ~0 ms (main -> this commit). A per-match finditer over non-ASCII
characters was slower than main on CJK (11.0 ms) and is not used.
2026-09-23 21:28:38 +05:30
kshitijk4poor
2ec8a89e53 test(redact): keep the ReDoS guard's wall-clock bound far from the fixed cost
300k digits took 0.36-0.73 s on the fixed tree against a 2 s ceiling, which is
flake territory on a loaded runner. 100k digits costs ~0.15 s fixed while the
unanchored pattern still takes tens of seconds.
2026-09-23 21:28:38 +05:30
Efe Büken
a1944881cf perf(codex): skip the per-char format-control scan for ASCII text
_neutralize_harmony_tokens walked every character of each string leaf that
contains both '<' and '|' through unicodedata.category() to look for format
controls. No ASCII code point is Cf and str.isascii() is an O(1) flag check,
so gate the scan on it. 160 KB ASCII code text: 85.7 ms -> 0.18 ms per call;
output unchanged.

Partial salvage of #98340: kept the isascii fast path; dropped the compiled
per-Unicode-version Cf pattern table, the benchmark script and the
table-drift/differential-matrix tests (source-reading change-detectors).

(cherry picked from commit 84e6561ce710b861ec1861e92558d2ae433dccab)
2026-09-23 21:28:38 +05:30
Riccardo Vecchi
36c2f05a00 perf(agent): skip surrogate regex for ASCII text
_sanitize_surrogates ran the surrogate regex over every string leaf of the
outbound messages and kwargs on each request. Surrogates are never ASCII and
str.isascii() is an O(1) flag check, so gate the scan on it; same gate on
_strip_non_ascii. 50 KB ASCII leaf: 1017 us -> 0.1 us; output unchanged.

Partial salvage of #83839: kept the isascii fast-path idea as a single gate
at the top of _sanitize_surrogates (main refactored to a generic fix-callable
walk, so the PR's per-site hunks no longer apply); dropped the monkeypatch
"never invokes regex" tests because they are change-detectors.

(cherry picked from commit 36d2575a12e2ffae1a3286208f269fb3d92acbf5)
2026-09-23 21:28:38 +05:30
illidan
bca2f0f3da perf(agent): avoid quadratic work in tool history pairing
_classify_tool_call_orphans rebuilt the surviving-result variant list and
scanned all of it for every declared tool call (O(calls x results)).
Orphan result variants are disjoint from every declared call, so matching
against the already-computed union of all result variants is equivalent
and O(1) per call. 1000 pairs: 50.9 ms -> 2.4 ms per classification.

Partial salvage of #105660: kept the _classify_tool_call_orphans union
reuse; dropped the messages[i+1:] slice -> range() hunk (sub-ms) and
tests/agent/test_tool_pairing_scaling.py (read-count assertions are a
change-detector). Hunk re-derived by hand: PR base was ~7.9k commits behind.

(cherry picked from commit 6879a69a7195c26162c63ba9474a07638e35e839)
2026-09-23 21:28:38 +05:30
Adolanium
c40655ebb6 perf(moa): weigh reference trim once per message instead of re-estimating per pop
The reference context-fit loop re-ran estimate_messages_tokens_rough on
head + body after every pop, paying one full memo walk per dropped
frame (O(n^2) on long histories). The estimator is a pure sum of
per-message weights, so weigh each message once, track a running
total, and subtract on pop. The pop sequence is unchanged and the trim
result is identical to the naive loop, pinned by a randomized
equivalence test against the original implementation plus an O(n)
estimator-call-count guard.

n=502 advisory frames trimmed to 106: 100.8ms -> 0.9ms per call.

cherry-picked-from: 226a8786f2cba9380be930d28604c81880d3f6f7
2026-09-23 21:28:38 +05:30
Xipong
774df367ee perf(prompt-caching): copy only mutable cache-plan rows
build_prompt_cache_plan deep-copied the whole request history on every request
just to strip stale markers and place at most four new ones. Take a shallow list
copy, copy-on-write only the rows the marker strip may touch (rows carrying
cache_control or list content), and deep-copy the rows that receive markers in
the tool-cache layout. The planner's output is byte-identical; the caller's
history is never mutated (#106061).

(cherry picked from commit a74fe3ed8af9fd38b4e4a7221a231b7484db9334)
2026-09-23 21:28:38 +05:30
kshitijk4poor
c191c34566 perf(tools): resolve register()'s existing-entry check without a merged copy
register() rebuilt the merged global+overlay dict for scoped registrations to
find the existing entry; route it through the same two-lookup helper as
get_entry (`_lookup`), keeping the scope-None path global-only as before.
`_toolset_entries` still walks the merged view because it needs every entry
of a toolset, not one name.

Adds the equivalence pin from #106092 (liuhao1024): get_entry must return
exactly what the merged view returns for every name/scope combination.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-23 21:28:38 +05:30
Xipong
30cd1e97d0 perf(tools): resolve registry entries without copying all tools
ToolRegistry.get_entry() built the full global+overlay merged dict just to
fetch one entry; per-tool callers (dispatch, tool_search_catalog.build_catalog)
turn that into quadratic work. Resolve against the profile overlay first, then
the global map, under the existing lock. Overlay shadows global, unknown scope
falls back to global, ambient scope still comes from current_scope_key().

Partial salvage of #106076: kept the get_entry fix, dropped the tracemalloc
allocation test because it is a fragile change-detector. Same fix was also
proposed in #106092 (#106062).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
(cherry picked from commit b8665ade15decefbfa1002405887bb32c94ef701)
2026-09-23 21:28:38 +05:30
Israel Lot
70f79c8cde fix(agent): anchor the Telegram token regex to a digit-run start (ReDoS)
`_TELEGRAM_RE` (`(bot)?(\d{8,}):([-A-Za-z0-9_]{30,})`) runs over every
redacted surface that contains a colon. On a long digit run with no
`:<token>` after it, the unanchored greedy `\d{8,}` is retried from every
digit of the run, so the scan is quadratic in the run length and holds
the GIL for its whole duration.

Real trigger: the foreground terminal tool spills a 2 MB `gh pr diff`
to disk and redacts the full spill (`_redact_spill_file`); the diff
carried a Unity `_typelessdata: 000...` line with ~1.5M consecutive
digits. That redaction pinned one core at 100% for hours after the
subprocess had already exited, so the command timeout could not fire,
and every session in the same `hermes acp` worker stalled. Any tool
output the agent reads (hex/decimal dumps, minified assets, logs) can
reach the same path.

A `(?<!\d)` lookbehind pins the bot id to the start of its digit run,
which is the only place a real token id can begin, so a non-matching
run is scanned once. Matching is unchanged for real tokens.

Measured on the unpatched pattern, `_typelessdata: ` + N zeros:
10k 0.41s, 20k 1.55s, 40k 6.2s (x4 per doubling); patched: 0.004s,
0.003s, 0.006s, 80k 0.014s.

(cherry picked from commit c48cb6bf93a15cd28fd253c5293413d72081fdc4)
2026-09-23 21:28:38 +05:30
kshitijk4poor
cf469e66f9 chore: map contributor emails for the perf salvage stack 2026-09-23 21:28:38 +05:30
kshitijk4poor
1572836184 perf(gateway): parse readiness and bundled-manifest YAML with the C loader
Same class as the config/manifest loader swap earlier in this stack. The
/ready probe re-read and pure-Python-parsed config.yaml on every poll, inside
the aiohttp handler (110 KB seeded config: 2332 -> 487 ms per probe on a
loaded runner). It now uses utils.load_yaml_file_readonly: the C loader, and
repeat probes reuse the parse until the file signature changes. Parse errors
are not cached, so an edit that breaks or fixes the file shows on the next
probe. The bundled-platform manifest reader was the one startup manifest
reader left on yaml.safe_load. managed_scope imports fast_safe_load at module
level next to file_signature instead of lazily.
2026-09-23 21:26:49 +05:30
Gianpietro Dal Zio
fae101eddc fix(profiles): serialize cold skill-count scans
_count_skills() captured `now` BEFORE walking the skills dir and stored it as the
cache timestamp, so a scan longer than the TTL published an already-expired entry
(the next caller re-walked immediately). Concurrent direct callers (profile info,
dump, non-lazy list) could also each run the same cold walk.

Stamp the cache entry after the walk, and take a per-skills-dir scan lock with a
re-check under it so overlapping callers for the same profile share one walk while
different profiles' background scans still run in parallel.

Partial salvage of #107137: kept the stamp-after-walk + locked double-check hunk,
replaced the single global _SKILL_COUNT_SCAN_LOCK with a per-key lock so N profiles'
scans don't serialize, and trimmed the threading.Event test file to one invariant
test (published timestamp >= scan end).

(cherry picked from commit 30aba1f737)
(cherry picked from commit c1ec369828)
2026-09-23 21:26:49 +05:30
ygd58
c1a59688c1 fix(constants): engage the warn-once latch on the first check, not only on the warning branch
_warn_profile_fallback_once() only set _profile_fallback_warned inside the warning branch,
so in the common case (no active_profile, or "default") every get_hermes_home() call with
HERMES_HOME unset re-ran the platform-default resolution + exists() stat + read_text().
Latch unconditionally after the first check; a profile activated mid-process after the
first call is out of scope for a one-shot startup-time check.

Fixes #90065.

Partial salvage of #90101: kept the hermes_constants.py latch hunk and one regression
test (latch engages on first check without a warning); dropped the call-count and
stderr-message tests to keep the item to a single invariant test.

(cherry picked from commit e033bae20abb88073ad58f3dd95ef0735603985c)
2026-09-23 21:26:49 +05:30
John Paul Soliva
99e8251e5b perf(gateway): capability probes read config without the defensive deepcopy
_slack_tools_loaded/_discord_tools_loaded run per turn via _ephemeral_change_key and
_get_platform_tools only reads the config, so the deepcopy in load_config() is pure
overhead on the event loop. Use load_config_readonly().

Partial salvage of #117993: kept the two gateway/session.py hunks, dropped
TestCapabilityProbeConfigLoader because it monkeypatched cfg.load_config to raise —
a change-detector on the symbol read, not an invariant.

(cherry picked from commit a487903180090c5a2fccd4749b05d302cf2740b1)
2026-09-23 21:26:49 +05:30
John Paul Soliva
16db511ed5 perf(config): route the config and manifest loaders through fast_safe_load
utils.fast_safe_load already exists, is pinned by tests/test_fast_safe_load.py, and
its comment names exactly these payers: 'startup parses config.yaml and every plugin
manifest, so the slow path cost ~0.9 s of cold start'. The migration was started —
hermes_cli/config.py uses it eight times, hermes_cli/main.py and hermes_cli/plugins.py
too — but the file-level loaders it was written for were never converted.

The cost is config SIZE, and the size is the installer's doing: it seeds config.yaml
by copying cli-config.yaml.example, 120,897 bytes of mostly comments. Nothing caches
load_gateway_config() and it has 238 production call sites.

Profiled before assuming a cause — reader.forward 34 ms, scanner.scan_to_next_token
28 ms, reader.peek 13 ms: the pure-Python PyYAML scanner, nothing else.

Measured A/B on one realistic pass (gateway config + every bundled plugin
description), medians of five runs, __pycache__ cleared between arms:
48.8 ms -> 2.46 ms. Per path: load_gateway_config 48.3 -> 1.84 ms, managed
config.yaml 44.4 -> 1.03 ms, 105 plugin.yaml manifests 59.3 -> 5.9 ms.

Same parse, same restricted tag set, same result — only the loader changes. Drops the
three 'import yaml' statements the swap orphaned.

(cherry picked from commit a4efd6a4506043159310eba02c32b673c455b88b)
2026-09-23 21:26:49 +05:30
kshitijk4poor
8646ddc19a chore: map contributor emails for the perf salvage stack 2026-09-23 21:26:49 +05:30
hermes-seaeye[bot]
6cd727bbfa fmt(js): npm run fix on merge (#120400)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 15:53:24 +00:00
hermes-seaeye[bot]
8191aa0660 fmt(js): npm run fix on merge (#120392)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 15:46:42 +00:00
doresa0
03544a73be fix(desktop): keep reasoning settings label on one line
Prevent short CJK labels like 推理 from wrapping awkwardly next to the
effort select in Model settings.

Closes #83118
2026-09-23 10:38:40 -05:00
teknium1
38c9611791 test(gateway): served-profile stop parks under the host (rebase over #120143)
test_pooled_served_profile_backend_unscoped pinned the pre-parking contract (stop on a served
profile refuses). With per-profile parking, stop on a served profile is the parking verb and
returns no refusal; start while unparked still refuses. Same one-line contract change already
made in test_gateway_multiplex_served_record.py.
2026-09-23 08:25:28 -07:00
teknium1
9ab2ad3d3f docs(gateway): parked vs standalone precedence; gap (1) of the standalone shim is closed by parking 2026-09-23 08:25:28 -07:00
teknium1
f6ce02eafa fix(gateway): parking composes with gateway.standalone and the dashboard Stop/Start twin
gateway.standalone wins over the parked marker: profile_lifecycle() returns
False for an opted-out profile, so `-p X gateway stop|start` keeps addressing
X's own gateway process and never writes gateway.parked, even while a stale
host record still lists X. Every installed-roster caller now threads both
kwargs (`include_standalone=True, include_parked=True`): the host-attach peer
walk and the standalone boot notice were reading the roster without parked
profiles.

Dashboard twin (#119886 class): `/api/gateway/stop?profile=X` on a served
profile no longer answers 409 — the spawned `hermes -p X gateway stop` parks
it; `/api/gateway/start` on a parked profile is allowed while a host
multiplexer is live (the child unparks it) and still refused when nothing can
serve it. Tests trimmed to the invariants: the marker-appearing case was a
subset of the boot-and-reconcile test.
2026-09-23 08:25:28 -07:00
Victor Kyriazakos
85d28d343a docs(gateway): stopping one profile without stopping the host
The stop, start and restart verbs on a profile served by the host multiplexer,
the gateway.parked marker (provisioning may pre-create it), reconcile timing,
and what parked means in gateway status.
2026-09-23 08:25:28 -07:00
Victor Kyriazakos
4c342c05de feat(gateway): stop, start and restart one profile under the host multiplexer
Under the host multiplexer every installed named profile came online because
its directory existed, and the only way to stop one profile's bots was to stop
the host, which stopped everyone's. Both are fleet-operator blockers.

`hermes -p X gateway stop` on a profile served by the host now parks it: it
writes `profiles/X/gateway.parked` first (so the 30s reconcile cannot re-add X
between the verb and the marker) and sends the new `unserve-profile` control
verb, which tears down X's adapters, reconnects and cron inside X's own scope
(the teardown `_unserve_profile` already used for deleted profiles). `start`
removes the marker and sends `serve-profile`, which runs the same add-path the
reconcile loop uses. `restart` cycles both without parking. The default
profile keeps today's whole-host meaning.

`profiles_to_serve()` skips parked profiles, so adapters, cron, ingress
membership and the served record follow from one chokepoint; roster callers
that mean "every installed profile" (plugin deps, Windows update, dashboard
listing and topology, migration inventory) pass `include_parked=True`.
Provisioning may pre-create the marker: an installed profile stays offline
until an operator starts it. Host boot logs one INFO per parked profile;
`gateway status` shows `parked (hermes -p X gateway start)`.

Tests: profiles_to_serve contract, control verbs through the runner
(round-trip and every refusal), CLI marker-before-socket ordering, reconcile
honouring the marker both ways, migration inventory retaining parked
profiles, two-home E2E through the real loaders.
2026-09-23 08:25:28 -07:00
teknium1
bd970b0588 fix(profiles): route-only launch pin; profile delete names its own unit under multiplex
Builds on tancou's #119129 (cherry-picked above): the pin now lives in
get_routing_process_hermes_home() and only the four routed-profile DECISIONS read it.
get_process_hermes_home()/get_hermes_home() keep following HERMES_HOME, so an env-only
home switch in a multiplexed process resolves as before.

- set_multiplex_active(True) pins the launch home only when no host pin exists, and
  set_multiplex_active(False) releases only the pin it created itself. A transient toggle
  (gateway_migrate._multiplex_read_mode, cron external-worker restore) no longer drops an
  embedding host's explicit pin_process_hermes_home(launch).
- profiles._cleanup_gateway_service binds set_hermes_home_override(profile_dir) beside the
  env write. Under the previous head, DELETE /api/profiles/<x> from a multi-profile dashboard
  resolved get_service_name() against the pinned launch home -> bare `hermes-gateway`, and
  disabled/stopped/unlinked the HOST multiplexer's unit. Same path serves rename_profile.

Tests (red on the previous head): explicit pin survives True->False; env readers follow the
env while pinned; two-home delete removes hermes-gateway-victim and leaves hermes-gateway.
2026-09-23 08:09:47 -07:00
tancou
c7c9c18ccf fix(profiles): pin the launch home so a mirrored HERMES_HOME cannot flip routed-profile decisions
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.

Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.

Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.

Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.

Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.

Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 08:09:47 -07:00
teknium1
11d5c710fd test(i18n): a served profile's turn binds the home override, not a mutated HERMES_HOME
`test_language_is_per_profile_under_multiplex` switched profiles by rewriting
`os.environ["HERMES_HOME"]` after `set_multiplex_active(True)`. That is the T1
standalone contract (environ IS the profile); a multiplexed turn binds
`set_hermes_home_override`, and since the launch home is pinned at activation
(#119242) a later env mutation is deliberately ignored. The two assertions are
unchanged; only the profile-switch mechanism now matches production.
2026-09-23 08:09:47 -07:00
teknium1
64c1cba883 chore: map contributor email for salvaged kanban commit 2026-09-23 08:09:47 -07:00