A SOUL.md in HERMES_HOME that *documents* the canonical injection phrase as
security guidance ("...content telling you to ignore previous instructions...")
was replaced wholesale by a [BLOCKED: SOUL.md ...] marker, so the agent ran
with no identity/constitution and only a one-line agent.log warning said so.
SOUL.md is the user's own file: agent writes to it always go through the
protected-instruction approval gate (tools/file_tools_write_guards.py) and no
repository checkout can plant it, so it sits in the same trust class as
config.yaml — not a cloned repo's AGENTS.md. `_scan_context_content` gains a
`user_authored` mode that still scans, logs the matched pattern(s) at WARNING
and loads the content; only `load_soul_md` uses it. Project-dir context files
(.hermes.md, AGENTS.md and the other project-dir files), subdirectory hints,
memory and tool-result scanning are unchanged and keep blocking. No threat
pattern is narrowed: the reported-speech form cannot be separated from real
attacks by regex ("I want you to ignore all previous instructions" is a
canonical payload).
The /context manifest reports such a file as `flagged` (loaded, ⚠ "review the
file") so the warning is visible on CLI, TUI, gateway and Desktop rather than
buried in the log.
Fixes#112570
The cp mock rejected before writing anything, so origin/main's direct copy into
the final target also left no target and the test was green on both sides. Have
the mock write a partial plugin.js into whatever destination it was handed before
rejecting: with a direct copy that half-tree sits at <appRoot>/media and the
assertion fails; with staged publication it lands in the staging sibling and is
swept. Verified red with origin/main's desktop-plugins-root.ts swapped in.
Extract the contributor's staging + rename sequence into
`publishDesktopTree` and make `installDesktopPluginFromGit` use it too:
its `copyDesktopTree` had the same mkdir + cp shape, so a copy that failed
half-way left an empty `<root>/<name>/` and every later install attempt
was refused with "already exists. Enable force reinstall to replace it".
One helper, both copy sites, same guarantee: a published folder is always
complete (marker included) or absent.
Docs: describe the staged copy and the marker-less/plugin.js rule in the
desktop plugin SDK guide.
_resolve_from_pool now calls pool.select(model=...) so the per-model
Anthropic cooldown can be honoured at resolve time. Two untouched sibling
test files stubbed the pool with `select=lambda: entry`, which rejects the
keyword and turned CI red (4 failures: TypeError unexpected keyword
argument 'model'). Widen both stubs to `lambda **_kw: entry`, the same
shape test_credential_pool_provider_boundary.py already uses.
Also carry the per-model 429 paragraph and the sibling-module bullet into
the zh-Hans credential-pools mirror, which the review flagged as stale.
Follow-up to the salvaged #111787 commit (@KoNit-K), same mechanism as #75578 (@adikpb).
What:
- Move the per-model cooldown logic out of the 2.7k-line credential_pool.py facade into
agent/credential_pool_model_cooldowns.py (mixin + module helpers), keeping only the
select/has_available/next_available_at hooks and the mark_exhausted_and_rotate branch in the facade.
- The model cooldown uses the same TTL policy as a credential-wide 429 (_exhausted_ttl: provider
reset_at wins, a sole credential keeps its 60s bench) instead of a flat 1h, so a single-credential
user is never benched LONGER for the failed model than before.
- Drop the `quota_scope == "account"` check: nothing in the tree produces that key.
- Also keep billing_unverified 429s credential-wide, matching the credential-wide branch.
- _rebind_primary_credential_pool read `rt` that was not in its scope (NameError on every
post-fallback restore); pass primary_model from the caller instead.
- resolve_anthropic_token(model=...) gates only model-aware callers; model-less diagnostics
(usage display, model discovery) keep the key as before.
- _anthropic_token_or_raise names the cooled model instead of claiming no credentials exist.
- Simplify the salvaged call sites (unconditional select(model=)/resolve_anthropic_token(model=)),
widen test stubs that lacked the new kwargs, trim the tests to two invariants on a real temp store.
- Docs: credential-pools.md documents per-model Anthropic 429 cooldowns.
Why: a generic Anthropic 429 is a per-model rate limit; benching the whole credential took every
other Claude model offline while the env/borrowed token path handed the same benched key straight
back (#111769, #61451).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: adikpb <67222969+adikpb@users.noreply.github.com>
`test_moved_lazy_pointers_resolve_to_the_split_off_siblings_object` resolves every
moved-lazy alias. `hermes_cli.web_server.PtyBridge` / `PtyUnavailableError` point at
`hermes_cli.pty_bridge`, which imports `fcntl`/`termios`, so on native Windows the
test recorded them as `unresolvable: ModuleNotFoundError('fcntl')` and failed (#112576).
When the failure is a POSIX-only stdlib module on `win32`, the test now checks that the
declared target module exists (`find_spec`, no import) and moves on; object identity for
those aliases is still checked on POSIX hosts, every portable alias keeps identity
coverage everywhere. Keyed on the exception's module name rather than a hardcoded alias
list so a future POSIX-only sibling does not need a test edit.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
`scan_plugin` stored `str(p.relative_to(plugin_dir))`, which is `sub\m.py` on
native Windows and `sub/m.py` elsewhere. The string is what compat notices print,
what `plugins_cmd` renders as `file:line`, and what the tests pin, so the same
plugin produced a different report per OS and
`test_scan_plugin_walks_dir_and_skips_tests` failed on Windows (#112576).
`.as_posix()` makes the recorded path stable and portable; nothing consumes the
native form.
delete/remove_file/unknown actions have no destination key, so the hint
rendered as "move that text to ." appended to an unrelated error. Return
'' for them. For patch, name old_string/new_string first so a failed
targeted edit is not steered toward the full-file rewrite the patch
error itself warns against.
When a batch op puts SKILL.md text in another action's slot (write_file's
file_content or patch's new_string on a create; file_content on a patch
with no old_string; content on a write_file), the validation error only
said the right key was "required". A 27B local model then replayed the
identical payload until the tool-loop guardrail tripped — the error never
told it where its text had landed.
Append one sentence at the single dispatch chokepoint naming every text
slot the op carries that the action never reads and where to move it,
driven by a key -> owning-action table (no per-action if-chain, no schema
growth). Batch atomicity and the plain missing-key wording are unchanged;
compliant calls are untouched.
Co-authored-by: astraltrekkin <164521089+rainbowgits@users.noreply.github.com>
Co-authored-by: micah-c02 <micah-c02@users.noreply.github.com>
The spec hardcoded a /tmp rig directory for its four screenshots and left
console.log tracing in place. Screenshots now go through
testInfo.outputPath() like chat.spec.ts so they land in test-results/ on
any machine and CI; the debug logging and its helper are gone.
registerRoutinesPane undismissed the pane on EVERY registration, and the
pane re-registers whenever a bot chat regains the workspace inside one Bots
session (open a group room, come back). So a user's ✕ was silently undone
without ever leaving Bot Mode. The remembered Close is now dropped only on
Bots-tab ENTRY — the sidebar-visible transition arms a one-shot flag that
the next registration spends — which is the case #102224 actually asked
for. The no-paneVisibility fallback keeps its single always-undismissed
registration.
Invariant test: boot into a bot chat → one undismiss; ✕ + group room and
back → still one; leave Bots and re-enter → two. Red on the previous head
(the round trip fired a second undismiss).
A job row in the Scheduled jobs (Routines) pane is a grid item, and grid items
default to `min-width: auto`, so a long nowrap title made the row as wide as
its text (~530px inside the 250px pane). The pane's overflow clipped the
enable/disable Switch and the delete control clean off the right edge and the
title was hard-cut with no ellipsis (#91623); the next-run label on the
metadata line lost its tail the same way (#89534). `min-w-0` on the card and
on both lines lets the row shrink to the pane, the title ellipsizes, and the
schedule pill / next-run label wrap onto two lines instead of being cut
mid-word.
Closing the pane with its ✕ remembered the dismissal in
`hermes.desktop.dismissedPanes.v1`, and nothing ever showed the pane again —
it only exists by being registered while Bot Mode is on screen, so the pane
was gone until a full layout reset (#102224). The plugin now drops that
remembered Close each time it registers the pane (new `host.undismissPane`
door — un-dismiss + adopt only, unlike `revealPane` it never fronts the pane
or un-collapses its zone), so re-entering the Bots tab brings the pane back
as the collapsed right-edge tab it first arrived as.
Live-tested in the real Electron app with e2e/bot-routines-pane-narrow.spec.ts
(fails on origin/main, passes here); the same run confirms the row's Switch
toggles the job in place (#95031 already fixed by 253b9d78c).
Co-authored-by: MicroWearld <MicroWearld@users.noreply.github.com>
Co-authored-by: worlldz <worlldz@users.noreply.github.com>
Co-authored-by: 686f6c61 <686f6c61@users.noreply.github.com>
The English environment-variables/configuration pages now state that
execute_code children keep the OS zone on Windows; the zh-Hans mirrors
still described the bare IANA override and contradicted it. The POSIX
control test duplicated test_code_execution.py::test_timezone_injected_when_set
(green against origin/main's code), so it guarded nothing new.
tests/tools/test_code_execution.py::test_timezone_injected_when_set asserted
child_env["TZ"] == "America/New_York" unconditionally, which encoded the
Windows defect as intended behaviour; on win32 it now asserts TZ is absent and
keeps the POSIX assertion.
Add a windows_only live-child test to test_code_execution_windows_env.py: with
timezone: configured, a real Windows child's time.timezone and astimezone()
offset must equal the test process's own OS-zone view — the reporter's exact
witness (time.timezone == 0, +01:00 instead of -07:00, tzname still correct).
Docs (environment-variables.md, configuration.md): HERMES_TIMEZONE / timezone:
reaches execute_code children as TZ on Linux/macOS only; Windows children keep
the OS zone because the Windows C runtime parses only POSIX-form TZ strings.
Sibling env builders swept: the remote-backend prefix in
code_execution_tool.py (TZ=... before `python3 script.py`) targets POSIX
backends (Docker/SSH/Modal) and is unchanged; no other spawner forwards
HERMES_TIMEZONE into TZ. HERMES_TIMEZONE itself stays stripped from the child
(adf23550f5: under the multiplexed gateway it holds only the default
profile's value).
Refs #112233
The Windows quarantine guards now resolve project_venv_dir(PROJECT_ROOT)
at call time, so the three tests that hard-coded PROJECT_ROOT/venv as the
venv-side interpreter went red on any host whose checkout uses .venv
(CI included: test_venv_launcher_ancestors_returns_venv_side_parent and
test_pause_kill_set_covers_venv_guard_abort_set). Mirror the guards'
lookup (and their resolved-path prefix match) in one helper instead.
The updater and Desktop preflight now accept the uv-default .venv layout;
say so where the Windows guards are documented, including which directory
wins when both exist.
`npm run check:lint` from apps/desktop reports the cherry-picked import order
as a perfectionist/sort-imports error and the new local in
`venvHermesShimPath` as a padding warning; both introduced by the pick.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Follow-up to the cherry-picked #112966 so the uv-default `.venv` layout is
supported end to end, not only at the lookup sites:
- `_ZIP_PRESERVED_TOP_LEVEL` gains `.venv`. The dirty-tree guard runs
`git status --ignored=matching`, so a gitignored `.venv/` surfaced as
`!! .venv/` and refused every ZIP fallback on such installs ("the working
tree has uncommitted changes or untracked files") — the live runtime was
being treated as user data the overlay would destroy.
- `_repair_venv_on_current_checkout` recreates the venv at the resolved
directory instead of a literal `venv`, so a broken `.venv` is rebuilt in
place rather than growing a second environment that `project_venv_dir()`
then prefers while `bin/hermes.cmd` still launches the old one.
- `_refuse_update_if_venv_foreign_owned` scans the resolved venv (the only
remaining `PROJECT_ROOT / "venv"` literal on the update path).
- windows.ps1 names the actual shim path in the lock-timeout message.
- Tests: extend the real-git ZIP guard test with the `.venv` case (red
before this commit); the holder-guard test now uses a kernel-runner child
whose cmdline lacks `hermes_cli.main`, so only the venv-prefix arm can match
it (red on origin/main); drop the `process.platform`-override vitest case,
which exercised the same resolver as the `.venv` case with a different
directory string.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
The GroupRow preview and the group transcript spelled the reader as the
English literal 'You' under every locale, so a zh roster read "You: …"
beside otherwise Chinese chrome. Add `group.you` to all four Bot Mode
bundles and use it at the three RENDER sites (bot-row preview,
group-chat-view speaker display and clicked label).
The persisted author marker written into the room log
(group-chat-parts.tsx, group-rounds.ts) stays the English sentinel:
group-activity.ts compares it by value, so translating it in place would
corrupt existing logs and break the comparison. The i18n.ts header now
documents that split.
The roster's right-click menu on a bot row still spelled `Pin to top` /
`Unpin`, `Hide` / `Unhide`, `Groups: …` / `Manage groups…`, the pin/hide
toasts, the metadata-load error toasts, the `This device` gateway label
and the attention-badge tooltips as English literals, so a zh/ja reader
saw English items inside an otherwise translated menu (#91667, residual of
#88798 / #91336 after the bundle landed in a7db531a2a).
Route every one of them through the plugin bundle (`b.bot.*`, all four
locales) and turn `BOT_ATTENTION_HINTS` into `botAttentionHint(reason)`,
which reads `botsText()` at render so the tooltip follows the active
locale. No behaviour change; strings only.
Tests: bot-row.test.tsx renders the menu under `zh` and asserts the
catalog strings appear and no English literal survives (red on base).
e2e/bot-mode-roster-localized.spec.ts drives the real Electron app with
`display.language: zh`: right-click menu + a fresh group row (red on base,
green here).
The group row's empty-room preview and its aria-label spelled
`${members.length} bots` and `${available} of ${total} available` as
literals, so a reader on ja, zh or zh-hant saw English in a row whose
neighbours were translated. The bundle already carries `group.memberCount`
in all four locales (a7db531a2a moved it there) and `group-chat-view.tsx`
renders through it; the row's rebuild (5afa487e93) reintroduced the literal
next to the bundle accessor it already holds.
Route both through the bundle: the preview and the aria-label's count use
`group.memberCount`, and a new `group.availableCount(available, total)`
carries the availability sentence in all four locales, for the aria-label
and the tooltip that share it. `'You'` in the last-message preview stays a
literal on purpose: it is the persisted author sentinel the bundle header
describes.
Tests: `i18n-test-helper` gains `translateBotsIn(locale)`, since a case that
has to tell a catalog string from an English literal cannot do it under
`en`. The new group-row case renders the same row under `en` and `ja` and
reads the preview and the accessible name; it fails on the previous source.
Review: the dialog test asserted a styling detail (`whitespace-pre-line`)
on top of the behavioural invariant; the textContent check already proves
the guard's paragraphs survive, so the class match only couples the test
to CSS. Also add the one missing sentence to the Desktop user guide so
"Switch anyway" / "Keep current model" is discoverable from the docs.
When the gateway answers `confirm_required` (large cached context,
expensive model, data-training tier), the Desktop asked for confirmation in
a warning toast with a single "Confirm" action and an ✕ — no button meant
"no", Enter/Esc did nothing, the four-item toast stack could evict the
pending question, and the guard's `\n`/`\n\n` paragraphs collapsed into one
run-on line. Answering after the session had moved on was a silent no-op.
`surfaceModelSwitchConfirm` (THE shared applier for both surfaces — the
composer picker via `config.set` and the Bots editor via
`profiles.configure`) now asks through `confirm()` from `@/store/confirm`,
i.e. the shell-mounted ConfirmDialog that apps/desktop/DESIGN.md names as
the only way to ask "are you sure": destructive "Switch anyway" vs "Keep
current model", Enter confirms, Esc/backdrop/✕ decline, the dialog owns
focus. Declining is free — nothing was applied before the answer. A stale
confirmed answer now toasts "Selection changed — the model switch was not
applied" instead of doing nothing. ConfirmDialog renders its description
`whitespace-pre-line` so backend-composed paragraphs survive. The applier
resolves `true`/`false` instead of returning a notification id; callers
fire-and-forget it. New i18n keys land in all six locales.
Slim redo of #112463 by @DavidMetcalfe (design: route through confirm(),
labels, stale notice). #112461 by @KoNit-K was the earliest filer (Cancel
action on the toast) and is superseded by the dialog.
Fixes#112458
Co-authored-by: DavidMetcalfe <80915+DavidMetcalfe@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The heading-sequence and code-block-count tests pinned the English page's exact shape, so any English-only edit of bot-mode.md (e.g. #113411 adds a heading) turned main red until the translation was re-synced. Replace them with two invariants of the translation itself: the page exists and routes like the English page (front matter id/slug agree), and every zh-Hans code block is a byte-accurate copy of a block that still exists in the English page, so only a genuinely stale command/config in the translation fails.
The zh-Hans build resolves ./x.md links against the locale tree and warns
when the target page is untranslated (desktop, multi-connection-desktop,
features/pets). Extensionless URL links resolve to the English fallback
page without a warning, matching the convention in the other zh-Hans
user-guide pages.
The salvaged translation (#94377) mirrored the English guide as of its
base commit; the English page has since gained 147 lines. Translate the
drift so zh-Hans readers see the same guide:
- Bots pane: row-click/tab-caption behaviour, the rewritten "Active now"
filter, and the archive-retires-a-Bot-Chat paragraph
(`sessions.auto_archive`).
- New section "Organize bots into sections" (`ui_meta`).
- Avatars: the geometric-face live-turn pose wording.
- Groups: queued-message lead paragraph, editable room identity
(picture + Group settings), room ordering (Move up/down), and the new
room bullets (one visible conversation, `@hermes` handoff, needs-you
prompts, rooms keep running via `groups.capabilities`, plugin hook
`on_room_member_activity`); member sessions no longer named
`Group: <name>`.
- Bot-to-bot messaging: live delivery ownership paragraph
(`runtime/bot_live_delivery/`), the "Staying silent" bullet, the
`SESSION_NOT_OWNED`/`target_busy` and quiet-CLI-turn paragraphs, the
side-by-side delivery sentence (`bot_mode.envelope_ttl_seconds`).
- `hermes peer run/status/stop` commands and paragraph, the NAT note.
- New sections "Transferring hosted room authority" (JSON-RPC payloads
kept byte-identical) and "Warm Bot Backends".
- Turning it off: Capabilities → Plugins → Bots Desktop switch.
Tests: drop the file-exists and internal-link checks (the Docusaurus
build already fails on broken links); keep heading parity and
code-block parity as the two invariants.
website/docs/user-guide/bot-mode.md had no zh-Hans mirror, so readers on
that locale silently fell back to the English page after #89250 wired up
Bot Mode's docs entry points. Translate the guide in full, preserving
heading order, commands, config keys, paths, and internal links, and add
a structural-parity test so the two stay in sync.
Fixes#94366
`botMetaKey` and `isDefaultBot` reached a row's owner through the strict
`botConnectionRoute`, which throws "Bot X has no connection owner" for a
source-scoped row whose connection is gone. Both helpers run while painting
(`bot-row.tsx` context menu, `hidden-bots.ts` default-bot selection over the
whole roster), and `botMetaKey`'s own docstring promises the nullable
`botSelectionKey` contract, so an orphaned row took the subtree down instead
of reading as a settled "no route" state.
Branch on `resolveBotConnectionRoute(...).status` the way `botRosterMeta`
already does: an `owner_removed` row keys by its degraded roster key and is
never "default" by route. Real dispatch keeps the strict wrapper.
Part of #110002 (the Bot Screen half lives in #108914).
The INFO-once flag `_compaction_prompt_drift_logged` lived on the AIAgent and
was never cleared, so on a CLI agent that /new, /resume or /branch-rebinds
sessions the line was INFO once per agent lifetime, not once per session as
the commit and PR text claim. reset_session_state is the per-session reset
point for agent-held counters (`_user_turn_count` lives there); clear the
flag alongside it.
After a no-progress stall the ladder's first rung (60s) was shorter than the
default idle stall window (120s), so the next oversized automatic turn
re-entered the same silent summary route about a minute after burning the
whole window (#112420: "made no progress ... continuing without
compression" followed by "context compression started" 62s later).
record_timeout_failure now floors every timeout-class cooldown at the
configured idle window; a window below the rung leaves the ladder untouched.
The "Compaction rebuilt a drifted system prompt" line fired on every compact
of a long session (19x/day observed). The rebuild stays mandatory; only the
first drift per session is INFO, later ones DEBUG.
Both hunks re-applied from PR #112504 (@JoaoMarcos44); its no-LLM prune on
stall is a product decision and was not ported.
Part of #112420
Fold the duplicated fallback worker into the existing snapshot worker (one
parameter, functools.partial) and replace the patched-compress_context test
with an invariant test that drives the real AIAgent facade, timeout wrap and
compress_context: the primary summary route stalls silently until the host
cancels its fence, the cancelled worker persists its stall_interrupted
backoff, and the pinned fallback_chain retry must still compress.
Why: the contributor test only asserted the kwarg handoff with compress_context
patched out, so it could not catch the real gate/cooldown race that #112387
reports (retry returned the caller's own list in 15 ms).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
A declared `security.fake_ip_ranges` block is trusted like `allow_private_urls`
for whatever it names, and `_resolve_fake_ip_ranges` accepted the entries
verbatim. Declaring 10.0.0.0/8, 127.0.0.0/8, 100.64.0.0/10, fc00::/7, or a
catch-all 0.0.0.0/0 / ::/0 therefore made loopback, RFC 1918, CGNAT and ULA
hosts dialable at pre-flight and at TCP connect, while the docs promised those
classes "stay blocked". Only the link-local/metadata floor actually held.
A local proxy owns none of those classes, so an entry overlapping one of them
(or the unspecified address, which is how the catch-alls are caught) is now
dropped with a warning instead of widening the guard; the other declared
entries keep working. The docs sentence now says so.
The browser hybrid-routing oracle `_url_is_private` classified the declared
sentinel as private via `ipaddress.is_private` without consulting the
declaration, so with a cloud browser provider and
`browser.auto_local_for_private_urls` (default on) every URL on a fake-ip
host was routed to the local sidecar. The sentinel is not private for routing
either: the name is public and the cloud browser resolves it itself.
Follow-up to the cherry-picked #112536 (@AYin-Z):
- The range cache now bypasses the process-global cache when a profile
override is active, exactly like _global_allow_private_urls — a multiplex
gateway serving several profiles must not apply the first profile's
declaration to later ones (A->B->A probe: True/False/True).
- _reset_allow_private_cache() resets both caches; the separate test-only
_reset_fake_ip_cache() is gone.
- _is_declared_fake_ip() is a plain predicate folded into the one
_resolved_ip_block_reason() condition; the metadata floor is still tested
first, so a declared block can never cover 169.254.0.0/16 or the metadata
IPs. Dropped the comma-split string parsing (a single string is still
accepted as one entry) and the debug log line.
- Tests trimmed from five to two invariants: declared sentinel dialable at
pre-flight and connect time; the declaration excuses only the declared
block (RFC 1918 + metadata still blocked, sentinel blocked again once the
declaration is removed).
A TUN proxy in fake-ip mode (Mihomo/Clash `fake-ip`, Surge enhanced) answers DNS with an
address from its own block — `198.18.0.0/15` by default — for every name outside its filter.
The SSRF guard resolves names with the local resolver and reads that sentinel as a private
destination, so on such a host every fetch fails before the request leaves the machine:
`web_extract`, gateway media downloads (`is_safe_url` has ~20 call sites, including the
Feishu/WeCom/Telegram/Slack/Discord attachment paths) and the browser relay all report
"URL targets a private or internal network address".
The existing hostname allowlist (`_TRUSTED_PRIVATE_IP_HOSTS`, the QQ multimedia case) does not
generalise to this: the sentinel is a property of the host's resolver, not of one name.
Fix: `security.fake_ip_ranges` declares the CIDR blocks the local proxy owns. Answers inside a
declared block are dialable even with private-IP blocking on. Empty by default, so no existing
host changes behaviour. Loopback, RFC 1918, link-local, CGNAT and the cloud-metadata floor are
untouched. Pre-flight and connect-time checks both go through `_resolved_ip_block_reason`, so one
exemption covers the class instead of the ~20 call sites.
Tests: `TestDeclaredFakeIpSentinelRanges` in tests/tools/test_url_safety.py — red on the unfixed
module (declared sentinel rejected, "Blocked request to private/internal address during connect:
example.com -> 198.18.0.55"), green after. The pre-existing benchmark/QQ hostname tests are
unchanged and still pass (66 passed).
Docs: website/docs/user-guide/security.md + zh-Hans translation. The key is registered in
hermes_cli/config_defaults.py so `hermes config set security.fake_ip_ranges` is not flagged as
unknown.
The agent-declared failure evidence was routed through
_summarize_cron_failure_for_delivery, whose substring heuristics re-diagnose
natural-language text: a child that "timed out after 30 minutes waiting on the
database" was delivered as "the AI model service did not respond in time … add
one with `hermes fallback add`", and a child that saw a vendor 401 became "the
AI model service rejected the sign-in. Sign in again with /login". The operator
got a wrong diagnosis and a wrong remediation; the real evidence only survived
in the saved run output.
The marker is now carried as a flag (_RunDelivery.agent_declared ->
_compose_run_delivery(agent_declared=...)) and composed like blocked_config
already is for the same reason: `⚠️ Cron '<name>' failed: <evidence>` plus the
run-log / re-run hint. Incident ledger, acked-suppression and the streak nudge
are unchanged.
Review follow-up on #113155.
Keep the positive (first-line marker fails the run) and negative (quoted marker stays
successful) invariants from #112431; drop the no_agent control and the prompt-text
presence check, which only re-read the constant and the hint string.
Add the user-facing description of the marker next to [SILENT] in the cron docs so the
control token is discoverable without reading the prompt hint.
Co-authored-by: hanamizuki <claude@hanamizuki.tw>
A bot created moments before the room runs its intro turn in the background;
when it lands the roster fronts that bot's chat tab and the reply text found
by getByText sits outside any room body (null closest()). Scope the locator
to a room body and re-select the room inside the wait.