Commit Graph

2619 Commits

Author SHA1 Message Date
kshitijk4poor
f0df5e05b1 fix(disk-cleanup): harden and gate the per-candidate git tracked check
_git_tracks hand-rolled `git ls-files` without windows_hide_flags, an
isolated git env or stdin=DEVNULL. It runs from the synchronous
post_tool_call hook, so on a windowless Windows host each test_*/tmp_*
candidate could flash a console (#54220/#56747 class), and an inherited
GIT_DIR/GIT_WORK_TREE would point it at the wrong repo. Reuse
hermes_cli.source_check._git_ok, which already does all three and returns
False on any failure. The timeout drops to 5s (ls-files needs no more).

git reads the argument as a pathspec, so an untracked `test_[1].py` or
`tmp_*` glob-matched a tracked sibling and was never cleaned. Prefix it
with `:(literal)`.

Only spawn git when a .git exists at or above HERMES_HOME (HERMES_HOME is
a checkout, or sits inside a dotfiles repo). Without one, no repo can
track the file, so a stock install now does a few stats and skips the
process spawn on every qualifying tool call.

Drop the docstring paragraph that repeated _git_tracks' rationale.
2026-09-27 01:36:34 +05:30
kshitijk4poor
50ca677d2b fix(disk-cleanup): ask git per candidate so enclosing repos are covered too
The tracked-file guard only ran when HERMES_HOME/.git existed, so a
HERMES_HOME nested in an enclosing repo (a ~/.git dotfiles repo tracking
~/.hermes/scripts/test_x.py) still had that committed file deleted by
quick() - the same bug class the guard was added for.

guess_category only reaches this check for test_*/tmp_* names, so a
per-candidate 'git -C <parent> ls-files --error-unmatch -- <name>' is
cheap. It finds whichever repo encloses the file, and it replaces the two
lru_caches plus the index-stat signature: with no cache there is nothing
that can outlive the index, and ls-files output no longer needs decoding.
An exact tracked check stays safe above HERMES_HOME, unlike the bare .git
probe that was dropped for being too broad.

Co-authored-by: David Crandall <david@convergentdesign.dev>
2026-09-27 01:36:34 +05:30
kshitijk4poor
cc92bea9e3 fix(disk-cleanup): decode ls-files output safely instead of swallowing every error
_git_tracks wrapped the lookup in a bare except Exception. The only real
failure it hid was UnicodeDecodeError from text=True on non-UTF-8 paths in
the ls-files output, which escapes the (OSError, SubprocessError) catch.
Decode with surrogateescape (matching how Python decodes filesystem paths)
and drop the catch-all so real bugs surface.

Co-authored-by: David Crandall <david@convergentdesign.dev>
2026-09-27 01:36:34 +05:30
David Crandall
9d6609a14b fix(disk-cleanup): the ls-files cache must not outlive the git index
`_git_tracked_index()` cached one `git ls-files` set per process, so a
long-lived process could keep classifying a file as untracked after git had
taken ownership of it. `guess_category()` runs on every post-tool-call, so the
gateway can cache the index while `test_scratch.py` is still untracked; an
auto-snapshot then stages and commits it, `quick()` re-validates the stored
"test" entry through `guess_category()`, the stale cache still reports the file
as untracked, and `_delete_item()` unlinks a file git now owns (`git status`
shows `D test_scratch.py`).

Key the cached set on the git index's stat signature as well as the home, so a
staged, committed or unstaged change is a new cache key and no invalidation
hook is needed anywhere. The index path comes from `git rev-parse --git-path
index`, cached per home — so the per-call cost stays a single `os.stat`, which
also covers a linked worktree (where `.git` is a pointer file) and an explicit
`GIT_INDEX_FILE`. Fail-open behaviour is unchanged: no index and no git mean the
empty set, and the guard never raises.

Regression test: classify an untracked `test_scratch.py`, then `git add` and
`git commit` it in the same process with no `cache_clear()`, and assert
`quick()` leaves it on disk while an untracked control beside it is still
cleaned.

(cherry picked from commit a3271750608ca4e02e5205311a09b906e16d33d5)
2026-09-27 01:36:34 +05:30
David Crandall
0e799da223 fix(disk-cleanup): a git-tracked file is never disposable, even when HERMES_HOME is the checkout
The `_inside_git_worktree()` guard sliced its parent chain at HERMES_HOME, so
`.git` entries strictly below the home were the only ones consulted. That misses
the case where HERMES_HOME IS a git checkout: on this install `~/.hermes` is the
userfiles repo, so `~/.hermes/scripts/` and any top-level `tmp_*` file sit inside
a worktree yet resolve to no `.git` below the home.

Observed live: the bundled disk-cleanup plugin classified two COMMITTED
regression tests in `~/.hermes/scripts/` as disposable session scratch and
unlinked them — `scripts/test_analyze_upstream_opportunities.py` and
`scripts/test_customization_protocol_v2.py`; the deletion was then committed by
the next auto-snapshot (ce8dc99), so the suite that would have caught a scanner
regression was silently gone. 13 tracked `test_*` files under `scripts/` plus 17
tracked top-level `tmp_*` files were in the same blast radius.

Fix: when HERMES_HOME itself carries `.git`, ask git whether it TRACKS the exact
path (`git ls-files`, one cached subprocess per process keyed on the home) rather
than treating the whole home as protected — a scratch file merely living beside
tracked ones must still be cleanable, or cleanup would be disabled entirely.

Tests: two regression tests — a git-tracked `test_*` file is never classified
disposable and a stale pre-fix tracked.json entry is dropped by quick()'s
re-validation instead of deleted; plus the control that an UNTRACKED `test_*`
file in the same repo is still cleaned. Existing suite 30 -> 32 passed.

(cherry picked from commit c5555940108286c4c1070e84d6d694ad9a35bb76)
(cherry picked from commit c7c39e60d4dc467bf9681c3dcb5781ac96557724)
2026-09-27 01:36:34 +05:30
kshitijk4poor
da4ab40ea9 refactor(telegram): gate chat-id pickers once in the callback dispatcher
The model-picker fix added a second copy of the choice picker's auth
check, and each copy rebuilt the callback ctx that _handle_callback_query
had already computed as `cb`. The chat-id picker loop is the only route
into both handlers. Gate there once with the existing cb, so a picker
added to that loop later is covered without anyone having to copy the
check again. Fail-closed behaviour is unchanged: _callback_authorized
returns True only for an authorized tapper and answers the denial text
otherwise. The back-button test calls the handler directly and no longer
needs its auth stub.

Co-authored-by: Dusk1e <yusufalweshdemir@gmail.com>
2026-09-27 01:34:15 +05:30
kshitijk4poor
a2ce2da095 fix(discord): gate model picker cancel on callback auth
Every other ModelPickerView callback (provider/model select, expensive
confirm, back) runs through the shared _HermesView._gate auth check, but
_on_cancel did not. A non-allowlisted member could tap Cancel on the
owner's picker, marking it resolved and clearing the view. No model
switch was possible, so impact is low, but the gate should be uniform.
Reuse the same _gate call the sibling handlers use.
2026-09-27 01:34:15 +05:30
Dusk1e
999b19b6ed fix(telegram): re-check auth for model picker callbacks
Model picker taps (mp:/mpg:/mpv:/mm:/mc:/mb/mx/mg:) never went through the
callback allowlist, so a group member who is not allowlisted could tap the
owner's /model picker and switch the owner's session or config model. Gate
the handler on _callback_authorized like the choice-picker, approval and
clarify buttons, before any picker state is read.

Ported from #33859 (gateway/platforms/telegram.py) to the plugin adapter.

(cherry picked from commit b002be8feacd09fde0ee9752de7b396d0aac69ba)
2026-09-27 01:34:15 +05:30
kshitijk4poor
e174f36ae2 refactor(cache): match api.x.ai by hostname and share the Grok slug prefixes
A substring check on the raw base URL also matches proxies such as
https://proxy.test/api.x.ai/v1, so is_slot_keyed_cache_route now compares
utils.base_url_hostname() exactly, like chat_completion_helpers and
codex_responses_adapter already do.

The aggregator Grok prefixes were spelled out in both prompt_cache_scope and
the OpenRouter profile; one GROK_AGGREGATOR_MODEL_PREFIXES constant keeps the
fork-scope check and the x-grok-conv-id header on the same model set.
2026-09-27 00:38:34 +05:30
Kyzcreig
d159b6e0c3 fix(cache): cache-parity forks get a derived cache scope on xAI
#109964 made same-model cache-parity forks (background review, /btw) inherit
the parent's resolved cache scope. That is the right call for
content-addressed caches (Anthropic, DeepSeek, Gemini) and for OpenAI's
routing-only prompt_cache_key. xAI is different: x-grok-conv-id /
prompt_cache_key select ONE server-side conversation slot, so a fork of
similar size that diverges early evicts the parent's slot and the parent's
next call reads cold.

Measured on grok-4.7 (xai-oauth Responses), ~160k context, fork of similar
size diverging right after the system prompt, parent's next call after the
fork:
  shared scope (current):   1,152 / 162,239 read (2/2 runs)
  derived scope (this PR): 162,176 / 162,239 read (2/2 runs)
  no fork (control):       177,536 / 177,639 read (2/2 runs)

build_cache_parity_fork tags the fork (_prompt_cache_fork_tag). The
resolver derives "<scope>::<tag>" for tagged agents on slot-keyed routes
only (xai, xai-oauth, api.x.ai, OpenRouter x-ai/grok-*). The inherited
scope is kept everywhere else. The tag is applied outside the memo, so a
mid-run provider fallback re-evaluates. OpenRouter's Grok x-grok-conv-id
now honours a fork scope over the ambient affinity/conversation scope
(chat_completions threads cache_scope_id to the profile). session_id and
transcript identity are unchanged.

(cherry picked from commit 40fbb39d978648c76940cfb7835f77a42d3719b0)
2026-09-27 00:38:34 +05:30
kshitijk4poor
40cb28e60c fix(agent,slack): close remaining compressor elision gaps from review
- salvage summary cap and _bound_oversized_record still composed bare
  truncation idioms in model-visible text; route them through elide /
  elide_middle so a copied marker is guard-visible.
- the active-task line repr()'d the elided text, escaping the marker's
  apostrophe when the user text held both quote kinds and hiding it from
  the guard; elide after repr instead (text within the cap stays whole).
- a leftover budget smaller than the marker produced a content-free,
  over-budget marker line in _build_verbatim_user_section and the Slack
  nested-attachment path; skip the item instead.
- _build_verbatim_user_section elided twice, reporting the wrong total;
  one elide at min(cap, remaining).
- drop redundant len() pre-checks before elide() and name verification
  stop's repeated 1200.

Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
2026-09-26 23:49:25 +05:30
kshitijk4poor
f7122daaab fix(slack,trajectory): mint the compression marker for agent-facing truncations
The Slack Block Kit payload dump and the nested-attachment text budget
are both fed to the agent, and trajectory_compressor's summarizer input
becomes training data; all three still used the imitable bare
"... [truncated]" idiom. Route them through elide()/elide_middle() with
module-level imports and extend the no-idiom invariant to scan
plugins/platforms/slack and trajectory_compressor.py.

Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
2026-09-26 23:49:25 +05:30
kshitijk4poor
d24aadfdd1 refactor(copilot): one GitHub effort clamp for the profile and main agent
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.

The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
2026-09-26 22:21:47 +05:30
kshitijk4poor
b17e037e35 fix(slack): a follow-up in a flat reply_in_thread: false channel keeps the silence fallback
With reply_in_thread: false the whole channel is one session and the bot
answers top-level, so an unmentioned top-level message there is a
follow-up in a conversation the bot is part of, like a thread reply.
reply_expected is now False for a free-channel message only when it starts
its own session (a new top-level thread), else None. The bot-id set is
built inside _slack_reply_expected, as _channel_gate_allows does.
2026-09-26 07:21:49 +05:30
kshitijk4poor
6c566fcd6d fix(slack): thread follow-ups and reaction triggers keep the silence-marker fallback
The Slack rule marked every admitted message that was not a DM, a mention
or a command as not addressed, so a plain "done?" in a thread the bot is part
of, or a reaction trigger, could end on a bare silence marker and vanish,
the case #111624 fixed (#110952).

reply_expected is now False only for a message that opens by @mentioning
someone else, or a top-level message a free-response channel admitted
without a mention. Reaction triggers and pipe-form self mentions count as
addressed; other thread replies are None (visible fallback). The
free-channel predicate moves into _slack_is_free_channel so the gate and
the rule read the same one. The test drives the real _handle_slack_message.
Docs describe the rule in its own note instead of the
ignore_other_user_mentions tip, and the messaging index documents the
human-turn fallback.
2026-09-26 07:21:49 +05:30
Victor Kyriazakos
b08bb5afb8 fix(gateway): a bare silence marker on a turn not addressed to the bot stays silent
Since 5ea8fb2b78 (#111624, for #110952) the gateway rejects a bare silence
marker on any human turn and delivers "The model returned only a silence
marker for a message that needed a reply" instead. That protects a human
who asked this bot something and got nothing back. It also fires on every
human message the adapter admitted without the bot being addressed at all:
a free-response channel, a thread follow-up under
`thread_require_mention: false`, or a message @-mentioning another person
or bot with `ignore_other_user_mentions: false`. A bot whose SOUL declines
peer-addressed turns with a deliberate marker now posts that notice on
every such message. A fleet running several bots in shared Slack threads
reported it as spam on v2026.9.21. #37940 established that intentional
silence must not be re-inflated. Both contracts hold once the turn knows
whether a reply was expected.

`MessageEvent.reply_expected` (True, False, None) is set by the adapter
where the message is admitted. Slack (`slack_reply_expected`): a 1:1 DM,
an @mention of this bot or a command is True, anything else it admits is
False. Other adapters leave None, which keeps today's behaviour, so nothing
changes for them until they are ported. `response_filters.silence_allowed`
holds the one rule (machinery turn, or reply not expected) and both call
sites use it: the live turn in `run_turn._hmwa_shape_agent_response` and
the crash-recovery redelivery from #120377 (1136f135dd), which reads the
flag back from the persisted turn metadata. The suppressed case logs one
DEBUG line naming platform and chat.

Operator workaround until this lands: `platforms.slack.extra.
ignore_other_user_mentions: true` drops peer-addressed messages before a
turn exists.

(cherry picked from commit 094439776ab898cccde303a1c2c911c8ab5bfb75)
2026-09-26 07:21:49 +05:30
Hermes Agent
94f3dbec9b fix(agent): close one-shot AIAgents on every exit path
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:

- batch_runner._process_single_prompt: one agent per prompt, N prompts
  per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
  gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.

Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.

Fixes #50197
2026-09-25 12:49:56 -05:00
brooklyn!
bfb1d952c0 fix(kanban): start the events stream at the board tail
Desktop opened /events with no since, and a missing cursor was read as
0, so every open replayed task_events history. Seed the socket from the
snapshot or this connection's last frame, and start a cursorless stream
at MAX(id). An explicit since still replays from there.

Fixes #81537
2026-09-25 11:26:40 -05:00
kshitijk4poor
e326520d50 refactor(codex): one credential/route authority for aux + image paths
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
  fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
  directly when the pool yields no token (no second uncached pool load,
  no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
  is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
  helper's error fallback reads the profile-scoped override, not the raw
  process env.
- model setup flow: the confirm guards get the resolved Codex base, not
  the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
2026-09-25 21:27:06 +05:30
kshitijk4poor
e87f673faa fix(codex): send catalog/image credentials only to their own route
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.

- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
  credential actually routes to (runtime_provider._pool_entry_mode_and_url:
  env > model.base_url while the row is canonical > row URL) instead of the
  ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
  resolved with the token from hermes_cli/models.py, the CLI default-model
  swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
  one pool selection; the image plugin, _build_codex_client and the raw
  Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
  chatgpt.com; a gateway key may probe its own gateway's /models.

Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.

Addresses @andrexibiza's review on #121508.
2026-09-25 21:27:06 +05:30
funky-xamarin
894db6a35b fix(image-gen): honor scoped Codex base URL for native images
(cherry picked from commit f09df8fffa7277bebe909abee5841fa3d0379237)
2026-09-25 21:27:06 +05:30
kshitijk4poor
fdec926ef5 fix(slack): replace a native stream in place only for a restyled final, share one commit path
- A mid-turn notify reply (/status, /approve, clarify answer) shares the
  stream's thread key; it no longer seals and overwrites the half-streamed
  answer. In-place replacement now requires the final to match the stream
  after normalizing mrkdwn markers and whitespace; anything else posts fresh
  and leaves the stream open.
- One _commit_stream helper for both seal-then-commit paths, so the rewrite
  path also falls back to chat.update when stopStream fails.
- A stream reopened after a server-side seal is seeded with only the text
  past the sealed message (tracked as 'base'), not the whole segment.
- Streams older than 15 min are sealed and dropped on the next start.
- _stream_key reuses _workspace_thread_key/scope_id_for_chat; the stream
  dict no longer duplicates chat/team ids.
2026-09-25 14:28:14 +05:30
EloquentBrush0x
d5a2d1f069 fix(slack): reopen a native draft stream when Slack seals it mid-turn, same as the task-card twin
5648f81431 fixed this exact server-side seal (Slack closes a native stream
after a few minutes of a long turn, live-observed at ~5m20s; the lifetime
is not documented) for the native task-card stream: on
message_not_in_streaming_state from appendStream, drop the dead ts and
start a fresh stream, seeded with the full current content so nothing is
lost.

send_draft — the plain-text native streaming path used when task cards
are not enabled — hits the identical seal but never got the fix: its
generic except block only recognizes the feature-gate markers
(not_allowed, missing_scope, ...) and otherwise just logs debug and
returns failure. gateway/stream_consumer_transport.py's
_send_draft_frame() docstring is explicit that "any failure permanently
disables drafts for this run" — so a long turn streaming as plain text
degrades to the edit-based fallback for its remainder exactly the way
the task-card bug did before 5648f81431.

Mirror the task-card fix: on message_not_in_streaming_state from
chat.appendStream, drop the dead ts and _start_stream() a fresh one
seeded with the full accumulated text (not just the delta), so the next
frame's delta still resumes correctly. One reopen per frame; a second
rejection propagates as a real failure, matching the twin's behavior.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 2b4ff4e23bdf684d2fde1a9512f11dcd7ef99c42)
2026-09-25 14:28:14 +05:30
kshitijk4poor
d5bddd7f00 fix(slack): replace a rewritten native-stream final in place instead of re-posting
A mrkdwn-rewritten turn-final (e.g. *Done:* -> _Done:_) no longer continues the
streamed text, so it was classified unrelated: the stale stream was sealed and
send() posted a second message (#95430 cause B). Seal, then chat.update the
sealed ts with the final; post fresh only if the in-place update fails.

Co-authored-by: liguoyu <guoyu.li@lcfuturecenter.com>
2026-09-25 14:28:14 +05:30
ms-elbdev
72893ca6a1 fix(slack): never re-post a successfully streamed answer
Problem: with native streaming (chat.startStream/appendStream/stopStream)
the same answer could land twice in a thread — once as the streamed
message, once as a fresh chat.postMessage — while the streamed message
kept its live-typing indicator.

Mechanism: `_try_finalize_stream` matched the turn-final against the
streamed text with a raw `startswith`. The agent strips `final_response`
and joins footers with `rstrip()`, so any surrounding-whitespace
difference made the finalize fall through to a plain post although the
open stream already showed the whole answer (and was never sealed). A
`chat.stopStream` failure took the same fresh-post path even when the
streamed text equalled the final. Streams were also keyed per `chat_id`
only, so two concurrent turns in two threads of one channel sealed or
overwrote each other's stream.

Fix:
- Key native streams per `(team_id, chat_id, thread_ts)`; the stream
  consumer stamps the same `thread_id` on every draft frame and on the
  turn-final `send()`, so both resolve to the same key.
- Honor the streaming contract (gateway/AGENTS.md): sends carrying
  `_interim_send` or `expect_edits` never seal a stream.
- Classify the final against the streamed text as equal / extends /
  unrelated with edge-whitespace tolerance (`_stream_relation`). The
  stopStream delta is sliced from the RAW final, so nothing inside the
  answer (blank lines, fences, tables) is dropped or repeated.
- Commit rule: one `chat.stopStream`, one retry only when no tail is
  appended (`markdown_text` APPENDS, so an ambiguous failure must not
  repeat it), then `edit_message(finalize=True)` on the stream ts as the
  idempotent in-place commit — it already owns format/truncate/Block Kit
  and the block-rejection retry. Only when both fail does `send()` post
  a fresh message (a duplicate beats a lost answer).
- Oversized tails and rewritten finals (`notify=True`) seal the stale
  stream on what is visible before falling back, so no stream is left
  with a live-typing indicator.
- `_seal_stream` takes the exact unsent delta instead of recomputing it
  from `final_text`; `disconnect()` and the stream API calls route
  through the stream's own team client.

Tests: tests/gateway/test_slack_native_streaming.py covers the
whitespace-only difference, the stopStream-failure commit path, the
bounded retry, the uncommittable fallback, interim/preview sends,
per-thread keying, oversized tails, rewritten finals and the
GatewayStreamConsumer end-to-end path.

(cherry picked from commit a64d10071ce7816b124467e29407d7d47bfdde8d)
2026-09-25 14:28:14 +05:30
Jony
f588166691 fix(google-chat): avoid guessing auth failure cause 2026-09-24 19:27:41 -05:00
ethernet
7c3d5e93e8 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 16:17:47 -04:00
Robin Fernandes
749220ef00 feat(web): serve managed search through Perplexity
Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.

The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.

Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
2026-09-24 16:17:16 -04:00
ethernet
0f65d698d0 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 13:10:14 -04:00
kshitijk4poor
857ad4b954 fix(dingtalk): tame the SDK's own retry-loop log storm and hand off on websockets incompat (#24851)
dingtalk-stream 0.24.3 start() retries forever inside the SDK, logging a
malformed logger.exception() every 3s, so the adapter breaker never saw the
error. Install a dedup filter on the SDK logger that can't raise on bad
format args, detect the websockets incompatibility (bare or chained
TypeError), log one ERROR with a pin-consistent hint, and hand off via
_set_fatal_error(retryable=False) + _notify_fatal_error(): only a
reinstall of the pinned versions and a restart fixes it. Breaker stays
tripped until the error type changes; constants moved to module level.
2026-09-24 22:29:59 +05:30
kshitijk4poor
3c6b995456 fix(dingtalk): surface persistent SDK TypeError as ERROR with upgrade hint (#24851)
The reconnect storm in #24851 is driven by a dingtalk-stream/websockets
incompatibility that raises TypeError on every start(); backoff never
recovers it. Log the first occurrence per error run at ERROR with an
upgrade hint instead of a generic WARNING.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-24 22:29:59 +05:30
Bartok9
c43ecdee95 fix(dingtalk): add circuit breaker to stream reconnect loop to prevent log storm (#24851)
DingTalk stream-mode reconnection storms the gateway: when start() raises
the same error every cycle, _run_stream logs a WARNING and reconnects at the
60s cap forever, generating hundreds of MB of identical log lines and hanging
the gateway.

Add a per-error-type circuit breaker: after 5 consecutive identical errors,
suppress the repeated WARNING (one ERROR summarises) and pause 300s instead of
spinning at 60s. Reset backoff + counters after a clean start() so a recovered
connection is treated fresh.

Salvage of #24881 re-implemented on current main: the original fix targeted
gateway/platforms/dingtalk.py, which has since been refactored to
plugins/platforms/dingtalk/adapter.py. Same logic, new path; tests import the
new module.

Closes #24851.

(cherry picked from commit 84201c58255ae3d3c9b269ac777ea0ff444d9229)
2026-09-24 22:29:59 +05:30
kshitijk4poor
8d959f965c fix(feishu): log rejected delete_message responses
Surface code/msg at debug when the recall API returns non-success
(matches the exception path) and trim the docstring. Drop the trivial
disconnected-client test.
2026-09-24 22:28:19 +05:30
686f6c61
11d5fb153d fix(feishu): delete truncated stream previews on fallback
Feishu had no delete_message, so a failed finalize-edit plus fallback
send left the truncated edit bubble next to the full final. Implement
the SDK delete and thread the fallback send to the originating message.

(cherry picked from commit c61add84ad40b1bc39288405a0a05b4f621e00fc)
2026-09-24 22:28:19 +05:30
David Metcalfe
9a8748796d fix(kimi): exclude brotli from Accept-Encoding to work around httpx/brotlicffi streaming decode bug (#59556)
httpx's brotlicffi backend (pinned for Discord attachment decoding) fails
on Kimi API's content-encoding: br SSE streaming responses with:

  brotli: decoder process called with data when can_accept_more_data() is False

This surfaces as 'API call failed after 3 retries: Connection error'
on Windows Desktop where the Electron-packaged Python venv includes
brotlicffi (#59556).  DeepSeek and Xiaomi are unaffected — only
moonshot endpoints (api.moonshot.ai, api.moonshot.cn) trigger the bug.

The fix forces Accept-Encoding: gzip so the Kimi API falls back to gzip
compression, which httpx handles reliably on all platforms.  This is the
same workaround already applied in tools/skills_hub.py for sitemap
fetches and the Hermes index download.

Closes associated issues: #28043, #48428

(cherry picked from commit 9c130133b0676138858b73847d1d7c7549339484)
2026-09-24 22:25:06 +05:30
ethernet
f91bd82286 fix: restore catalog Hindsight and harden installers and test guards 2026-09-24 02:03:59 -04:00
ethernet
16652eea18 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	gateway/config.py
#	gateway/config_loader.py
#	gateway/readiness.py
#	hermes_cli/managed_scope.py
#	hermes_cli/plugin_python_deps.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/update_cmd_maint.py
#	plugin-catalog/hindsight.yaml
#	plugins/plugin_loader.py
#	providers/__init__.py
#	scripts/run_tests.sh
#	tests/gateway/test_control_socket_windows_live.py
#	tests/gateway/test_gateway_streaming_nested_config.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_plan_reconciliation_windows_live.py
#	tests/hermes_cli/test_update_apply_shallow_count.py
#	tests/hermes_cli/test_update_concurrent_quarantine.py
#	tests/hermes_cli/test_update_shim_self_lock.py
#	tests/hermes_cli/test_verify_console_scripts.py
#	tests/tools/test_lazy_deps.py
#	tests/tui_gateway/test_subprocess_encoding.py
#	tools/lazy_deps.py
2026-09-23 15:26:34 -04:00
SilverNine
709aaf9764 fix(slack): bound nested attachment block text across the whole unfurl array
a74e0155b6 made attachments[].blocks[] reach the agent through
_append_link_unfurls, but rendered each attachment's blocks with no ceiling.
Slack allows 20 attachments per message, so one alert could project 20x what
a single attachment does (measured: 3,247 chars for 1 -> 64,855 for 20 with
8x400-char rich_text sections each), while the top-level blocks path caps
once at 6000.

Share one budget (_SLACK_UNFURL_BLOCKS_MAX_CHARS, the same 6000 the top-level
path uses) across the array: the first attachment keeps its body, later ones
are truncated against the remainder, and a spent budget still leaves every
header visible. After: 3,247 -> 6,659 chars at 20 attachments.

(cherry picked from commit 3efc1e4c532c38a220d4c6ea53745d59d334aa3e)
2026-09-23 21:47:15 +05:30
Michael Versluis (Berry)
aca906df2c fix(gateway): keep status caches bounded across awaits (#87479)
Telegram's per-(chat_id, status_key) status-message cache grew without
bound; give it the same _STATUS_MESSAGE_IDS_MAX=2000 FIFO half-trim the
Slack adapter already has. In both adapters, guard the post-await write-back
after a successful edit with a compare-before-write (only re-store the id if
the cached entry is still the one we edited) so an eviction or replacement
that happened during the await is not undone.

Partial salvage of #87480: kept the Telegram bound and both compare-before-write
guards (re-applied by hand, 17586 behind), defined the max as a class attr
like Slack instead of an instance attr, dropped the 4 new tests.

(cherry picked from commit 43ae95e98e)
2026-09-23 21:47:15 +05:30
devorun
b331e5d282 fix(buzz): persist channel cursors off the event loop
_handle_events called _save_cursors inline whenever a batch moved the
cursor. It ends in atomic_json_write (mkstemp + fsync + os.replace), and
on the WebSocket transport _handle_events runs once per inbound EVENT
frame, so every message paid an fsync on the gateway's event loop,
stalling every other adapter and in-flight turn for its duration.

Split the snapshot from the write: the payload is still built on the
loop (_channel_state is loop-owned), and the write goes through
asyncio.to_thread. The snapshot is taken under an asyncio.Lock so a
slower, older write can never land after a newer one and regress the
durable cursor. connect() keeps the synchronous _save_cursors.

(cherry picked from commit 299929741429262f497b86a67940ffb37faa1694)
2026-09-23 21:47:15 +05:30
John Paul Soliva
16db511ed5 perf(config): route the config and manifest loaders through fast_safe_load
utils.fast_safe_load already exists, is pinned by tests/test_fast_safe_load.py, and
its comment names exactly these payers: 'startup parses config.yaml and every plugin
manifest, so the slow path cost ~0.9 s of cold start'. The migration was started —
hermes_cli/config.py uses it eight times, hermes_cli/main.py and hermes_cli/plugins.py
too — but the file-level loaders it was written for were never converted.

The cost is config SIZE, and the size is the installer's doing: it seeds config.yaml
by copying cli-config.yaml.example, 120,897 bytes of mostly comments. Nothing caches
load_gateway_config() and it has 238 production call sites.

Profiled before assuming a cause — reader.forward 34 ms, scanner.scan_to_next_token
28 ms, reader.peek 13 ms: the pure-Python PyYAML scanner, nothing else.

Measured A/B on one realistic pass (gateway config + every bundled plugin
description), medians of five runs, __pycache__ cleared between arms:
48.8 ms -> 2.46 ms. Per path: load_gateway_config 48.3 -> 1.84 ms, managed
config.yaml 44.4 -> 1.03 ms, 105 plugin.yaml manifests 59.3 -> 5.9 ms.

Same parse, same restricted tag set, same result — only the loader changes. Drops the
three 'import yaml' statements the swap orphaned.

(cherry picked from commit a4efd6a4506043159310eba02c32b673c455b88b)
2026-09-23 21:26:49 +05:30
ethernet
c13ea774e6 refactor: make install-stamp.json the single runtime version identity
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.

Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).

Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
2026-09-23 11:41:01 -04:00
ethernet
063560696e Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	tests/gateway/test_launchd_exit_timeout_drain_cap.py
2026-09-23 10:52:20 -04:00
teknium1
6bbbc65b45 fix(image_gen): resolve OpenAI and Meta image base URLs through the secret scope
Sibling sites of #119986's class: the openai and meta-ai image backends
resolved their API key through get_secret but the base URL through
os.environ, so on a multiplexed gateway a routed profile's key was sent to
the launch profile's endpoint. Both fields now come from the same scope
(get_secret_str, like the DeepInfra video backend after #119986).
2026-09-23 07:49:35 -07:00
John Paul Soliva
bb3011a202 fix(video_gen): resolve OpenRouter and DeepInfra credentials per profile, not from os.environ
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:

- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
  the credential pool, not the environment. Chat and image_gen/openrouter find
  it through resolve_runtime_provider. video_gen reported OpenRouter
  unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
  routed profile's video jobs were submitted, polled and downloaded with the
  launch profile's key and billed to that account. A profile whose key lived
  only in its own .env could not use the backend at all.

The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.

OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
2026-09-23 07:49:35 -07:00
John Paul Soliva
707852d92f fix(web): rebuild the Exa and Parallel clients when their API key changes
cached_sdk_client returned the client cached on tools.web_tools before it
read the key, so the Exa, Parallel and AsyncParallel clients kept the key
they were first built with for the life of the process. A key fixed in .env
and applied with /reload still sent the old key (401s until a restart), and
on a gateway serving multiplexed profiles every profile's Exa and Parallel
calls went out on whichever profile's key built the client first, billed to
that account. A key removed from the environment also kept being used.

Resolve the key on every call and reuse the cached client only when it was
built with that key; the slot now holds (key, client) as one value so two
builds racing under different keys cannot record one key beside the other
key's client. Firecrawl already compares its credential before reusing its
client; this brings the two SDK-backed providers in line.
2026-09-23 07:49:35 -07:00
John Paul Soliva
b5a300fe34 fix(email): pairing, decline and gateway grants reach the gateway instead of dying in the adapter pre-gate
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.

The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.

Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
  decline) must authenticate its From:, open access or not: the
  pairing code or refusal is mailed back to that address. A granted
  sender still needs it short of open access, since a pairing grant
  keys on From: just as the allowlist does. Open access follows the
  gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
  GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
  longer exempts a listed address from From: authentication either
  (on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
  registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
  list grants a stranger nothing there, so the env flag alone no
  longer exempts one from From: authentication (that path mailed a
  pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
  dropped. The gateway's check also matches an address by its bare
  local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
  (a chat username) would admit or pair alice@<any domain>. The lists
  are parsed as the gateway parses them, JSON list literals included,
  or '["alice"]' would slip past this guard.

_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.

Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
  went from 0 events reaching the gateway to 1. pair mails a pairing
  code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
  GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
  stranger@<domain> through, under ignore or pair. Without the
  local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
  with a forged From: under allow-all beside an EMAIL_ or
  GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
2026-09-23 07:48:32 -07:00
ethernet
c6b358592d Merge origin/main into ethie/pm-clean 2026-09-23 10:43:45 -04:00
ethernet
f4a38b1ee8 Merge origin/main into ethie/pm-clean 2026-09-23 10:13:26 -04:00
Austin Pickett
a1838ea87a fix(telegram): keep update receipts across adapter rebuilds and restarts (#120257)
Completed update IDs lived only in the adapter's memory. The gateway
reconnect watcher builds a new TelegramAdapter and connects it with
is_reconnect=True, which keeps Telegram's pending queue, and a new PTB
Updater polls from offset 0. Telegram then resends every update whose
acknowledgement (the next getUpdates offset, or the cleanup call in
Updater.stop) never landed, and the fresh adapter admitted them again.

Write completed IDs to telegram_update_receipts_<bot_id>.json in the
adapter's Hermes home and seed admission from it once per bot. Receipts
older than 24h are dropped: the Bot API keeps unconfirmed updates no
longer than that, and it keeps the lookup clear of the random ID restart
Telegram may do after a week without updates. Writes are coalesced and
run off the loop; disconnect waits for the last one.

Refs #68502

Co-authored-by: Joe Githler <5716896+NoTimeforInfinity@users.noreply.github.com>
2026-09-23 10:05:17 -04:00