Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):
- the OAuth context-length probe (agent/model_metadata.py) and the
/model picker's live discovery (hermes_cli/codex_models.py) both GET
https://chatgpt.com/backend-api/codex/models with
Authorization: Bearer <gateway key> whenever model.context_length is
not pinned — the key is sent to a service it does not belong to and
cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.
Fix, mirroring the quota probe's existing gate in auth_codex:
- both catalog sites now decline to probe non-JWT credentials (real
Codex access tokens are JWTs; a gateway key is not one) and fall
back to the static table / offline sources — same outcome as the
doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
/models instead of chatgpt.com (catalog URLs are built from the
resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
same way the text client does.
Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.
(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
The session-hygiene compressor runs on the loop's default executor
(run_in_executor(None, ...)), which the self._executor quiesce never joins.
It was only registered with _track_deferred_agent_worker once a timeout,
turn-hold or unwind deferred it, so a gateway stop() that reached the
SessionDB close while the awaiting turn was still waiting saw zero deferred
workers and closed/checkpointed state.db under the worker's late write
(#101064 shape).
Track the future right after run_in_executor, mirroring
run_codex_hygiene_compaction. The existing close guard (#102198) now skips
close:session_db while the worker is live, and shutdown's interrupt pass
reaches the hygiene agent. _defer_agent_cleanup_until_future_done is kept
for cleanup; re-tracking the same future is idempotent (dict key).
The default executor is deliberately NOT shut down: stop() still uses
asyncio.to_thread afterwards (terminal runtime-status flush).
Regression test adapted from @pmaho's
test_default_executor_worker_is_seen_by_the_close_guard (#121360), now
driving the real _hmwa_hygiene_detached_attempt submission site.
Co-authored-by: pmaho <42786356+pmaho@users.noreply.github.com>
Two invariant tests through the real text-fallback picker and real admission
(plugin_world), replacing the helper-level tests from #121623 by
@victor-kyriazakos: open-and-exit writes nothing, and a session that ticks a
legacy-name-disabled platform and unticks another writes exactly those two rows
and keeps an enabled entry for a plugin the picker does not show. Both fail on
main by behaviour; the second also fails on the ported fix alone (the
manifest-name disable survived the tick).
Follow-up to the ported picker fix (#121623):
- Ticking a row re-enables it even when plugins.disabled holds the manifest
name (e.g. telegram-platform, written by pickers before #40190). The save
only dropped the key and its bare leaf, so the gate still matched the
manifest name and the plugin stayed off. It now purges every alias like
`hermes plugins enable/disable` (_apply_activation, shared with
_set_plugin_enabled), with one discovery scan per save.
- The success line counts the rows actually turned on/off instead of treating
every unticked row as "disabled"; the unused new_enabled return is gone.
- _entry_status shares the per-row status call between the picker
preselection and `hermes plugins list --enabled`.
The bare `hermes plugins` picker preselected only rows listed in
plugins.enabled. Bundled platforms, backends and model providers are active
without a list entry, so they opened unticked, and the save on exit wrote every
unticked row into plugins.disabled: opening the picker and leaving without a
change disabled every messaging adapter on the next gateway restart.
Rows now open ticked by the load-time rule (_plugin_status), and only rows the
user flipped are written.
Ported onto the plugins_cmd_toggle sibling and the admission-authority save
(the original targeted the pre-split plugins_cmd.py). The nested
canonical-key composite test now ticks its row explicitly: it handed the menu
a checkbox state that contradicted the config and relied on the old
rebuild-everything save. The original helper-level tests are rewritten
against real admission in a follow-up commit.
(cherry picked from commit 3a0d5f1b8bce3b23b115ba4995b8c23dfbba9ad0)
A session whose started_at is corrupt and has no in-window activity or
message timestamp still returned the raw cell as last_active, and the
order_by_last_active fallback sorted it above every healthy session.
Both fallbacks now go through the same window as the UNION ALL values;
a session with no trusted timestamp gets NULL.
27df3b8847 dropped the exclusion 65e79880c8 added, so the 30-minute e2e
job (shallow checkout, no bwrap, 900s per file) ran tests/e2e/core/upgrade
next to its own e2e-upgrade job. Its install/update files then timed out
and the job was cancelled at the step limit on main and every Python PR.
WSTransport.close() scheduled ws.close(code=1011) on every call, so
handle_ws's normal teardown reported 1011 before its own close. Move the
off-loop socket close into a one-shot WSTransport.abort() that the fanout
overflow path calls; close() is byte-identical to main again. Rename
_close_stalled_socket -> _close_socket(code, reason) with accurate log
text, share an _on_loop() helper with write(), and make the overflow test's
slow peer a real WSTransport whose socket close must be awaited with 1011.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Replace the loose 'closed or any detach/overflow token' oracle with a hard
'slow peer transport closed' assertion; fold the close-raises case into the
healthy-keeps-streaming test; merge the close/detach non-closing checks.
- WSTransport.close no longer special-cases 'ws is self'; the SocketClient
test helper now passes a real ASGI-ws stand-in instead.
- call_soon_threadsafe on a closed loop is suppressed explicitly.
- FanoutTransport docstring: closing an overflowing socket also drops other
sessions multiplexed on it; reconnect + replay recovers them.
A bounded mailbox overflow dropped membership without closing the socket
or emitting a control frame, so heartbeats kept succeeding and the pane
froze. Close only the overflowing peer after the lock is released.
Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit 3153460ad40f8452bfae6611d70d1425c24a6e3a)
- A mid-turn notify reply (/status, /approve, clarify answer) shares the
stream's thread key; it no longer seals and overwrites the half-streamed
answer. In-place replacement now requires the final to match the stream
after normalizing mrkdwn markers and whitespace; anything else posts fresh
and leaves the stream open.
- One _commit_stream helper for both seal-then-commit paths, so the rewrite
path also falls back to chat.update when stopStream fails.
- A stream reopened after a server-side seal is seeded with only the text
past the sealed message (tracked as 'base'), not the whole segment.
- Streams older than 15 min are sealed and dropped on the next start.
- _stream_key reuses _workspace_thread_key/scope_id_for_chat; the stream
dict no longer duplicates chat/team ids.
Keep test_whitespace_only_difference_does_not_duplicate (a whitespace-only
final never re-posts a streamed answer) and
test_two_threads_finalize_their_own_streams (per-thread stream key), plus
test_stop_and_update_both_fail_falls_back_to_fresh_post as the successor of
main's stop-failure test. Drop the helper-level, logging and
routing change-detectors.
The reopen test from 2b4ff4e23b indexed _active_streams by chat id, but
streams are now keyed (team, chat, thread_ts). Assert through _open_streams,
check the reopened stream anchors to the same thread, and keep one test for
the item (drop the reopen-failure variant).
5648f81431 fixed this exact server-side seal (Slack closes a native stream
after a few minutes of a long turn, live-observed at ~5m20s; the lifetime
is not documented) for the native task-card stream: on
message_not_in_streaming_state from appendStream, drop the dead ts and
start a fresh stream, seeded with the full current content so nothing is
lost.
send_draft — the plain-text native streaming path used when task cards
are not enabled — hits the identical seal but never got the fix: its
generic except block only recognizes the feature-gate markers
(not_allowed, missing_scope, ...) and otherwise just logs debug and
returns failure. gateway/stream_consumer_transport.py's
_send_draft_frame() docstring is explicit that "any failure permanently
disables drafts for this run" — so a long turn streaming as plain text
degrades to the edit-based fallback for its remainder exactly the way
the task-card bug did before 5648f81431.
Mirror the task-card fix: on message_not_in_streaming_state from
chat.appendStream, drop the dead ts and _start_stream() a fresh one
seeded with the full accumulated text (not just the delta), so the next
frame's delta still resumes correctly. One reopen per frame; a second
rejection propagates as a real failure, matching the twin's behavior.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 2b4ff4e23bdf684d2fde1a9512f11dcd7ef99c42)
A mrkdwn-rewritten turn-final (e.g. *Done:* -> _Done:_) no longer continues the
streamed text, so it was classified unrelated: the stale stream was sealed and
send() posted a second message (#95430 cause B). Seal, then chat.update the
sealed ts with the final; post fresh only if the in-place update fails.
Co-authored-by: liguoyu <guoyu.li@lcfuturecenter.com>
Problem: with native streaming (chat.startStream/appendStream/stopStream)
the same answer could land twice in a thread — once as the streamed
message, once as a fresh chat.postMessage — while the streamed message
kept its live-typing indicator.
Mechanism: `_try_finalize_stream` matched the turn-final against the
streamed text with a raw `startswith`. The agent strips `final_response`
and joins footers with `rstrip()`, so any surrounding-whitespace
difference made the finalize fall through to a plain post although the
open stream already showed the whole answer (and was never sealed). A
`chat.stopStream` failure took the same fresh-post path even when the
streamed text equalled the final. Streams were also keyed per `chat_id`
only, so two concurrent turns in two threads of one channel sealed or
overwrote each other's stream.
Fix:
- Key native streams per `(team_id, chat_id, thread_ts)`; the stream
consumer stamps the same `thread_id` on every draft frame and on the
turn-final `send()`, so both resolve to the same key.
- Honor the streaming contract (gateway/AGENTS.md): sends carrying
`_interim_send` or `expect_edits` never seal a stream.
- Classify the final against the streamed text as equal / extends /
unrelated with edge-whitespace tolerance (`_stream_relation`). The
stopStream delta is sliced from the RAW final, so nothing inside the
answer (blank lines, fences, tables) is dropped or repeated.
- Commit rule: one `chat.stopStream`, one retry only when no tail is
appended (`markdown_text` APPENDS, so an ambiguous failure must not
repeat it), then `edit_message(finalize=True)` on the stream ts as the
idempotent in-place commit — it already owns format/truncate/Block Kit
and the block-rejection retry. Only when both fail does `send()` post
a fresh message (a duplicate beats a lost answer).
- Oversized tails and rewritten finals (`notify=True`) seal the stale
stream on what is visible before falling back, so no stream is left
with a live-typing indicator.
- `_seal_stream` takes the exact unsent delta instead of recomputing it
from `final_text`; `disconnect()` and the stream API calls route
through the stream's own team client.
Tests: tests/gateway/test_slack_native_streaming.py covers the
whitespace-only difference, the stopStream-failure commit path, the
bounded retry, the uncommittable fallback, interim/preview sends,
per-thread keying, oversized tails, rewritten finals and the
GatewayStreamConsumer end-to-end path.
(cherry picked from commit a64d10071ce7816b124467e29407d7d47bfdde8d)
Seeding the disk-watch baseline on first observation also swallowed the
case where the process started before login: file absent, then an
external `hermes mcp login` writes it, and the provider never reloaded.
Only skip the reload when the provider already holds tokens in memory.
- drop _run_stream_subscribers; the sweep reads _RunStream.subscribers
- per-write timeout via asyncio.timeout (no Task per token), force_close kept
- _sse_frame(id=) instead of hand-prepended id: lines
- reconnect queue gets headroom for the replay length
Keep the concurrent-subscribers and reconnect-replay contracts from #69817;
drop the three tests that pin _RunStream internals (overflow bounds, sweep
bookkeeping, response-boundary transport cleanup).
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Strict OpenAI clients reject the named hermes.tool.progress SSE frames. Setting
tool_progress_events: false under platforms.api_server (loaded into
PlatformConfig.extra by from_dict) now drops them; default stays on.
Reimplements the intent of #42640 against the adapter config actually read in
production. Overlaps #49069 (erikerosev).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Recovery paths (guardrail halt, partial_stream_recovery) can return a
final_response without firing any content delta; /v1/chat/completions
streaming then closed with an empty body. Mirror _ResponsesStream.collect_result.
Co-authored-by: fmercurio <15571697+fmercurio@users.noreply.github.com>
interval_hours <= 0 made should_run_now() true on every idle tick,
re-running the review pass each time. Route it through the same
floor-with-default helper as the day counts (renamed _bounded_count),
and log the fallback warning once per (key, value) since the dashboard
status endpoint polls these getters.
get_stale_after_days()/get_archive_after_days() accepted any int from
curator.stale_after_days/archive_after_days. archive_after_days: 0 sets
archive_cutoff to now, so apply_automatic_transitions() (runs unconfirmed
on an idle tick, curator on by default) archives every skill with any past
activity on the next pass; a negative value builds a future cutoff. The
manual path already refuses the same value (_cmd_prune: "--days must be
>= 1"), and "0 disables" is the repo convention elsewhere.
Fix: a value < 1 falls back to the default with one warning naming the
key, the same bound _cmd_prune enforces. Same class of fix as b01b1c8b
(bound kanban gc retention so -N/0 cannot mass-delete).
(cherry picked from commit 6685565a8c5147c8d7c23602f371f60571c44074)
An SSH auth-failed boot error carried none of the local, host-key or
reauth latch tags, so it stayed retryable: every getConnection/api call
re-ran startHermes, re-emitted running: true and hid the boot-failure
overlay before its Gateway settings button could be clicked. Classify
the rejection (kind/sshError tag, or the message once stringified),
latch it like a host-key change, and keep it out of the renderer's
auto-retry loop. reset/repair/apply-config still release the latch.
Co-authored-by: x7peeps <xtpeeps@qq.com>
After the crash-loop budget trips on STATUS_STACK_BUFFER_OVERRUN, relaunch
once with GPU off instead of leaving a blank window. Sandbox stays intact.
Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit 02c096d5b2f0e0872b7dc35f408b70e1e51ead2e)
windowsHide on a GUI-subsystem Electron parent does not stop git.exe from
allocating a console. Route those spawns through a console-subsystem
python.exe host that starts git with CREATE_NO_WINDOW (0x08000000) and
forwards the git argv unchanged, including simple-git review probes.
The Windows tree-kill check ran at the top of backendShutdown and threw
before the graceful teardown, pool stop and straggler reap. It now takes
the owned child handles up front, runs after the reap on the children that
are still running, and surfaces a failure only once cleanup is done. A lock
whose delete fails is kept and logged instead of throwing out of close, and
exitAfterBackendShutdown still exits when shutdown reports a failure.
The reasoning half only touched a docstring, and its tests pinned main's
existing continuation (reasoning-off retry, then the 'No visible answer'
ceiling). Nothing changed, so the PR stays on the Windows close/stop fix.
finish_reason=length with empty visible content and a non-empty reasoning
or reasoning_content field uses the existing thinking-budget abort. No
model id is consulted. Empty content with no side channel still continues.
Desktop close/stop no longer discards Windows taskkill failures. After the
same tree-kill, owned PIDs are inventoried and only unheld gateway locks
are cleared.