38 Commits

Author SHA1 Message Date
Austin Pickett
6397a84a91 fix(sessions): un-hide an auto-archived chat when it is resumed or compressed
The idle sweep archives a whole compression lineage. When the chat later
resumed and compressed, the new tip was inserted with archived=0 under the
archived root, and the sidebar admits a lineage by its root's flag, so the
live chat stayed hidden.

Record sweep provenance in a new sessions.auto_archived column. Publishing a
compression child or reopening a session under a sweep-only archive
un-hides the lineage; a deliberate archive (sidebar, API, CLI) clears the
provenance and keeps the chat hidden. Compression children now inherit the
parent's archive state so a manually archived lineage stays uniform.

Fixes #117713
2026-09-28 21:03:09 -04:00
kshitijk4poor
dd66c132f2 fix(sessions): keep tool-call uids paired through rewrites, folds of repeated ids and chunked copies
Found by the pre-merge correctness and hermes-specific review of this stack:
- A fold after a row whose one response repeated a provider id (one shared uid) appended the
  absorbed turn's uid in the wrong slot. The shared uid now fills each of the row's own
  occurrences before the fold appends (per_occurrence_tool_call_uids). A single response keeps one
  shared uid for a repeated id: every result carries it, so no call looks unanswered.
- A digest-less re-flush of a restored fold (replay heal, adopt after a digest mismatch) let the
  stored pre-fold map overwrite the longer live list, so the later occurrence's results paired
  with nothing. A live list that starts with every stored occurrence is kept (stored uid or list).
- A dict that lost its uid map (a clone) rewrote tool_call_uids to NULL. The row-addressed
  rewrite now fills stored uids for the calls the dict still names; a call it dropped keeps none.
- append_messages_batch(chunk_rows=...) reset the pairing index per chunk, so a result in the chunk
  after its call was stored without tool_call_uid (branch/seed copies). The whole batch is paired
  before it is chunked.
- _merge_assistant_into keeps its dict guard on the absorbed turn's map (a plugin-built dict with a
  non-dict map no longer raises).
2026-09-29 04:15:21 +05:30
kshitijk4poor
d9808867c7 fix(sessions): an assistant fold keeps every occurrence uid of a reused provider id
Folding two assistant turns that reuse a provider id (llama.cpp-style constant ids) leaves
both calls in tool_calls, but the {id: uid} union kept only the later uid, so the first
occurrence lost its identity. The id now maps to a list of uids, one per occurrence in
tool_calls order (merge_tool_call_uids); index_tool_call_uids pairs each call with its own
occurrence, so a following result still pairs with the nearest call. Provider ids are never
rewritten. Persisted and restored as-is (the JSON map accepts list values).

Reported by @andrexibiza on #126307.
2026-09-29 04:15:21 +05:30
kshitijk4poor
9d680cb95c docs(sessions): the bounded legacy backfill and branch copies keeping message_uid 2026-09-29 04:15:21 +05:30
Eva
c2e2150d78 fix(sessions): close the identity gaps the review found and prove the recovery paths
- Strip persistence-only fields from a select_context() selection: the request copy was stripped before
  the hook, but an engine handing back the conversation_messages clones put timestamp, _row_id and the
  new ids on the wire (pre-existing for the old fields).
- Every row carries a message_uid whoever wrote it: an AFTER INSERT trigger mints one for a row inserted
  without it (an older build writing into a v31 store); the one-time backfill is gated by its own
  state_meta marker so a store whose schema version cannot advance (no FTS5) never rescans.
- One merge-witness helper (record_absorbed_message) for every fold of a durable message: alternation
  repair's user and assistant merges, the compressor's in-flight restatement onto the carrier, the real
  user anchor folded into a scaffolding turn (the anchor's uid leads), micro-compaction's adjacent-user
  merge.
- tool_call_uids follows tool_calls: it is a payload column, so an assistant merge's unioned map is
  rewritten with the row instead of being overwritten by the stale stored map on adoption.
- A result pairs with the nearest call naming its provider id: each assistant row's ids shadow earlier
  occurrences in the batch, restore and live-list indexes, a call row without a map leaves its result
  unpaired, and nothing pairs across a user turn.
- A composite rewind returns and installs the replacement row's uid on the live scaffold.
- Tests: import_foreign_history, hermes sessions recover and backup import carry the ids; the ACP
  restore test accounts for the uid on restored rows.

(cherry picked from commit b30636a49cb29fdd21ba4100eef9d810897338be)
2026-09-29 04:15:21 +05:30
Eva
76d0ac6caa feat(sessions): give tool calls and their results a per-occurrence uid
Provider tool-call ids repeat: Hermes mints deterministic call_<12hex> ids for identical calls and models
reuse ids, so tool_call_id cannot identify an occurrence. An assistant row now carries
messages.tool_call_uids, a {tool_call_id: uuid4} map for its tool_calls, and each tool-result row carries
the matching messages.tool_call_uid. The provider-facing tool_calls JSON is untouched, so row identity,
display identity and the CAS digest are unchanged and nothing nested reaches the wire.

Minted at the assistant row's first insert; paired onto the result in the same batch, from the live list
when the result is flushed later, or on restore from the preceding assistant row; copied by every clone
and re-insert; restored as the live _tool_call_uids / _tool_call_uid; stripped from provider requests.

(cherry picked from commit aa49c0f59788676d4beca288181711c558fca165)
2026-09-29 04:15:21 +05:30
Eva
f51d0a7233 feat(sessions): persist the consecutive-user merge witness on the survivor row
A merge survivor's _absorbed_message_uids lived only in memory, so an engine that restarted after a
merge was back to parsing "\n\n" to learn which stored messages a composite user row contains. Store
the list in messages.absorbed_message_uids (an owned column: the survivor's row-addressed rewrite and
every re-insert carry it), restore it as the live key, and round-trip it through export/import.

(cherry picked from commit 0bc05f2656229b1010f330e1998684b04d5f963f)
2026-09-29 04:15:21 +05:30
Eva
751d8526e3 feat(sessions): give every persisted message a durable message_uid that reaches context engines
The physical messages.id is re-issued by every copy (in-place compaction generations and their
concurrent-tail clones, rotation children, replace_messages), and _row_id is opt-in on restore, so a
context engine that keeps its own per-message state had nothing stable to key on across a restart or a
compaction boundary and re-identified rows by content and timestamp.

Add messages.message_uid (uuid4 hex), minted once at a row's first insert and stamped on the caller's
dict, copied by every SQL clone, kept by every re-insert of the same dict, left alone by row-addressed
rewrites, restored on every projection, and stripped from provider requests. A consecutive-user merge
survivor keeps its own uid and records the absorbed uids in _absorbed_message_uids. Schema v31 backfills
a uid onto rows that predate the column.

(cherry picked from commit 20d67ffb5e50e0ed956e56ddc7f8e28d36059e2e)
2026-09-29 04:15:21 +05:30
teknium1
16fe260aab fix: read-only open waits its busy budget once, not per retry
After #120386 raised the read-only busy timeout to 5 s, retrying a lock
inside _open_read_only multiplied the wait to ~20 s on blocking callers
(TUI profile loop, exit epilogue, hermes status). The connection already
waited the read budget; only transient disk-I/O errors are retried now.

Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still
classified as a transient lock.
2026-09-23 11:35:07 -07:00
teknium1
8ac45786bf fix(state): SessionDB open waits out a lock lost inside the FTS constructor
In rollback-journal (DELETE) mode a sibling process can take the write lock
between schema load and the messages_fts probe. FTS5's xConnect then fails its
%_config read and SQLite reports SQLITE_BUSY with the text "vtable constructor
failed: messages_fts". Every state.db lock classifier matched on the words
"locked"/"busy", so:

- a writable SessionDB() failed after 1s instead of waiting out the lock with
  _WRITE_PATIENCE_S, and callers disabled persistence for the run;
- a read-only open (dashboard, `hermes sessions list`, cross-profile readers)
  failed on the first busy timeout with no retry at all;
- the error read as not transient (dashboard 500, not 503) and as persistence
  cause "unknown" instead of "locked".

Add hermes_state_errors.is_sqlite_lock_error: SQLITE_BUSY/SQLITE_LOCKED by
result code when SQLite supplies one, text only when it does not (our own
re-raised messages, RPC-wrapped strings). Route the writer open patience loop,
the _execute_write retry, the reconcile re-raise, the WAL->DELETE flip, the
maintenance holder probe, is_transient_sqlite_error and
classify_persistence_error through it. The read-only open retries a lock
inside its existing bounded retry budget, next to the transient IOERR case.
2026-09-23 11:35:07 -07:00
teknium1
2182f51d7c fix: name the process holding the state.db write lock when a writer times out
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.

SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.

Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
2026-09-20 10:01:42 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
0c24dd56cb fix: pin the auxiliary wire's route-scoped reasoning_details strip; document OpenRouter/Nous-only replay
Replaces the direct transport.build_kwargs test with one driving prepare_chat_messages
(groq client strips, OpenRouter client keeps), so dropping the base_url kwarg goes red.
2026-09-19 09:28:45 -07:00
teknium1
04a151c3e0 docs(gateway): identity survives restore, relay, callbacks and thread hops 2026-09-19 02:28:50 -07:00
teknium1
f39ae9000c docs(sessions): say what user_id holds for desktop and dashboard sessions 2026-09-18 10:16:46 -07:00
teknium1
eb11e7eac9 fix(agent): reasoning promoted on a reasoning-only clean stop is never persisted as an ordinary reply
A clean ``stop`` with empty content and reasoning text still promotes the
reasoning to ``final_response`` (the vLLM nemotron parser files the whole
answer as reasoning; re-running the empty-response ladder re-billed the prompt
8x, #109205). What changes is what the assistant ROW carries: ``content`` stays
empty, the text lives in ``reasoning``/``reasoning_content``, and the promoted
text rides the ``api_content`` sidecar. ``build_api_messages`` substitutes the
sidecar on the wire, so the next request replays the answer byte-identically
(same bytes as before this change) and role alternation is intact, while
state.db, ``session.history`` and every other history surface no longer show
chain-of-thought as a reply indistinguishable from a real one.

The "Reasoning-only clean stop" log moves INFO -> WARNING and names model,
provider, api_calls and tool_turns: a model that ends every turn this way is
stalled (planning monologue, zero tool calls) while the turn reported
"complete".

The thinking-prefill interim row (``_thinking_prefill``) is ephemeral
scaffolding already skipped by the flush and popped before the final answer, so
it never persisted a content==reasoning row; no change there.

Live probe (real AIAgent, real SessionDB, temp HERMES_HOME, DeepSeek-shaped
reasoning_content only): before, the assistant row had content == reasoning
== reasoning_content with api_content NULL and an INFO log; after, content is
empty, api_content carries the text, the replayed second turn sends the same
assistant content on the wire, the WARNING carries the route, and a normal
"Hi there." reply persists unchanged.

Part of #111761

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-16 17:21:57 -07:00
teknium1
3147ccb1ad docs(session-storage): guard raises, not warns; database-location section follows get_hermes_home(); zh-Hans mirror
Corrections on top of #110617: hermes_state._ensure_test_isolation raises RuntimeError (the
draft said 'warning'); the older Database Location section still told readers the default
path is ~/.hermes/state.db, which the new section says never to hard-code. Mirror the new
section into the zh-Hans page.
2026-09-15 05:30:32 -07:00
KoNit-K
71dade7e34 docs(website): document profile state isolation 2026-09-15 05:30:32 -07:00
Teknium
3114916ee4 fix(gateway): carry accepted-input ownership through persistence
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.

Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
2026-09-07 14:11:18 -07:00
Teknium
f120ea7149 fix(gateway): exclude observations from accepted-turn ownership 2026-09-07 14:11:18 -07:00
Teknium
136d80d040 fix(gateway): retain one durable owner for failed input turns 2026-09-07 14:11:18 -07:00
Teknium
a9ef4a7625 fix(codex): keep transport echoes out of durable user history
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.

Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.

Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698

Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
2026-09-07 08:09:57 -07:00
Teknium
026e3e84ea fix: reject unavailable desktop profile and session targets 2026-09-07 06:26:22 -07:00
Teknium
12891ea3bd fix(desktop): scope tool changes and commit rebuild ownership together 2026-09-07 06:26:22 -07:00
Teknium
42f8389987 fix: restore persisted assistant replies alongside tools and reasoning 2026-09-07 04:56:22 -07:00
Ahmad Al-Faqih
9d1eec0ce0 fix: retain Responses assistant replies in RPC session history
Carve the history-only implementation and regression from #104754
(b4bfa76facb43111a863b1256c89875e3f01fde0) by PLASMA-FR; omit
the unrelated locale-picker changes. Complement merged #104523.

Real serve + WebSocket reconnect: SQLite retains two rows; before,
resume/activate/history return only the user; after, both survive.
Plain-content control retains both rows in both arms. Native macOS
sleep and the full desktop symptom are not established by this probe.

Refs #68321
2026-09-07 04:56:22 -07:00
Teknium
94ff4fe8f9 test: verify profile rebuild persistence through live serve 2026-09-07 04:54:49 -07:00
Teknium
4fb332b427 docs: sync developer-guide agent-core docs with the facade/siblings layout (#102117) 2026-09-04 00:07:14 -07:00
Teknium
2b55ded1ac perf(state): keep delegate-child transcripts out of the trigram FTS index (schema v30)
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.

Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.

The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
2026-09-03 02:35:37 -07:00
Teknium
5dfd1e77a8 fix(recovery): register lazy state.db tables in one schema map; cover .recover lane and count-mismatch loss
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.

Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.

Addresses #100313
2026-09-02 00:00:55 -07:00
spfcraze
8f91e249e4 perf(state): add messages(session_id, id) index for window/ordering queries
Every ORDER BY id query on the messages table sorted or scanned the
whole session: get_messages_around's window seek, latest_message_row_id
(LIMIT 1), and get_messages' full-load ordering all paid O(session
history) per call — hot mid-turn via session_search and reactions.
messages.id is an original column (INTEGER PRIMARY KEY AUTOINCREMENT),
so the index lives in SCHEMA_SQL next to idx_messages_session — no
legacy-column migration hazard (the kanban lesson from #28776 does not
apply).

Measured (real schema, one 20k-message session, median of 30):
get_messages_around 7.08 -> 0.22 ms (32x), latest_message_row_id
3.37 -> 0.011 ms (307x), get_messages full load 111.6 -> 98.6 ms
(1.13x — remaining cost is row deserialization, not the sort).
Window results byte-identical at probe points across the session.

Tests: VM-step pin (get_messages_around bounded work, calibrated
~12 vs ~855 handler calls, threshold 300 — fails without the index)
and window parity with/without the index. No EXPLAIN/plan text
(behavior contracts, AGENTS.md).
2026-08-02 21:16:17 +05:30
Teknium
43874d1a96 docs: accuracy sweep + coverage for 2 months of shipped features
Accuracy pass (all 373 pages audited against code, 13 parallel audits):
- configuration.md: 12 stale defaults/keys (file-sync rewrite, clarify
  timeout, streaming knobs, iteration budget, TTS/STT enums)
- reference/: commands/env-vars/toolsets/tools synced with
  COMMAND_REGISTRY, argparse tree, OPTIONAL_ENV_VARS, TOOLSETS
  (28 env vars added, 3 phantom removed, mcp__ naming, webhook
  platform restricted toolset)
- features/, messaging/, developer-guide/, guides/: ~60 factual fixes
  (web_extract truncation, dashboard auth fail-closed, delegation
  blocked tools, adapter signatures, session schema v23, phantom
  Matrix env vars, hermes setup tts, auth spotify, webhook --skills)
- zh-Hans: explicit heading IDs fix 2 broken WSL2 anchors

New coverage for features shipped in the last 2 months (verified
against code before writing):
- compression.in_place, verify-on-stop (+v31/v32 migration reality),
  ${env:VAR} SecretRef, display.timestamp_format, session:compress
  hook + thread_id/chat_type fields
- /journey learning timeline, per-channel model/system-prompt
  overrides, /sessions search, clarify multi-select, -z --usage-file,
  uninstall --dry-run, config get/unset
- MCP elicitation, extra_headers, discover_models, api-server run cap,
  Bedrock cachePoint, Discord reasoning_style, Google Chat clarify
  cards, vibe reactions, resume cwd restore, Yuanbao forwarded
  messages, api_content sidecar, roaming pet, tool_progress log mode,
  WhatsApp polls/locations, kanban per-task model + lifecycle hooks
2026-07-29 08:48:05 -07:00
Teknium
176a98c39a docs: refresh salvaged audit fixes against current main
PR #36051's values went stale since May 31: session-store SCHEMA_VERSION
is now 21 (PR said 14), and the dashboard ships 8 built-in themes
(PR said 7). Also document the v16/v18/v20 data migrations added since.
2026-07-16 04:47:40 -07:00
kocaemre
a710becd6c docs: fix 25 documentation/code inconsistencies (audit round 3)
Cross-checked website/docs against the source at main HEAD and corrected
documented commands, env vars, config keys, headers, and default values
that don't match the code. Docs-only; no behavioral changes.

Refs #36048

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 04:47:40 -07:00
Teknium
289cc47631 docs: resync reference, user-guide, developer-guide, and messaging pages against code (#17738)
Broad drift audit against origin/main (b52b63396).

Reference pages (most user-visible drift):
- slash-commands: add /busy, /curator, /footer, /indicator, /redraw, /steer
  that were missing; drop non-existent /terminal-setup; fix /q footnote
  (resolves to /queue, not /quit); extend CLI-only list with all 24
  CLI-only commands in the registry
- cli-commands: add dedicated sections for hermes curator / fallback /
  hooks (new subcommands not previously documented); remove stale
  hermes honcho standalone section (the plugin registers dynamically
  via hermes memory); list curator/fallback/hooks in top-level table;
  fix completion to include fish
- toolsets-reference: document the real 52-toolset count; split browser
  vs browser-cdp; add discord / discord_admin / spotify / yuanbao;
  correct hermes-cli tool count from 36 to 38; fix misleading claim
  that hermes-homeassistant adds tools (it's identical to hermes-cli)
- tools-reference: bump tool count 55 -> 68; add 7 Spotify, 5 Yuanbao,
  2 Discord toolsets; move browser_cdp/browser_dialog to their own
  browser-cdp toolset section
- environment-variables: add 40+ user-facing HERMES_* vars that were
  undocumented (--yolo, --accept-hooks, --ignore-*, inference model
  override, agent/stream/checkpoint timeouts, OAuth trace, per-platform
  batch tuning for Telegram/Discord/Matrix/Feishu/WeCom, cron knobs,
  gateway restart/connect timeouts); dedupe the Cron Scheduler section;
  replace stale QQ_SANDBOX with QQ_PORTAL_HOST

User-guide (top level):
- cli.md: compression preserves last 20 turns, not 4 (protect_last_n: 20)
- configuration.md: display.platforms is the canonical per-platform
  override key; tool_progress_overrides is deprecated and auto-migrated
- profiles.md: model.default is the config key, not model.model
- sessions.md: CLI/TUI session IDs use 6-char hex, gateway uses 8
- checkpoints-and-rollback.md: destructive-command list now matches
  _DESTRUCTIVE_PATTERNS (adds rmdir, cp, install, dd)
- docker.md: the container runs as non-root hermes (UID 10000) via
  gosu; fix install command (uv pip); add missing --insecure on the
  dashboard compose example (required for non-loopback bind)
- security.md: systemctl danger pattern also matches 'restart'
- index.md: built-in tool count 47 -> 68
- integrations/index.md: 6 STT providers, 8 memory providers
- integrations/providers.md: drop fictional dashscope/qwen aliases

Features:
- overview.md: 9 image models (not 8), 9 TTS providers (not 5),
  8 memory providers (Supermemory was missing)
- tool-gateway.md: 9 image models
- tools.md: extend common-toolsets list with search / messaging /
  spotify / discord / debugging / safe
- fallback-providers.md: add 6 real providers from PROVIDER_REGISTRY
  (lmstudio, kimi-coding-cn, stepfun, alibaba-coding-plan,
  tencent-tokenhub, azure-foundry)
- plugins.md: Available Hooks table now includes on_session_finalize,
  on_session_reset, subagent_stop
- built-in-plugins.md: add the 7 bundled plugins the page didn't
  mention (spotify, google_meet, three image_gen providers, two
  dashboard examples)
- web-dashboard.md: add --insecure and --tui flags
- cron.md: hermes cron create takes positional schedule/prompt, not
  flags

Messaging:
- telegram.md: TELEGRAM_WEBHOOK_SECRET is now REQUIRED when
  TELEGRAM_WEBHOOK_URL is set (gateway refuses to start without it
  per GHSA-3vpc-7q5r-276h). Biggest user-visible drift in the batch.
- discord.md: HERMES_DISCORD_TEXT_BATCH_SPLIT_DELAY_SECONDS default
  is 2.0, not 0.1
- dingtalk.md: document DINGTALK_REQUIRE_MENTION /
  FREE_RESPONSE_CHATS / MENTION_PATTERNS / HOME_CHANNEL /
  ALLOW_ALL_USERS that the adapter supports
- bluebubbles.md: drop fictional BLUEBUBBLES_SEND_READ_RECEIPTS env
  var; the setting lives in platforms.bluebubbles.extra only
- qqbot.md: drop dead QQ_SANDBOX; add real QQ_PORTAL_HOST and
  QQ_GROUP_ALLOWED_USERS
- wecom-callback.md: replace 'hermes gateway start' (service-only)
  with 'hermes gateway' for first-time setup

Developer-guide:
- architecture.md: refresh tool/toolset counts (61/52), terminal
  backend count (7), line counts for run_agent.py (~13.7k), cli.py
  (~11.5k), main.py (~10.4k), setup.py (~3.5k), gateway/run.py
  (~12.2k), mcp_tool.py (~3.1k); add yuanbao adapter, bump platform
  adapter count 18 -> 20
- agent-loop.md: run_agent.py line count 10.7k -> 13.7k
- tools-runtime.md: add vercel_sandbox backend
- adding-tools.md: remove stale 'Discovery import added to
  model_tools.py' checklist item (registry auto-discovery)
- adding-platform-adapters.md: mark send_typing / get_chat_info as
  concrete base methods; only connect/disconnect/send are abstract
- acp-internals.md: ACP sessions now persist to SessionDB
  (~/.hermes/state.db); acp.run_agent call uses
  use_unstable_protocol=True
- cron-internals.md: gateway runs scheduler in a dedicated background
  thread via _start_cron_ticker, not on a maintenance cycle; locking
  is cross-process via fcntl.flock (Unix) / msvcrt.locking (Windows)
- gateway-internals.md: gateway/run.py ~12k lines
- provider-runtime.md: cron DOES support fallback (run_job reads
  fallback_providers from config)
- session-storage.md: SCHEMA_VERSION = 11 (not 9); add migrations
  10 and 11 (trigram FTS, inline-mode FTS5 re-index); add
  api_call_count column to Sessions DDL; document messages_fts_trigram
  and state_meta in the architecture tree
- context-compression-and-caching.md: remove the obsolete 'context
  pressure warnings' section (warnings were removed for causing
  models to give up early)
- context-engine-plugin.md: compress() signature now includes
  focus_topic param
- extending-the-cli.md: _build_tui_layout_children signature now
  includes model_picker_widget; add to default layout

Also fixed three pre-existing broken links/anchors the build warned
about (docker.md -> api-server.md, yuanbao.md -> cron-jobs.md and
tips#background-tasks, nix-setup.md -> #container-aware-cli).

Regenerated per-skill pages via website/scripts/generate-skill-docs.py
so catalog tables and sidebar are consistent with current SKILL.md
frontmatter.

docusaurus build: clean, no broken links or anchors.
2026-04-29 20:55:59 -07:00
nerijusas
81e01f6ee9 fix(agent): preserve Codex message items for replay 2026-04-25 18:22:06 -07:00
Teknium
5b0243e6ad docs: deep quality pass — expand 10 thin pages, fix specific issues (#4134)
Developer guide stubs expanded to full documentation:
- trajectory-format.md: 56→233 lines (JSONL format, ShareGPT example,
  normalization rules, reasoning markup, replay code)
- session-storage.md: 66→388 lines (SQLite schema, migration table,
  FTS5 search syntax, lineage queries, Python API examples)
- context-compression-and-caching.md: 72→321 lines (dual compression
  system, config defaults, 4-phase algorithm, before/after example,
  prompt caching mechanics, cache-aware patterns)
- tools-runtime.md: 65→246 lines (registry API, dispatch flow,
  availability checking, error wrapping, approval flow)
- prompt-assembly.md: 89→246 lines (concrete assembled prompt example,
  SOUL.md injection, context file discovery table)

User-facing pages expanded:
- docker.md: 62→224 lines (volumes, env forwarding, docker-compose,
  resource limits, troubleshooting)
- updating.md: 79→167 lines (update behavior, version checking,
  rollback instructions, Nix users)
- skins.md: 80→206 lines (all color/spinner/branding keys, built-in
  skin descriptions, full custom skin YAML template)

Hub pages improved:
- integrations/index.md: 25→82 lines (web search backends table,
  TTS/browser providers, quick config example)
- features/overview.md: added Integrations section with 6 missing links

Specific fixes:
- configuration.md: removed duplicate Gateway Streaming section
- mcp.md: removed internal "PR work" language
- plugins.md: added inline minimal plugin example (self-contained)

13 files changed, ~1700 lines added. Docusaurus build verified clean.
2026-03-30 20:30:11 -07:00
teknium1
d87a1615ce docs: add ACP and internal systems implementation guides
- add ACP user and developer docs covering setup, lifecycle, callbacks,
  permissions, tool rendering, and runtime behavior
- add developer guides for agent loop, provider runtime resolution,
  prompt assembly, context caching/compression, gateway internals,
  session storage, tools runtime, trajectories, and cron internals
- refresh architecture, quickstart, installation, CLI reference, and
  environments docs to link the new implementation pages and ACP support
2026-03-14 00:29:48 -07:00