Commit Graph

41 Commits

Author SHA1 Message Date
kshitijk4poor
41303b379e refactor(sessions): export guard always counts every row; export_all reuses _active_clause
assert_export_safe grew an include_inactive flag whose only production caller
(the console export guard) always passes True, leaving the live-only branch
and its default unused. Drop the parameter and count every row, which is what
the transfer export materializes; the console call and docstring follow. The
existing guard tests seed live rows only and are unchanged.

export_all re-derived the live-row clause by hand; use the existing
_active_clause(include_inactive, False) so "live row" stays defined in one
place. Behaviour is unchanged.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-26 22:27:13 +05:30
kshitijk4poor
deb7bc7343 fix(sessions): import coerces archived-row flags and splits live rows once
A transfer row's active/compacted flags were read by truthiness, so a
hand-edited or foreign JSONL row with "active": "0" (truthy string) imported
live - putting compaction-archived turns back into model context - and
"active": null archived a live row. Coerce both flags with the same
_coerce_or(..., int, default) used for the int session columns: a missing,
null, or unparsable active flag means live (older exports carry none), and
compacted defaults to 0.

The same pass partitions rows into live and archived up front, replacing the
id()-set complement, the dead "_row_id" guard (_insert_message_rows always
sets it), and the subtract-and-reparse counter fixup: message_count and
tool_call_count now come straight from the live rows, with identical values
for well-formed exports. The kept round-trip test re-imports its payload with
stringified/null flags and asserts every row keeps its state.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-26 22:27:13 +05:30
joaomarcos
85deb4553a fix(sessions): transfer export round-trips compaction-archived turns
The console `sessions export` backup went through export_session /
export_all, whose batched read filters `AND active = 1`, so an
export-then-import restore silently dropped every compaction-archived
turn (display 9 -> 3 in the C13 probe). export_all now takes
include_inactive (storage-order rows with their active/compacted flags,
which import_sessions restores as archived), the console export passes
it for both single-session and all-session exports, and
assert_export_safe sizes the same projection it guards.

Ported from #123267's transfer/backup hunks; its import-normalisation
hunks are superseded by #122680's checkpoint-correct import.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-26 22:27:13 +05:30
John Paul Soliva
0325349f96 fix(sessions): archived rows stay out of the import's checkpoint pruning
import_sessions inserted every row live and pruned shadowed checkpoints before restoring the archived flags. A newer archived carrier (a rewound turn) then stripped the newest live checkpoint, and every archived row lost its own. Rows written before #102374 pruning each carry one, so adopting such a donor fed the model a context without its checkpoint while the donor was retired as unrecoverable.

The import now prunes once the archived rows are archived again, keyed on the live rows only; every other writer still prunes on insert. The blank line at the end of the adoption test file is dropped.

(cherry picked from commit f9ad62b37517cbbdc27f75e9e58900048258bd7b)
2026-09-26 22:27:13 +05:30
John Paul Soliva
0af6b2121e fix(sessions): stranded-session adoption carries compacted history before retiring the donor
session.resume with a `profile` for an id that profile's store lacks
adopts the lineage from the launch store and retires the donor with
end_reason=adopted_by_profile, which is not recoverable.

Adoption copied live rows only. In-place compaction (the default)
archives earlier turns under the same id (active=0, compacted=1), so
those turns reached neither the profile copy nor any live row: a
15-turn chat with one compaction showed 30 messages before and 22
after, with the rest behind the unrecoverable archive. The divergence
guards compared live counts, so they could not see the loss.

Adoption now exports every row with its active/compacted flags
(export_session*/_lineage gain include_inactive), and import_sessions
restores a row exported archived as archived, so the copy shows the
same display history and feeds the model the same live context. Its
counters count live rows only. Both divergence guards compare all rows,
so a donor is never retired while any of its rows is missing here.

fe8b643db6 kept adoption on live rows because import_sessions inserted
every row as live context. That still holds for the display projection
(include_compacted), which the JSON export, /save json and the
dashboard keep excluding; adoption uses the new storage-order export.

(cherry picked from commit c1ae532f973277a9a10da79ada3534e827a492dc)
2026-09-26 22:27:13 +05:30
John Paul Soliva
fe8b643db6 fix(sessions): transcript exports keep compacted turns, so --delete-after-verified no longer deletes them unseen
In-place compaction is the default. It soft-archives every earlier row of
a session under the same id (active = 0, compacted = 1), and Desktop and
the dashboard still show those turns. The md/qmd export read the session
through export_session -> get_messages with the default live-only clause,
so it wrote only the compaction summary and the carried tail.
verify_export_file then compared the file with that same dict, and
delete_session removed every row of the session, including the archived
turns that never reached the file.

The md/qmd export now reads the display history (include_compacted), for
a single session and for --lineage logical. export_session and
export_session_lineage take include_compacted, off by default:
import_sessions inserts every message as live context, so the JSON export
and stranded-session adoption keep reading live rows only.

The other transcripts people read had the same hole without the delete:
`/save md` and `/save html` (CLI and gateway), `sessions export --format
html` (one session or all of them) and `--only user-prompts` each held 1
of 6 answers on a six-turn session after one compaction. They now read
the display history too (SAVE_TRANSCRIPT_FORMATS; export_all gains
include_compacted and reads per session then, since the display read
dedupes per session). `/save json`, JSONL and the dashboard's JSON export
stay live-only for the import reason above.

Before deleting, the verify step also re-counts the store's display rows
for every session the file covers and refuses on a mismatch. A message
that lands while the files are written, or a later export change that
reads a narrower view, now refuses the delete instead of being removed
unseen. Like the adoption retire loop, the re-count runs just before
delete_session, not inside its transaction. Rewind rows (undone turns,
the superseded originals of a carried tail) are still deleted without
being exported, as `hermes sessions delete` does: they are not part of
the history the session shows.

Measured through the real CLI on a session with 6 turns and one default
in-place compaction (15 rows, 13 shown): before, 3 messages were exported
and all 15 rows deleted, with answers 1-5 missing from the file; after,
13 messages are exported in display order, then deleted.

(cherry picked from commit adeaff1e33f1ae2b8a066ef2374bb29ade50b3b8)
2026-09-23 22:22:32 +05:30
teknium1
d566442b56 fix(session-export): lineage timings span the merged messages; import ignores derived size
export_session_lineage spread segments[-1] over the merged dict, so the top-level `timings`
described only the last compression segment while `messages` spanned the whole lineage — a
reader would see a 500 ms wall clock over a lineage that ran for hours. Compute the block over
the merged message list (segments keep their own).

_validate_import_session measured the raw session JSON, so the derived `timings.intervals`
(one entry per message pair) counted toward the 5 MiB per-session limit and could reject a
long lineage whose actual content fits. Strip `timings` before measuring; it is rebuilt from
the messages on the next export anyway.
2026-09-15 03:51:07 -07:00
teknium1
2a92d6be4d fix: session-export timings tolerate corrupt timestamp cells
`_export_timings` parsed message timestamps with a bare `float()`, so a
pre-existing out-of-range cell (e.g. 8.4e252) reached
`datetime.fromtimestamp` and raised OverflowError, killing the whole
`export_session`/`export_all` call — a regression caught by
tests/hermes_state/test_corrupt_row_robustness.py. Route the cell through
`coerce_epoch`, the reader every other timestamp surface already uses, so a
bad row degrades to a `missing` count (with the usual warning naming the
session) and the export still lands.
2026-09-15 03:51:07 -07:00
Teknium
f25f655559 feat(session-export): include timing evidence
Port from nearai/ironclaw#7735 by deriving a text-free timing summary from persisted Hermes message timestamps in session exports.
2026-09-15 03:51:07 -07:00
teknium1
d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
teknium1
991b23ad8d refactor(sessions): one session-id minter; QQ update-prompt key from build_session_key
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.

Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".

gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.

Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
2026-09-13 05:21:02 -07:00
teknium1
496eb13bd7 fix(state): one corrupt timestamp row no longer kills sessions list, export or insights
SQLite dynamic typing lets a TEXT cell ('not-a-timestamp'), inf/nan or a
garbage double (8.4e252 salvaged from a damaged page) sit in a REAL
timestamp column. Every reader called datetime.fromtimestamp()/float
arithmetic on the raw cell, so ONE bad row raised TypeError/OverflowError
out of the row loop and took down the whole `hermes sessions list`/browse
table (#102399), all three exporters — JSONL/MD, QMD, HTML (#102352) —
and `hermes insights` (#99959).

Fix the class with ONE helper, hermes_cli.timefmt.coerce_epoch(): a
stored cell becomes float epoch seconds inside a sane 1970..2103 window
or None after a WARNING that names the session id. Every reader routes
through it — relative_time (list/browse/resume picker), format_epoch
(prune/candidates tables), the three exporters' timestamp formatters,
insights' _get_sessions/_day/period range — so a bad row renders as
'?'/'N/A'/raw text for that one cell and the command completes.

Write side: hermes_state_messages._coerce_timestamp (append_message,
append_messages_batch, import) and the import path's started_at now use
the same window, so a new out-of-range timestamp falls back to now()
instead of being persisted — new bad rows cannot be written by Hermes.

Reported-by: #102399, #102352, #99959 reporters; kokhlo's insights
analysis pointed at every reporting site, not just line 860.
2026-09-11 06:24:54 -07:00
Xipong
78f85112d9 perf(state): batch message hydration in export_all 2026-09-11 06:24:54 -07:00
Adolanium
9186e3ebc5 feat(desktop): session import view for foreign coding-agent transcripts
Browse the backend host's foreign CLI session logs, preview a bounded read-only
transcript, and continue a copy in Hermes under the selected profile. Reuses the
hermes_cli.foreign_sessions parsers and the portability validator/writer;
imports are transactional and deduplicated on the recorded origin.
2026-09-06 09:09:41 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
cb6cc64700 refactor(state): search/schema/registry/portability/telegram/usage/titles — inline single-use helpers, contextlib.suppress ladders, pack wrappers around unchanged SQL literals 2026-09-02 21:57:21 -07:00
Teknium
2c2ced0c52 refactor(hermes_state_portability): per-session import validator raising ValueError (fuzz-verified parity) 2026-09-02 19:13:44 -07:00
Teknium
c1620901ae refactor(hermes_state): AST-neutral packing of state mixins (120 cols) 2026-09-02 19:10:31 -07:00
Teknium
5e8931b356 refactor(hermes_state_portability): alias public rich-row getter, reuse safe_json_loads, flatten coercers 2026-09-02 18:59:38 -07:00
Teknium
d2614f435e refactor(hermes_state): AST-neutral line packing across state mixins 2026-09-02 18:30:39 -07:00
Teknium
410f9a0fa6 refactor(hermes_state_portability): fold coerce wrappers, shared _rich_rows, compact docstrings 2026-09-02 18:21:56 -07:00
Teknium
d86ebe66c0 refactor(state): CJK range table; shared _coerce_or for import numeric coercion 2026-09-02 16:29:12 -07:00
Teknium
a138533157 refactor(state): fold docstring closers (whitespace only) 2026-09-02 16:21:39 -07:00
Teknium
5340109e61 refactor(state): compact verbose docstrings (rules/invariants kept) 2026-09-02 16:14:33 -07:00
Teknium
95714f6d93 refactor(state): AST-neutral line packing; derive v22 session_model_usage DDL from the heal DDL 2026-09-02 16:13:05 -07:00
Teknium
3530e4e024 refactor(state): split import_sessions into validate/insert/attach-parents helpers; shared rich SELECT builder 2026-09-02 16:04:59 -07:00
nftpoetrist
c7429f60ca fix(state): close the SessionDB lock gate's blind spot on its own mixin files
test_no_locked_readers_gate.py (#97676) parses hermes_state.py's SessionDB
class body with ast and flags any method that holds the writer lock
around a pure-read query — Pattern C, where every concurrent turn's
persistence convoys behind an unrelated read. #97676 converted 39 such
methods and closed detection blind spots for alias/variable-SQL readers.

But SessionDB is declared as
`class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)`,
and the gate only ever opened hermes_state.py — it never parsed the three
mixin files those base classes are defined in, so a locked reader
declared there was structurally invisible to the scanner regardless of
how good the alias/variable-SQL detection got.

Applying the gate's exact scanning logic to the three mixin files
directly turns up 9 genuine pure-read methods still holding the writer
lock, none in #97676's converted list:

- hermes_state_search.py: _fts_teardown_trash_step, fts_optimize_available,
  optimize_fts_storage, list_recent_user_messages
- hermes_state_portability.py: distinct_session_cwds, list_cron_job_runs,
  _get_session_rich_rows_batch, list_skill_scaffolded_sessions,
  get_first_assistant_text

_get_session_rich_rows_batch is a hot path: it backs list_sessions_rich's
compression-tip resolution and the web server's session-search hydration
across every gateway install — its own docstring already claims "same
read-your-writes guarantee as list_sessions_rich", but list_sessions_rich
was already using _read_ctx() (its guarantee comes from flush_token_counts()
before the read, not from holding the writer lock) while this method's
implementation never caught up to match.

Converted all 9 to `with self._read_ctx() as conn:`, the exact pattern
#97676 used, verified each is a genuine pure read with no hidden writes
by tracing every helper call it makes.

Extended the gate itself (_ALL_STATE_SOURCES) to scan all three mixin
files under their own class names, plus hermes_state.py, so this blind
spot can't silently reopen. Added a regression test
(test_scan_all_state_sources_visits_every_mixin_file) that plants a
synthetic violation in a mixin-shaped file and asserts the scanner still
catches it — a change that reverts the file list back to one file passes
the existing sabotage test but fails this one.

Mutation-verified: with the gate's new scope but the old (unconverted)
mixin sources, test_no_locked_pure_readers fails and names all 9 real
violations with correct file/line. Restored the fix; it passes clean.
2026-09-03 04:29:06 +05:30
Teknium
d15c61b5dc refactor(state): split SessionDB into domain mixins and free-function modules; unify SQL boilerplate
hermes_state.py 17,220 -> 6,442 LOC. Behavior-neutral: every moved body is
AST-identical to the original, verified per extraction.

SessionDB core
- _write_sql / _write_rowcount / _read_one / _read_all replace ~120 copies of
  the `def _do(conn): conn.execute(...)` + `_execute_write(_do)` and
  `with self._read_ctx() as conn: row = conn.execute(...).fetchone()` shapes.
- _set_lineage_column replaces four copies of the recursive compression-lineage
  UPDATE (archived / pinned / hidden / last_read_at).
- _read_session_number unifies the three compression counter readers.
- Dead (zero refs repo-wide): restore_rewound, delete_gateway_routing_entries,
  _is_duplicate_replayed_user_message, SessionPortabilityMixin.get_first_assistant_text.

New mixins bound onto SessionDB via the MRO (logger name stays "hermes_state"):
  hermes_state_messages    SessionMessagesMixin       48 methods
  hermes_state_compression SessionCompressionMixin    30
  hermes_state_gateway     SessionGatewayMixin        26
  hermes_state_maintenance SessionMaintenanceMixin    13
  hermes_state_usage       SessionUsageMixin          12
  hermes_state_titles      SessionTitlesMixin         13
  hermes_state_telegram    SessionTelegramTopicsMixin 11
Origin-internal symbols resolve through a lazy `from hermes_state import ...`
inside the few methods that need them (no import cycle).

New free-function modules, every name re-imported into hermes_state so
`hermes_state.<name>` (and test monkeypatches on it) keep working; intra-module
calls to patched helpers go through the lazy origin import:
  hermes_state_repair   repair/backup/preflight (43 defs)
  hermes_state_wal      journal-mode / PRAGMA policy (33 defs)
  hermes_state_dbfile   header probes, zeroed-db quarantine, stats, holders (21 defs)

Existing mixins: search — shared FTS MATCH/LIKE builders, unified rebuild
status/step/finish engines, state_meta helpers; schema — one legacy/v23 FTS init
branch, shared _live_pk_columns, Row/tuple dual access dropped; portability —
shared _PREVIEW_RAW_SUBQUERY_SQL and _rich_row; common — single
stat_db_file_identity (was 3 copies), AUTO_VACUUM_MIN_FREELIST_RATIO.

Docstrings/comments hand-compacted (AST-identical) keeping every invariant,
ordering rule, failure mode and WHY. Schema SQL, migration order and PRAGMAs
untouched. test_repair_path_has_no_bare_connects repointed to hermes_state_repair.
2026-09-02 13:32:13 -07:00
kshitijk4poor
dc50f02090 fix(adoption): re-check donor growth at retire time, not just at export
Closes the TOCTOU window flagged in review on #93369 (merged via
#93430): the divergence guard compared EXPORT-TIME message counts, but
another backend can append donor messages between the export snapshot
and the retire loop — that growth would be stamped behind the
non-recoverable adopted_by_profile archive, the exact H2 class the
guard exists to prevent, just via a narrower race.

The retire loop now re-reads live donor vs local counts immediately
before end_session and leaves the donor unretired (donor_retired=False,
warn-logged) on any donor-ahead signal; the next resume's export-time
guard then handles the divergence normally. Equal-count CONTENT
divergence (donor rewind+rewrite) remains invisible to count comparison
— documented as accepted: bytes stay in the donor store either way.

New red-first-verified regression simulates the exact race by appending
to the donor from inside an export_session_lineage wrapper.
adoption+ownership suites: 25 passed; ruff clean.
2026-08-24 15:07:18 +05:30
kshitijk4poor
80b202f53a harden(adoption): review findings — exact-id donors only, divergence guard, honest donor_retired
Review batch (3 reviewers) on the final diff surfaced:
- H1: title-based donor matching could adopt AND non-recoverably retire
  an UNRELATED default-store conversation (bot titles collide by design;
  get_session_by_title has no archived filter/ordering). Donor probe is
  now exact-id only — the stranded repro always has the id.
- H2: re-adoption after a partial run could retire a donor that had
  accumulated NEWER messages than the profile copy (skip-based
  idempotency never merges). New divergence guard compares message
  counts and refuses retirement when the donor is ahead (still adopts).
- M1: donor_retired reported True even when every retirement step
  failed under suppress. Now per-segment tracked + warn-logged;
  True only when all applied.
- M3: adopted=False (e.g. import validation limits) was silent — now
  warn-logged with import errors.
- M4: archived donors are never re-adopted (no cross-profile cloning).
- Dead 'from pathlib import Path' dropped; contextlib no longer needed.

5 new red-first-verified regressions (title-collision immunity,
archived-donor immunity, non-vacuous owns_db gating with a real donor
seeded, divergent-donor retirement refusal, donor_retired truthfulness).
tests/tui_gateway: 578 passed. ruff clean.
2026-08-23 18:58:40 -07:00
kshitijk4poor
26a4f89ada fix(gateway): adopt stranded bot sessions from the default store on profile resume
Pre-#93296, the desktop routed session RPCs by the focused tile, so a
profile bot's turns executed on the default backend and its canonical
session accumulated in the DEFAULT profile's state.db. Post-fix, the
profile backend correctly receives the resume — but its store has never
seen the session, so the same chat 4001s forever (unreachable instead
of misrouted). Live repro: Teknium's Developer bot, session c93770.

- hermes_state_portability: SessionDB.adopt_session_lineage_from() —
  composes the existing export_session_lineage()/import_sessions()
  primitives; donor rows are archived (never deleted) with
  end_reason=adopted_by_profile, which is deliberately NOT in
  RECOVERABLE_END_REASONS so canonical-lookup resurrection cannot undo
  an adoption. Idempotent (already-present ids skip).
- tui_gateway/methods_session: profile-scoped session.resume falls back
  to adoption from the default store right before the 4007; ids unknown
  to BOTH stores still 4007 exactly as before, and launch-profile
  resumes never consult the fallback.
- tests: 10 new (7 unit on the primitive incl. compression-lineage
  unit adoption + non-resurrectable archive; 3 handler-level through
  server.handle_request incl. the live repro shape); db-ownership
  leak test taught that the shared launch handle probe is by design.

Follow-up to #93296/#93311; part of #93091.
2026-08-23 18:58:40 -07:00
abundantbeing
a2a23a8f7e fix(clients): hide compaction carriers across surfaces 2026-08-21 21:53:11 +05:30
Teknium
26aa12337a feat: /save exports the current session as json, md, or html on all platforms
Rework of salvaged PR #6372 (@ag9920) onto current main:

- /save promoted from CLI-only JSON snapshot to a cross-platform session
  export: `/save [json|md|html] [filename] [redact]` on CLI and every
  gateway platform (sent as a document via adapter.send_document).
- Rendering routes through the existing shared renderers
  (hermes_cli/session_export.py + session_export_html.py) instead of the
  PR's new hermes_state formatter — new helpers normalize_save_format /
  render_session_for_save / default_save_filename are shared by both
  surfaces.
- `redact` arg runs the export through the force-mode secret redaction
  pass (session_export_md.redact_session_data) before writing.
- Gateway handler awaits AsyncSessionDB correctly, sanitizes user-supplied
  filenames with basename, and lands in gateway/slash_commands.py (the
  handlers moved out of gateway/run.py since the PR was authored).
- /export stays profile export (name collision resolved: session export
  lives on /save).
- Slack 50-slash cap curation: /platform moves to the /hermes-only set to
  free a native slot for /save (parity test updated rationale comment).
- Folds in PR #62268 (@briandevans): None title/model coalescing in the
  single-session HTML export.

Closes #4249. Closes #51200.
2026-08-15 02:04:49 -07:00
ag9920
23ad35b7fa feat: Add /export command for session export (Markdown/JSON) to CLI and Gateway 2026-08-15 02:04:49 -07:00
embwl0x
7d066c3c56 fix(state): deduplicate session system prompts 2026-08-03 20:37:17 +05:30
kshitij
7db2827520 refactor(state): chunk the batched tip-row IN clause at 900 ids
Simplify-pass fold: SQLITE_MAX_VARIABLE_NUMBER is 999 on pre-3.32\nbuilds (which the repo still supports — the trigram-availability\nmachinery exists for exactly that class), and limit=10000\nlist_sessions_rich callers exist in web_server. Chunk inside the\nbatch helper — the single choke point — so no call site can overflow.
2026-08-03 17:32:17 +05:30
jasoisjaso
adcdf9dc63 perf(state): batch compression-tip row fetch in list_sessions_rich
list_sessions_rich()'s compression-root projection called
_get_session_rich_row() once per root — a separate single-row query per
compression root on every session-list render. Resolve every tip id
first, then fetch all tip rows in one WHERE id IN (...) query via the
new _get_session_rich_rows_batch().

_get_session_rich_row() is now a thin wrapper over the batch method, so
the enriched SELECT (preview + last_active) lives in exactly one place —
future column changes (e.g. #42196's include_system_prompt) only touch
one query.

get_compression_tip()'s chain walk is untouched; it's a genuine
per-session graph walk with branch/delegate-exclusion and race handling,
and batching it safely is out of scope here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 17:32:17 +05:30
Teknium
038c1ad872 docs(gateway): precise watchdog scope, explicit import-resets-activity contract (review S4)
- Config docs now describe session_stall_timeout precisely: a RECOVERY
  notifier for an in-process AIAgent with an adapter-queued follow-up —
  not a general gateway/session stall detector — with a per-AIAgent scan
  cadence (not globally coordinated per durable session).
- import_sessions documents the deliberate export-includes /
  import-resets asymmetry for the activity fields (no resurrected
  'working' labels on machines where no agent runs), with a regression
  pinning both halves.
- Strip trailing whitespace in contributors/emails/fangliquan@qq.com
  (git diff --check housekeeping).

PR #76354 review, scope/contract items + housekeeping.
2026-08-02 16:16:36 -07:00
fangliquanflq
c2088efe9e feat(gateway): session activity watchdog, stall notify, compress timeout (#72424)
Three mechanisms to detect and notify when gateway sessions stall silently:

1. Mid-turn activity heartbeats stamped to SessionDB so hermes sessions list
   and hermes status show progress during long turns without new message rows.

2. Stall watchdog: when a busy session has pending inbound and the shared
   activity clock is idle past agent.session_stall_timeout (default 300),
   log a WARNING and notify the user once to try /new. Notify-only; does
   not kill the turn.

3. Compaction timeout: fenceless compress_context callers get a progress-aware
   host budget (compression.context_timeout_seconds default 120 idle,
   compression.context_total_ceiling_seconds default 600 ceiling). On timeout,
   cancel via commit fence, skip compaction without dropping messages, and
   continue the turn.

Closes #72016 (slices 1-3; slice 4 cumulative SSE stream-retry deadline
remains a follow-up).

Cherry-picked from PR #72424 by @fangliquanflq.
2026-08-02 16:16:36 -07:00
teknium1
21c7ae8563 refactor: split SessionDB into Search/Schema/Portability mixins (mechanical move, ~2.9K LOC out of hermes_state.py) 2026-07-29 10:14:19 -07:00