Commit Graph

1091 Commits

Author SHA1 Message Date
teknium1
bd46352705 feat: catalog card image slot is 2:1; version pill reads as a pill
repo social banners are 2:1, so that shape fills the slot without a crop;
the recommended size is in the README/docs. Desktop AgentPluginRow gains
catalog_version to match the generated contract.
2026-09-16 14:18:39 -07:00
teknium1
034313e7cd feat: plugin catalog entries carry an optional version label and card image
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).

Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.

Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
2026-09-16 14:18:39 -07:00
brooklyn!
4e9d3c713a feat(desktop): browse catalogs as cards with a saved list option 2026-09-16 13:55:13 -05:00
brooklyn!
8b551a1ab3 docs(catalog): restore native browsing and install guidance 2026-09-16 13:55:13 -05:00
teknium1
9b8cc18364 Revert "docs(catalog): explain shared feeds and Desktop install links"
This reverts commit aea0580eb7.
2026-09-16 09:25:14 -07:00
teknium1
9d67c7c29a Revert "fix(desktop): show skill install progress and completion in the shared dialog"
This reverts commit 25be09822d.
2026-09-16 09:25:14 -07:00
brooklyn!
25be09822d fix(desktop): show skill install progress and completion in the shared dialog 2026-09-16 04:33:07 -05:00
brooklyn!
aea0580eb7 docs(catalog): explain shared feeds and Desktop install links 2026-09-16 04:33:07 -05:00
brooklyn!
fb56a7e06d docs(desktop): the wake-word ear and speaker live in the mic's fan
The desktop guide and the wake-word page told users to click an ear in the
composer row; it now fans out of the microphone on hover. The folded voice
menu, not the ear, carries the silent-mic hint.
2026-09-16 02:56:57 -05:00
teknium1
6e6550e181 docs(kanban): document worker_output on crashed / protocol_violation events 2026-09-15 22:47:03 -07:00
outpoints
c2f743f9a1 fix(honcho): [verified] preserve seeded branch title provenance 2026-09-15 22:30:11 -07:00
outpoints
9e6c79f5be docs(honcho): document workspace routing and provider context 2026-09-15 22:30:11 -07:00
outpoints
5237cab756 fix(honcho): thread logical cwd through agent construction
(cherry picked from commit b1d7207c45311be658592c6ad34ee84634fed0ee)
2026-09-15 22:30:11 -07:00
outpoints
3cbdc32565 fix(honcho): don't let auto-generated session titles override sessionStrategy
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.

Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.

Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.

Fixes #24740

(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
2026-09-15 22:30:11 -07:00
teknium1
3abeca16e6 fix(kanban): 5xx and timeouts requeue the worker instead of spending its retry budget
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
2026-09-15 22:05:38 -07:00
teknium1
23863ccbaf fix(cli): one-shot chat -q exits non-zero on failure; 75 covers upstream 429 and overload
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
2026-09-15 21:46:41 -07:00
teknium1
2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1
9a1b0a7b06 docs(kanban): document the worker exit-code contract (1 failure, 75 quota wall) 2026-09-15 19:28:32 -07:00
teknium1
abdb402701 fix(mcp): carry the lazy status across the TUI wire, tests and docs
Follow-up to the ported status fix:

- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
  closed wire enum; `mcp.servers.status` would raise `ContractViolation`
  on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
  contract files.
- `ui-tui` session panel: an unknown status fell through to the red
  `failed` branch; render `lazy` with its cached tool count (inline
  branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
  yields `status: lazy` with the cached tool count and a summary without
  `failed` (eager control stays `configured`, live control stays
  `connected`); a lazy-only run neither warns nor re-arms the startup
  retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
  `cli-config.yaml.example`, the MCP config reference and the MCP guide.
2026-09-15 19:06:54 -07:00
teknium1
05fac10a75 docs(api-server): MCP trust-gate consent surfaces as approval.request on /v1/runs
Document that an untrusted-server write-capable MCP tool now parks a run in
waiting_for_approval and is resolved through POST /v1/runs/{id}/approval,
the same bridge dangerous-command approvals already use.

Part of #111526
2026-09-15 19:06:27 -07:00
teknium1
ee1bfef857 fix(mcp): NO_PROXY for MCP servers uses the repo matcher; trim tests to two invariants
Follow-up to the salvaged #111796 commit:

- NO_PROXY matching goes through `agent.proxy_bypass.should_bypass_proxy` (the one
  matcher the LLM transport and the gateway adapters already use), so CIDR ranges and
  `*.host` patterns bypass the proxy for MCP servers exactly as they do for the model
  endpoint. The stdlib `proxy_bypass` stays for the OS bypass list (Windows
  ProxyOverride / macOS exceptions). Live probe: NO_PROXY=10.255.255.0/24 still routed
  the MCP request through the proxy before this commit, direct after.
- Drop the try/except around `getproxies()` / `proxy_bypass()`: the stdlib guards its
  own registry/sysconf reads and httpx calls the same functions unguarded.
- Trim the six contributor tests to two invariants (mount + NO_PROXY incl. CIDR; both
  client builders carry mounts next to the body-cap transport). Fixture uses the stdlib
  `getproxies_environment` / `proxy_bypass_environment` instead of a hand-rolled copy and
  skips when the mcp SDK is absent.
- Docs: one sentence on the MCP page about proxy resolution for HTTP/SSE servers.
- contributors/emails mapping for the PR author.
2026-09-15 19:05:57 -07:00
teknium1
037771a692 docs: heartbeat and loop ticks stay with the surface that registered them
Users pointed at the Desktop app as a workaround-breaker ("don't keep the shared session open"); state
the ownership rule so the behaviour is discoverable: a chat-registered heartbeat or /loop is fired by the
gateway and replies into the chat even while a TUI / Desktop viewer has the same session open.
2026-09-15 18:57:09 -07:00
DavidMetcalfe
3e833fd56f docs(browser): cover logins across scheduled and unattended runs
Real-profile browsing and cron's per-job toolsets are both documented, but
nothing connected them: browser.md never mentioned scheduled runs and cron.md
never mentioned login state, so the constraints of driving a login-gated site
from a cron job were only discoverable in source.

Adds a Scheduled and unattended runs subsection to the real-profile section:
the browser.use_real_profile prerequisite (off by default), the credentials a
login form or a fresh 2FA challenge needs saved ahead of time, the auth re-sync
behaviour and its Windows consequence. Plus a pointer from cron.md's toolset
section.
2026-09-15 18:46:54 -07:00
teknium1
6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1
80f76edfaf fix(web): rescue eligibility asks the provider whether the ring was walked
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
2026-09-15 18:43:05 -07:00
teknium1
45a4db2225 fix(web): key extract cache on metadata.sourceURL too, pin redirect case
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.

Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
2026-09-15 18:42:38 -07:00
teknium1
339fa6d918 fix(gateway): bounded redacted result preview on tool.completed run events
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.

The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.

Salvages #111821 (@KoNit-K), part of #111815.

Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
2026-09-15 18:42:10 -07:00
teknium1
4c4d54554d docs(kanban): stop-nudge scope covers delegate children and in-process cron
State that the turn-end guard fires only for the dispatcher-owned worker, not
for delegate_task children or cron runs that inherit HERMES_KANBAN_TASK.
2026-09-15 18:41:19 -07:00
teknium1
996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
teknium1
0f38867a1b docs(dashboard): say lifetime-capping proxies still close the chat socket
The keepalive only defeats idle timeouts; proxies that cap total socket
lifetime (some tunnels) still close it and the chat reattaches on its own.
2026-09-15 18:38:16 -07:00
teknium1
a8a36c461b docs(dashboard): describe the chat PTY keepalive and hidden-tab reconnect pause 2026-09-15 18:38:16 -07:00
teknium1
bd63866253 fix(kanban): give a finished worker a grace window before the terminal reaper signals it
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.

Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
2026-09-15 18:35:32 -07:00
teknium1
aa5817d9be fix(kanban): reap workers that outlive their finished run
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).

Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.

Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.

Fixes #111791
2026-09-15 18:35:32 -07:00
teknium1
0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1
c7f4bc5bd7 fix(kanban): text dispatch output and both "dispatcher stuck" warnings name the hold reason
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.

- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
  task plus rate_limited / skipped_locked / memory_pressure for one or more
  DispatchResults, so the CLI daemon and gateway warnings share one wording:
  `Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
  tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.

Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>

Part of #111910
2026-09-15 18:34:11 -07:00
teknium1
9013fcdc87 fix(cron): an unreadable cron toolset restriction fails the run instead of granting every tool
_resolve_cron_enabled_toolsets returned None when _get_platform_tools
raised, and AIAgent reads None as "load every toolset": a malformed
platform_toolsets block (or a stale-module import error after an update)
turned the operator's cron restriction into the full default set, with
only a log warning. Unattended jobs process untrusted text, so that is a
privilege widening, not a safety net (#111380).

The resolver now raises a RuntimeError naming the cause; run_job's
existing failure path records it on the job (last_error, failure streak,
incident) and the agent is never constructed. Per-job enabled_toolsets
(unknown names included) and the MCP merge path never touch the
platform resolver and are unchanged; the disabled-toolset resolver has no
fail-open branch.

Live: platform_toolsets: oops -> before: run ok, enabled_toolsets=None,
98 tool names selected; after: run fails "Cron toolset resolution
failed, so this run was refused rather than given every tool", agent
never constructed. normal / unknown-per-job / mcp-merge shapes: identical
before and after.

Co-authored-by: Austin Bell <10687162+robertaustinbell@users.noreply.github.com>
2026-09-15 18:31:55 -07:00
teknium1
0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1
c1bbcf9712 fix(agent): hint-preview truncation log names the real remedy; invariant tests
Follow-up to the two salvaged commits (#111777, #111781 by @KoNit-K):

- agent/prompt_builder.py::_truncate_content — with queue_warning=False the
  logged line no longer tells the operator to "pin a larger
  context_file_max_chars, or use a larger-context model": the subdirectory
  hint cap is a constant neither knob raises. It now points at the read_file
  recovery the marker already discloses.
- tests/gateway/test_startup_environment_probe.py — replace the
  call-detection test with the behavioural invariant: an oversized SOUL.md in
  HERMES_HOME and a warm-up leave the truncation-warning queue empty for the
  next default-executor task (the api_server turn path runs on that executor
  without copy_context, which is how the boot warning reached a foreign
  session).
- tests/agent/test_subdirectory_hints.py — fold the new drain assertion into
  the existing oversized-hint test (same fixture) and pin that the log carries
  no context_file_max_chars advice.
- agent/AGENTS.md, website/docs/.../context-files.md — the hint cap is 32,000
  (docs said 8,000) and is fixed; document that it is logged, not surfaced as a
  chat warning.
2026-09-15 18:19:26 -07:00
teknium1
c5c71ea1ad fix(docs): document the 10 s SSE keepalive comment for custom parsers 2026-09-15 18:18:30 -07:00
teknium1
eb562b10ad docs(plugin-catalog): allow maintainer-curated sweep entries alongside owner submissions
Teknium ruled that maintainers may add batches of community plugins from a
reviewed sweep instead of waiting for each owner to submit. Rule 5 and the
user-guide checklist now say so, and give authors the explicit right to adjust
or remove a swept-in entry via their own PR.
2026-09-15 12:55:47 -07:00
teknium1
651168d7b2 docs(credential-pools): env: rows may name any variable, not only the declared one 2026-09-15 11:50:48 -07:00
teknium1
d84ece48b8 fix(mcp): Figma OAuth login completes despite the omitted iss parameter
Figma's authorization-server metadata advertises
authorization_response_iss_parameter_supported and its redirect omits iss,
so the mcp SDK's RFC 9207 check discarded every valid code and login never
finished. For that one issuer the provider fills a missing iss with the
discovered issuer and warns; a mismatching iss still fails and every other
server keeps the strict rule.

Fixes #111135
2026-09-15 09:29:00 -07:00
teknium1
6f24245532 fix(kanban): gate create-with-parents like link; archived parent is terminal
create_task(parents=[open parent]) — the reporter's actual incident path —
parked the card in todo with only a `created` event, and kanban_create's
payload carried no `gated`, so the board still showed an unexplained todo
while only the link surface was fixed. create_task now appends the same
dependency_wait {reason: parent_not_done, parent} event and kanban_create
returns gated/gated_by, mirroring kanban_link.

link_tasks gated on `status != 'done'`, but _parents_satisfied and
recompute_ready treat `archived` as terminal: linking a ready child under an
archived parent demoted it to todo with a false parent_not_done event and the
next recompute promoted it straight back. Gate on not in ('done','archived').

Review finding: create_task(parents=...) emitted no dependency_wait/gated; link under an archived parent flapped ready->todo->ready with a false reason.
2026-09-15 06:25:42 -07:00
teknium1
89ef145254 fix(kanban): dashboard link reports the gate; docs; trim to two invariant tests
The dashboard's POST /links is the fourth writer of link_tasks (CLI, tool,
dashboard, plus the graph builder); return the same ``gated`` flag so every
surface that can create the deadlock can see it. Document the
``dependency_wait`` payload the link path emits and the delegation rule the
reporter derived (never link a support card under the card it unblocks).

Drops the CLI output test (a change-detector on prose); the two DB-level
invariants (event emitted on demotion / none for a done parent) stay.
2026-09-15 06:25:42 -07:00
Konstantin Khlopkov
35b1609fc3 fix(kanban): surface the link-time demotion of a ready child to todo
A ready child linked under an unfinished parent drops to todo with no
event and no operator signal; the only trace used to be claim_rejected
after a forced promote. Record a dependency_wait event when the demotion
fires, return the gate from link_tasks, warn in the CLI link command,
report gated in the kanban_link tool, and document the gate.
2026-09-15 06:25:42 -07:00
teknium1
a919c414ce docs(kanban): worker sessions are named after their card, no title model call 2026-09-15 06:08:20 -07:00
teknium1
738d63a34d fix(kanban): automatic stale-claim reclaims count toward the failure breaker
A claim that expired without a worker ever spawning (worker_pid NULL) was
reclaimed and immediately re-claimed on every dispatcher tick, with
consecutive_failures stuck at 0 — nothing could trip the breaker. Route the
reclaim through _record_task_failure (own txn after the reclaim commit, same
shape as enforce_max_runtime) instead of the salvaged raw counter increment,
so per-task max_retries / kanban.failure_limit and the gave_up event apply
and last_failure_error carries the stale lock. reclaim_task (operator path)
still resets the counter; the live-worker extend path never reaches it.

Trims the salvaged tests to one invariant that walks the breaker to its trip.
2026-09-15 06:06:32 -07:00
teknium1
238720e928 test(journey): one end-to-end invariant for foreground-created skills; docs + list-unmanaged label
Replace the predicate unit test with an invariant on the real builder: a skill recorded
by a foreground create is in build_learning_graph() with zero uses, and an unmarked
never-used local skill is not. `hermes curator list-unmanaged` prints the actual
marker (created_by:learn) instead of hard-coding created_by:null. Docs: curator.md and
memory.md describe the learn marker and what the journey shows.
2026-09-15 05:38:31 -07:00
fangliquan
ebb6dc6e70 fix(stt): prefer cached local whisper models 2026-09-15 05:36:37 -07:00
teknium1
0d1a3e704a docs(vault): manager items fill on every saved website origin 2026-09-15 04:56:01 -07:00