18 Commits

Author SHA1 Message Date
Austin Pickett
3d5cf17123 test(agent): execute_code replay streak notice fires on warn-only desktop config 2026-09-27 13:20:28 -05:00
Daniel.lu
0cc43227d9 fix(agent): ignore execute_code per-call metadata in the loop-detector result hash 2026-09-27 13:20:28 -05:00
teknium1
b2ecd3518f test: purge low-value tests, lane py04 (294 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
9814fb9286 fix(guardrails): mark every file-tool loop refusal and exempt it on the live failure seam
The salvaged commit adds the `guardrail_refusal` marker to the read-dedup block
and exempts it in `classify_tool_failure`. That classifier is only the fallback:
the executor (`agent/tool_executor.py::_commit_tool_result`) passes
`failed=_detect_tool_failure(...)` from `agent/display.py` into
`after_call`, so on the real path the refusal was still counted and the
escalation to `repeated_exact_failure_block` still fired (live probe: block at
call 8 of an identical read after 5 refusals, on the salvaged head as on main).

- `agent/tool_result_classification.py`: `GUARDRAIL_REFUSAL_KEY` +
  `is_guardrail_refusal()` so the marker has one definition and both
  classifiers share the predicate.
- `agent/display.py::_detect_tool_failure`: exempt refusals (the live seam).
- `tools/file_tools.py`: the two sibling loop refusals -- the consecutive-read
  BLOCK in `read_file_tool` and the repeated-search BLOCK in `search_tool` --
  carry the same marker; they are emitted via `tool_error` too and escalated
  the same way.
- Tests trimmed to two invariants covering all three refusals on both
  classifiers, plus the negative (a genuine error still counts; a non-boolean
  marker is ignored).

Approval-denied write blocks (`file_tools_write_guards`) deliberately keep
counting: a user's denial retried unchanged is exactly the loop the streak
exists to stop.
2026-09-18 10:03:04 -07:00
rlaope
c962c698f8 fix(guardrails): a harness refusal is not a tool failure
The read-dedup block in `_dedup_stub_or_block` returns its message through
`tool_error`, so the body carries `"error"` for the model's benefit. That is
exactly the substring `classify_tool_failure` keys on, so the guardrail counted
its own refusal as a failed call: refusing a repeated read raised the failure
streak that produces the next, harder refusal, and the resulting
`repeated_exact_failure_block` told the model N identical calls had failed when
none had.

Observed on a real session (55 patches, 73 reads, 13 searches, 14 terminals):
13 of the 14 results the classifier called failures were this block. One tool
genuinely failed -- a `gh` call with exit 8. The model was told to change
strategy over a count the harness had produced itself, and could not diagnose it
because the message did not describe anything that had happened.

The block now carries `guardrail_refusal`, and the classifier exempts a body
marked that way. The exemption is keyed on the marker rather than on wording, so
a body that genuinely failed still counts and the streak that stops a real loop
is unchanged.

Signed-off-by: rlaope <piyrw9754@gmail.com>
(cherry picked from commit b3b09822f3db5e70a181e7dd19725afe1b1583a7)
2026-09-18 10:03:04 -07:00
Teknium
25d954c2cf fix(guardrails): hard stops catch replays, never legitimate iteration
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:

- Edit -> re-run is progress. A successful mutating call (write_file/patch,
  a green terminal/execute_code, browser actions, job/message/cron/memory/
  skill mutations) marks progress for every failing signature still being
  counted this turn; the next identical retry restarts its streak instead
  of accumulating toward exact_failure_block_after. A pure replay never
  mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
  (terminal, execute_code, process pollers, browser_navigate, web_extract)
  same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
  task loops with a live parent/client and do real edit -> re-run work.

Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
  unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
  this commit:        COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
2026-09-02 00:26:57 -07:00
Teknium
76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
benbenwyb
cd2d3089fb fix(agent): guard repeated skill reads
Treat skill_view and skills_list as idempotent read-only tools so the existing no-progress guardrail can warn or block repeated identical skill loads. This prevents large skill outputs from being re-added to the context in tool loops.

Add regression coverage for repeated skill_view results under hard-stop guardrails.
2026-09-02 00:26:57 -07:00
João Vitor Cunha
384fc4bf83 fix(guardrails): preserve interactive platform defaults 2026-09-02 00:26:57 -07:00
João Vitor Cunha
ee2147f9e6 fix: hard stop tool loops on non-interactive platforms 2026-09-02 00:26:57 -07:00
Teknium
6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
teknium1
cb06017b1d refactor(guardrails): make runaway-loop caps per-turn, not session-total
Per Teknium: the caps should bound a single agent loop, not accumulate
over the whole session. Rename SessionCapConfig -> LoopCapConfig and the
config section session_caps -> loop_caps; move the counters into
reset_for_turn (invoked per turn via turn_context) so each turn starts
with a fresh budget; retune defaults 200 -> 50 (a single turn issuing 50
web searches / spawning 50 subagents is already pathological). Block
codes session_*_cap -> loop_*_cap and messages updated to drop the
/new-resets-the-budget guidance (irrelevant now that it resets per turn).
Tests flipped: the old persists-across-turn-resets assertion becomes
resets-each-turn.
2026-07-26 21:27:45 -07:00
teknium1
b68787ad25 Inspired by Claude Code: session-wide runaway-loop caps for web_search and delegate_task
Add per-session lifetime caps on web_search calls and subagent spawns
(defaults 200/200, matching Claude Code v2.1.212). Unlike the existing
per-turn tool-loop guardrails, these count over the whole session and
reset only when a fresh agent is built (/new, /clear). Hitting a cap
blocks the offending call and halts the turn cleanly.

- agent/tool_guardrails.py: SessionCapConfig + session counters on the
  controller (in __init__, not reset_for_turn, so they persist across
  turns). before_call() enforces caps first, independent of
  hard_stop_enabled. delegate_task batches count each task.
- hermes_cli/config.py: tool_loop_guardrails.session_caps defaults.
- docs + tests (unit + E2E validated against a real AIAgent).
2026-07-26 21:27:45 -07:00
Juniper Bevensee
fb0217c656 fix(agent): tolerate lone UTF-16 surrogates in tool-guardrail hashing
Tool results scraped from the web/social platforms can carry unpaired
UTF-16 surrogates (e.g. half of a mathematical-bold character pair).
_sha256() did a strict utf-8 encode, which raises UnicodeEncodeError on
that input and took down the whole conversation loop — the hash only
needs deterministic bytes, not valid UTF-8, so encode with
surrogatepass instead.
2026-07-18 02:08:39 -07:00
ooovenenoso
d759a67c0f fix: add recovery hints to loop guard warnings 2026-05-19 00:12:12 -07:00
GodsBoy
da0ddbf88a fix: classify landed file mutations with diagnostics 2026-05-13 06:46:23 -07:00
Mind-Dragon
0704589ceb fix(agent): make tool loop guardrails warning-first 2026-04-30 20:43:15 -07:00
Mind-Dragon
58b89965c8 fix(agent): add tool-call loop guardrails 2026-04-30 20:43:15 -07:00