Per-finding verdicts from the cross-vendor review of fix/status-fix
(#97655/#97654):
[1] Flash NIT (real, cheap) — FIXED. Added test_error_without_failed_flag_
marks_failed: an error string with the 'failed' key ABSENT (not False)
must still be status=failed + exit_reason=error. The branch order
(result.get('failed') or result.get('error')) already handles this; the
test pins the error-alone path.
[2] GPT-OSS SHOULD-FIX — PINNED. Added test_empty_error_with_summary_is_
completed: error='' is falsy so result.get('error') falls through to the
summary-presence heuristic => status=completed. No code change; the
existing branch is correct and the new test locks it in.
[3] GPT-OSS SHOULD-FIX — VERIFIED, NO CHANGE. Grepped every delegation
exit_reason consumer:
* tools/delegation_live_log.py finalize() prints exit_reason generically
and only special-cases == 'max_iterations' for a readable suffix.
* tools/process_registry.py derives truncated as
(truncated or exit_reason == 'max_iterations') — gated, not exhaustive.
* tools/async_delegation.py passes exit_reason through generically.
The gateway/status.py, cron/scheduler.py and run_agent.py 'exit_reason'
hits are a DIFFERENT field (turn_exit_reason / gateway exit reason), not
the delegation result's exit_reason. No exhaustive if/elif over the enum
missing an 'error' case, so nothing to add.
[4] GPT-OSS NIT — DONE. Enriched _run_single_child's docstring to enumerate
status in {completed, interrupted, failed} and exit_reason in {completed,
max_iterations, interrupted, error}, and added a compact enum comment at
the result-entry construction. Verified the process_registry.py renderer
comment (truncated <= exit_reason == 'max_iterations') still holds — the
truncation flag is derived exactly that way, so no contradiction.
[5] GPT-OSS NIT — REJECTED. The proposed 'fallback for legacy dicts that
explicitly set failed=False' is not adopted. No consumer produces a result
dict with an explicit failed=False and no summary while relying on
completed semantics: run_agent.py sets failed=True only on genuine failure
and omits the key on success (no failed=False producer). Also, the
proposed elif would reintroduce ambiguity (explicit failed=False + no
summary => 'completed'?) and diverge from the conservative else => 'failed'.
result.get('failed') is falsy for both explicit-False and absent, so no
distinction exists to preserve; the else is the correct default.
Tests: 301 passed, 7 skipped (tests/tools -k 'delegate or process_registry').
TestDelegateFailedChildStatus: 6 passed.