Files
hermes-agent/tests/agent/test_oneshot_footprint.py
teknium1 a79d1d3a71 feat: one-shot runs drop the self-improvement footprint (no skill authoring, fewer process skills, delegation cap)
A finite `hermes chat -q` / `--oneshot` run has no later session in its HERMES_HOME
to learn for, yet it ran the full interactive self-improvement loop. Measured over
21 one-shot benchmark trajectories: 7 skills created and a bundled one patched
mid-task, 37 of ~215 tool calls on skill_view/skill_manage, skill text = 34% of all
tool-result bytes fed back into context, plus reviewer subagents spawned on the
agent's own diff (one task: 5 delegations, 62 subagent API calls, each re-paying a
cold system prompt).

Keyed on the existing HERMES_SINGLE_QUERY_SESSION marker (approval gate, delegation
dispatcher), so interactive and gateway sessions are byte-identical:

* agent/oneshot_footprint.py (new sibling): skill_manage is pruned from the tool
  set; the ## Skills block keeps the index + skill_view but drops the record/patch/
  offer-to-save coaching and the "load process skills for work you already know"
  push (SKILLS_GUIDANCE follows because it is gated on skill_manage).
* delegation.oneshot_max_children (default 2, 0 = unlimited): total children a
  one-shot run may spawn; past it delegate_task returns a tool error telling the
  model to finish inline.
* requesting-code-review skill: reviewer/fixer subagents (Steps 5 and 7) are
  interactive-only; one-shot applies the checklist inline.

Live: `hermes chat -q "list tools starting with skill_"` on the portal — base
"skill_manage, skill_view, skills_list", fix "skill_view, skills_list"; a real
AIAgent under a temp HERMES_HOME shows the interactive prompt unchanged (6643 chars
both) and the one-shot prompt without skill_manage / offer-to-save.
2026-09-19 09:21:18 -07:00

66 lines
3.1 KiB
Python

"""One-shot sessions (``hermes chat -q``) drop the self-improvement footprint.
Across 21 one-shot benchmark trajectories the agent created 7 skills and patched a bundled one
mid-task, spent 37 of ~215 tool calls on skill_view/skill_manage, and spawned review subagents of
its own work. None of that has a consumer in a finite run. The marker is the same
``HERMES_SINGLE_QUERY_SESSION`` the approval gate and delegation dispatcher read, so an interactive
session — the control in every test here — keeps the full surface.
"""
import pytest
from agent import oneshot_footprint
from agent.prompt_builder import build_skills_system_prompt
@pytest.fixture
def oneshot(monkeypatch):
monkeypatch.setenv("HERMES_SINGLE_QUERY_SESSION", "1")
def _tools(*names):
return [{"type": "function", "function": {"name": n}} for n in names]
def test_oneshot_hides_skill_manage_and_skill_authoring_coaching(oneshot, interactive_prompt, tmp_path):
"""-q: no skill_manage tool and a skills prompt that neither asks to save/patch skills nor pushes process
skills; skill reading stays. The interactive prompt for the same skills dir is the control."""
kept = {t["function"]["name"] for t in oneshot_footprint.prune_oneshot_tools(
_tools("skill_manage", "skill_view", "skills_list", "terminal"))}
assert "skill_manage" not in kept and {"skill_view", "skills_list", "terminal"} <= kept
prompt = build_skills_system_prompt(available_tools={"skill_view", "skills_list"}, skills_dir_override=_skills_dir(tmp_path))
assert "demo-skill" in prompt and "skill_view" in prompt
assert "skill_manage" not in prompt and "offer to save as a skill" not in prompt
assert "skill_manage" in interactive_prompt and "offer to save as a skill" in interactive_prompt
def _skills_dir(tmp_path):
d = tmp_path / "skills" / "misc" / "demo-skill"
d.mkdir(parents=True, exist_ok=True)
(d / "SKILL.md").write_text("---\nname: demo-skill\ndescription: Demo skill for tests.\n---\n# Demo\n", encoding="utf-8")
return tmp_path / "skills"
@pytest.fixture
def interactive_prompt(tmp_path, monkeypatch):
monkeypatch.delenv("HERMES_SINGLE_QUERY_SESSION", raising=False)
prompt = build_skills_system_prompt(available_tools={"skill_view", "skills_list", "skill_manage"},
skills_dir_override=_skills_dir(tmp_path))
monkeypatch.setenv("HERMES_SINGLE_QUERY_SESSION", "1")
return prompt
def test_oneshot_delegation_budget_charges_total_children_then_refuses(oneshot, monkeypatch):
from tools import delegate_tool
monkeypatch.setattr(delegate_tool, "_get_oneshot_max_children", lambda: 2)
parent = type("P", (), {})()
assert delegate_tool._oneshot_spawn_budget(parent, 1) is None
assert delegate_tool._oneshot_spawn_budget(parent, 1) is None
err = delegate_tool._oneshot_spawn_budget(parent, 1)
assert err and "oneshot_max_children" in err
# Interactive sessions are never charged, whatever the count.
monkeypatch.delenv("HERMES_SINGLE_QUERY_SESSION")
assert delegate_tool._oneshot_spawn_budget(parent, 50) is None