Files
hermes-agent/agent/oneshot_footprint.py
teknium1 a79d1d3a71 feat: one-shot runs drop the self-improvement footprint (no skill authoring, fewer process skills, delegation cap)
A finite `hermes chat -q` / `--oneshot` run has no later session in its HERMES_HOME
to learn for, yet it ran the full interactive self-improvement loop. Measured over
21 one-shot benchmark trajectories: 7 skills created and a bundled one patched
mid-task, 37 of ~215 tool calls on skill_view/skill_manage, skill text = 34% of all
tool-result bytes fed back into context, plus reviewer subagents spawned on the
agent's own diff (one task: 5 delegations, 62 subagent API calls, each re-paying a
cold system prompt).

Keyed on the existing HERMES_SINGLE_QUERY_SESSION marker (approval gate, delegation
dispatcher), so interactive and gateway sessions are byte-identical:

* agent/oneshot_footprint.py (new sibling): skill_manage is pruned from the tool
  set; the ## Skills block keeps the index + skill_view but drops the record/patch/
  offer-to-save coaching and the "load process skills for work you already know"
  push (SKILLS_GUIDANCE follows because it is gated on skill_manage).
* delegation.oneshot_max_children (default 2, 0 = unlimited): total children a
  one-shot run may spawn; past it delegate_task returns a tool error telling the
  model to finish inline.
* requesting-code-review skill: reviewer/fixer subagents (Steps 5 and 7) are
  interactive-only; one-shot applies the checklist inline.

Live: `hermes chat -q "list tools starting with skill_"` on the portal — base
"skill_manage, skill_view, skills_list", fix "skill_view, skills_list"; a real
AIAgent under a temp HERMES_HOME shows the interactive prompt unchanged (6643 chars
both) and the one-shot prompt without skill_manage / offer-to-save.
2026-09-19 09:21:18 -07:00

49 lines
2.6 KiB
Python

"""What a finite one-shot session (``hermes chat -q`` / ``--oneshot``, ``hermes -z``) does NOT do.
A one-shot run has no later session in its HERMES_HOME to learn for: the process answers one query and
exits. The interactive self-improvement loop is pure overhead there, and a measured one — across 21
one-shot benchmark trajectories the agent authored 7 new skills and patched a bundled one mid-task, 37 of
~215 tool calls were ``skill_view``/``skill_manage``, and skill text was 34% of every tool-result byte fed
back into context. Three rules follow, all keyed on the same session marker the approval gate and the
delegation dispatcher already read (``HERMES_SINGLE_QUERY_SESSION``), so interactive sessions are untouched:
* ``skill_manage`` is not offered (``skills_list``/``skill_view`` stay: reading a domain skill can still win);
* the ## Skills prompt drops the "record it / patch it / offer to save" coaching and the "load process skills
even for tasks you already know" push, keeping only "load a skill when it adds knowledge you lack";
* delegation is capped per session (``delegation.oneshot_max_children``): subagents each re-pay a cold
system prompt and re-explore the repo, and the observed spawns were mostly "independent review of my own
work" rather than parallel work.
"""
from __future__ import annotations
from typing import Any, Dict, Iterable, List
ONESHOT_HIDDEN_TOOLS = frozenset({"skill_manage"})
def is_single_query_session() -> bool:
"""The finite ``-q`` marker, read through the session env so gateway-bound sessions never see it."""
try:
from gateway.session_context import get_session_env
except Exception:
import os
get_session_env = os.environ.get
return str(get_session_env("HERMES_SINGLE_QUERY_SESSION", "") or "") == "1"
def prune_oneshot_tools(tools: Iterable[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""*tools* minus ``ONESHOT_HIDDEN_TOOLS``; identity when the session is not one-shot."""
tools = list(tools)
if not is_single_query_session():
return tools
return [t for t in tools if (t.get("function") or {}).get("name") not in ONESHOT_HIDDEN_TOOLS]
ONESHOT_SKILLS_LOAD_GUIDANCE = (
"## Skills\n"
"Scan the skills below and load one with skill_view(name) only when it carries domain knowledge you lack "
"for THIS task (an API, a tool's commands, a project's conventions). Do not load general process skills "
"(testing, debugging, review methodology) for work you already know how to do, and do not create or edit "
"skills: this is a one-shot run with no later session to reuse them.\n"
)