Files
hermes-agent/scripts/toolperf_abeval/run_all.sh
Teknium ced8e30217 feat(scripts): reproducible core-toolset A/B eval harness (toolperf_abeval)
Ships the hard A/B evaluation used for the August 2026 core-toolset
performance batch (#77056) as a reusable harness: 9 error-inducing trap
tasks derived from measured production waste classes, two-arm
PYTHONPATH-only comparison, ATOF-trace-based scoring, resume-safe
batteries.

Hardened from the original one-off: paths de-hardcoded (ABEVAL_ROOT /
ABEVAL_HOME), encoding= on all file IO, startup crashes retry on resume
instead of polluting cells, post-hoc grading fix for err_inline_script
baked in. Live-smoked end to end (baseline arm, qwen3-coder-30b,
err_multi_dir: exit 0, correct on-disk verification, resume record
written).
2026-08-05 13:43:30 -07:00

38 lines
1.4 KiB
Bash
Executable File

#!/bin/bash
# Full A/B eval: N models x 2 arms x 9 tasks x R reps.
#
# Usage:
# ./run_all.sh <baseline-tree> <fixes-tree> [reps] [model ...]
#
# baseline-tree checkout of the code WITHOUT the changes (e.g. a worktree
# of origin/main)
# fixes-tree checkout WITH the changes (e.g. your integration branch)
# reps repetitions per cell (default 3)
# model ... models to test (default: the Aug 2026 pair)
#
# Requires: ABEVAL_HOME pointing at a configured Hermes home (see README.md),
# and this script run with the python that has hermes-agent's deps installed.
set -euo pipefail
BASE=${1:?usage: run_all.sh <baseline-tree> <fixes-tree> [reps] [model ...]}
FIXES=${2:?usage: run_all.sh <baseline-tree> <fixes-tree> [reps] [model ...]}
REPS=${3:-3}
shift $(( $# >= 3 ? 3 : 2 ))
MODELS=("$@")
if [ ${#MODELS[@]} -eq 0 ]; then
MODELS=("anthropic/claude-sonnet-4.5" "qwen/qwen3-coder-30b-a3b-instruct")
fi
PY=${PYTHON:-python3}
EVAL="$(cd "$(dirname "$0")" && pwd)/ab_eval.py"
export ABEVAL_ROOT=${ABEVAL_ROOT:-$PWD/abeval-workspace}
for model in "${MODELS[@]}"; do
echo "=== $model / baseline ==="
"$PY" "$EVAL" run --arm baseline --model "$model" --reps "$REPS" --pythonpath "$BASE"
echo "=== $model / fixes ==="
"$PY" "$EVAL" run --arm fixes --model "$model" --reps "$REPS" --pythonpath "$FIXES"
done
echo "=== ALL RUNS DONE ==="
"$PY" "$EVAL" report --models "$(IFS=,; echo "${MODELS[*]}")"