Files
hermes-agent/tests/gateway/test_api_server_responses_history_dedupe.py
teknium1 d19963782b fix(api-server): anchor the Responses current turn on this turn's user row, not a history prefix match
`_response_messages_turn_start_index` located the current turn by matching
`result["messages"]` against `conversation_history + [user]` row by row. The
loop repairs host-fed history before its first call (consecutive assistant or
user rows merge, orphan tool results drop) and compaction rewrites it, so the
transcript legitimately stops sharing a prefix with the client history; the
match then returned 0 and the WHOLE transcript was treated as the current
turn: earlier turns' function_call / function_call_output items were replayed
as this turn's `output` / `run.completed` turn_messages, and the stored
`previous_response_id` chain grew by another copy of the history every turn.

Anchor on the loop's canonical current-turn user index instead
(`agent.turn_context.reanchor_current_turn_user_idx`: last user row that
says this turn's text, else the last user-originated row), and keep the
semantic prefix match only for suffix-only results without a user row
(mocked/legacy paths). The boundary logic moves to the topical sibling
`api_server_turn_boundary.py`; the routes mixin delegates.

The dedupe test that asserted "divergent transcript => append history +
transcript" encoded the duplication itself; it now asserts the fallback for
the case it actually exists for (a suffix-only result).

Fixes #89891
Co-authored-by: heyf <tonyheyifan@gmail.com>
2026-09-19 12:11:23 -07:00

46 lines
2.1 KiB
Python

"""Stored /v1/responses transcripts must not duplicate history across chained turns.
The agent's ``result["messages"]`` copies of the prior history carry ``timestamp`` /
``_db_persisted`` / ``reasoning`` / ``finish_reason`` that the API layer's bare
``{"role", "content"}`` dicts never have, so a whole-dict prefix check failed every turn and
the stored history grew 3 -> 8 -> 17 instead of 2 -> 4 -> 6 (#95137, #101644, #82513).
"""
from agent.agent_runtime_helpers import repair_message_sequence
from agent.message_metadata import append_message
from gateway.platforms.api_server import APIServerAdapter
class _Agent:
api_mode = "chat_completions"
session_id = "s"
def _agent_turn(history, user_message, answer):
"""Real producers: turn_context appends the user row via append_message (stamps timestamp)."""
messages = list(history)
append_message(messages, {"role": "user", "content": user_message})
repair_message_sequence(_Agent(), messages)
append_message(messages, {"role": "assistant", "content": answer, "reasoning": "r", "finish_reason": "stop"})
return {"messages": messages}
def test_chained_responses_turns_store_each_message_once():
build = APIServerAdapter._build_response_conversation_history
history, counts = [], []
for turn, prompt in enumerate(["Remember TOKEN-ABC.", "What token?", "Third"]):
answer = f"answer {turn}"
history = build(history, prompt, _agent_turn(history, prompt, answer), answer)
counts.append(len(history))
assert counts == [2, 4, 6]
assert [m["role"] for m in history] == ["user", "assistant"] * 3
def test_suffix_only_result_still_appends_to_history():
build = APIServerAdapter._build_response_conversation_history
prior = [{"role": "user", "content": "a"}, {"role": "assistant", "content": "b"}]
# Mocked/legacy paths return only this turn's rows (no user row): appended after the input.
result = {"messages": [{"role": "assistant", "content": "d"}]}
stored = build(prior, "c", result, "d")
assert [m["content"] for m in stored] == ["a", "b", "c", "d"]