Files
hermes-agent/website/docs/user-guide
teknium1 004c8a51f0 feat(api-server): reasoning on non-streaming replies; echoed reasoning input items are ignored
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:

- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
  `choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
  completed `reasoning` output item ahead of the message (and ahead of that step's
  `function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
  from the assistant messages the agent already persisted (`build_assistant_message`
  stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
  in non-stream mode the agent fires the callback per provider delta AND once more with
  the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
  `conversation_history`. Responses SDK clients replay a prior response's `output`
  list as the next `input`; before, the item became an empty `user` message (in
  `input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
  (`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
  `output_item.done` item and the non-streaming shape.

Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
2026-09-19 00:46:31 -07:00
..