14 Commits

Author SHA1 Message Date
kshitijk4poor
d24aadfdd1 refactor(copilot): one GitHub effort clamp for the profile and main agent
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.

The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
2026-09-26 22:21:47 +05:30
Jash Lee
7a597324a1 fix(copilot): preserve Astra reasoning effort
(cherry picked from commit 108f20c553852246deb0ce86fd5474d5d7992dc1)
2026-09-26 22:21:47 +05:30
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
kshitijk4poor
8c6cb6f995 refactor(agent): drop the Optional import orphaned by the LM Studio helper removal 2026-09-14 20:35:28 +05:30
kshitijk4poor
488f2fc86d refactor(agent): drop the orphaned hand-rolled summary kwargs builder
`_chat_summary_attempt` now builds through `_build_api_kwargs`, leaving
`_iteration_summary_chat_kwargs` (56 lines mirroring the transport by hand)
and its only consumer `AIAgent._resolve_lmstudio_summary_reasoning_effort`
without a caller. The transport already owns every quirk they re-derived
(fixed temperature, LM Studio `reasoning_effort`, portal tags, provider
preferences, pareto router plugin), so there is nothing to keep in sync.
2026-09-14 20:35:28 +05:30
Teknium
7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
c13d0b023e refactor(agent/reasoning_params,rate_limit_credits): module-level helpers keep AIAgent attribute surface identical 2026-09-02 19:22:52 -07:00
Teknium
922717d538 refactor(agent/G_small): compact docstrings (keep WHY/invariants), inline constant comments 2026-09-02 19:16:38 -07:00
Teknium
270128294e refactor(agent/G_small): single-line signatures/calls, github effort fallback table, any() slot scan 2026-09-02 19:08:38 -07:00
Teknium
ca5f9f342a refactor(agent/G_small): hoist echo-family import, dataclass pending item, adopt_credits_state helper 2026-09-02 18:58:29 -07:00
Teknium
d038249b12 refactor(agent/G_small): second code pass — merged guards, table-driven compact display, joined wrapped calls 2026-09-02 18:51:24 -07:00
Teknium
d49ec401e3 refactor(agent/reasoning_params): unify LM Studio/Ollama probe caches, flatten gate ladder 2026-09-02 18:25:16 -07:00
Teknium
0ab2e9672c refactor(run_agent): extract VisionMessagePrepMixin and ReasoningParamsMixin 2026-09-02 13:29:40 -07:00