One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
176 lines
5.8 KiB
Python
176 lines
5.8 KiB
Python
"""Regression test for the `/model` picker confirmation display.
|
|
|
|
Bug (April 2026): after choosing a model from the interactive `/model` picker,
|
|
``HermesCLI._apply_model_switch_result()`` printed ``ModelInfo.context_window``
|
|
straight from models.dev, which always reports the vendor-wide value (e.g.
|
|
gpt-5.5 = 1,050,000 on ``openai``). That ignored provider-specific caps — in
|
|
particular, ChatGPT Codex OAuth enforces 272K on the same slug. The sibling
|
|
``_handle_model_switch()`` (typed ``/model <name>``) was already fixed to use
|
|
``resolve_display_context_length()``; the picker path was missed, causing
|
|
"sometimes 1M, sometimes 272K" for the same model across sibling UI paths.
|
|
|
|
Fix: both display paths now go through ``resolve_display_context_length()``.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
from unittest.mock import patch
|
|
|
|
from hermes_cli.model_switch import ModelSwitchResult
|
|
|
|
|
|
class _FakeModelInfo:
|
|
context_window = 1_050_000
|
|
max_output = 0
|
|
|
|
def has_cost_data(self):
|
|
return False
|
|
|
|
def format_capabilities(self):
|
|
return ""
|
|
|
|
|
|
class _StubCLI:
|
|
"""Minimum attrs ``_apply_model_switch_result`` reads on ``self``."""
|
|
def _stage_and_swap_model(self, result, old_model):
|
|
# Staging + in-place swap lives in a helper; run the real one on this stub.
|
|
import cli as _cli_mod
|
|
return _cli_mod.HermesCLI._stage_and_swap_model(self, result, old_model)
|
|
|
|
agent = None
|
|
model = ""
|
|
provider = ""
|
|
requested_provider = ""
|
|
api_key = ""
|
|
_explicit_api_key = ""
|
|
base_url = ""
|
|
_explicit_base_url = ""
|
|
api_mode = ""
|
|
_pending_model_switch_note = ""
|
|
|
|
|
|
def _run_display(monkeypatch, result):
|
|
import cli as cli_mod
|
|
|
|
captured: list[str] = []
|
|
monkeypatch.setattr(cli_mod, "_cprint", lambda s, *a, **k: captured.append(str(s)))
|
|
# Avoid writing to ~/.hermes/config.yaml during the test.
|
|
monkeypatch.setattr(cli_mod, "save_config_value", lambda *a, **k: None)
|
|
cli_mod.HermesCLI._apply_model_switch_result(_StubCLI(), result, False)
|
|
return captured
|
|
|
|
|
|
def test_picker_path_uses_provider_aware_context_on_codex(monkeypatch):
|
|
"""``_apply_model_switch_result`` must prefer the provider-aware resolver
|
|
(272K on Codex) over the raw models.dev value (1.05M for gpt-5.5).
|
|
"""
|
|
result = ModelSwitchResult(
|
|
success=True,
|
|
new_model="gpt-5.5",
|
|
target_provider="openai-codex",
|
|
provider_changed=True,
|
|
api_key="",
|
|
base_url="https://chatgpt.com/backend-api/codex",
|
|
api_mode="codex_responses",
|
|
warning_message="",
|
|
provider_label="ChatGPT Codex",
|
|
resolved_via_alias=False,
|
|
capabilities=None,
|
|
model_info=_FakeModelInfo(), # models.dev says 1.05M
|
|
is_global=False,
|
|
)
|
|
with patch(
|
|
"agent.model_metadata.get_model_context_length",
|
|
return_value=272_000,
|
|
):
|
|
lines = _run_display(monkeypatch, result)
|
|
|
|
ctx_line = next((l for l in lines if "Context:" in l), "")
|
|
assert "272,000" in ctx_line, (
|
|
f"picker-path display must show Codex's 272K cap, got: {ctx_line!r}"
|
|
)
|
|
assert "1,050,000" not in ctx_line, (
|
|
f"picker-path display leaked models.dev's 1.05M for Codex: {ctx_line!r}"
|
|
)
|
|
|
|
|
|
def test_picker_path_falls_back_to_model_info_when_resolver_empty(monkeypatch):
|
|
"""If ``get_model_context_length`` returns nothing (rare — truly unknown
|
|
endpoint), the display still surfaces ``ModelInfo.context_window`` so the
|
|
user sees *something* rather than a silent blank.
|
|
"""
|
|
result = ModelSwitchResult(
|
|
success=True,
|
|
new_model="some-model",
|
|
target_provider="some-provider",
|
|
provider_changed=True,
|
|
api_key="",
|
|
base_url="",
|
|
api_mode="chat_completions",
|
|
warning_message="",
|
|
provider_label="Some Provider",
|
|
resolved_via_alias=False,
|
|
capabilities=None,
|
|
model_info=_FakeModelInfo(), # context_window = 1_050_000
|
|
is_global=False,
|
|
)
|
|
with patch(
|
|
"agent.model_metadata.get_model_context_length",
|
|
return_value=None,
|
|
):
|
|
lines = _run_display(monkeypatch, result)
|
|
|
|
ctx_line = next((l for l in lines if "Context:" in l), "")
|
|
assert "1,050,000" in ctx_line, (
|
|
f"resolver-empty path should fall back to ModelInfo, got: {ctx_line!r}"
|
|
)
|
|
|
|
|
|
def test_global_switch_clears_context_pin_owned_by_previous_route(monkeypatch):
|
|
import cli as cli_mod
|
|
|
|
writes = []
|
|
monkeypatch.setattr(cli_mod, "_cprint", lambda *_a, **_k: None)
|
|
monkeypatch.setattr(
|
|
"utils.atomic_roundtrip_yaml_update",
|
|
lambda path, key, value: writes.append((key, value)),
|
|
)
|
|
cli = _StubCLI()
|
|
cli.model = "shared-model"
|
|
cli.provider = "custom"
|
|
# Runtime may already diverge from persisted config through a session override.
|
|
cli.base_url = "https://small.example/v1"
|
|
result = ModelSwitchResult(
|
|
success=True,
|
|
new_model="shared-model",
|
|
target_provider="custom",
|
|
provider_changed=False,
|
|
api_key="",
|
|
base_url="https://small.example/v1",
|
|
api_mode="chat_completions",
|
|
warning_message="",
|
|
provider_label="Custom",
|
|
resolved_via_alias=False,
|
|
capabilities=None,
|
|
model_info=_FakeModelInfo(),
|
|
is_global=True,
|
|
)
|
|
|
|
configured = {
|
|
"model": {
|
|
"default": "shared-model",
|
|
"provider": "custom",
|
|
"base_url": "https://large.example/v1",
|
|
"context_length": 1_048_576,
|
|
}
|
|
}
|
|
with (
|
|
patch(
|
|
"agent.model_metadata.get_model_context_length",
|
|
return_value=256_000,
|
|
),
|
|
patch("hermes_cli.config.read_user_config_raw", return_value=configured),
|
|
):
|
|
cli_mod.HermesCLI._apply_model_switch_result(cli, result, True)
|
|
|
|
assert ("model.context_length", None) in writes
|