Files
hermes-agent/agent/context_breakdown.py
Teknium 9bcbe7b5df feat(i18n): pluggable, layered language packs across core, Desktop and TUI (#126296)
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales

* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter

ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.

* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs

* chore(tui): split en catalog siblings by lane (slash sibling)

* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports

* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts

* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list

* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes

* feat(plugins): report language-pack layers in the mid-run activation summary

* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n

StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.

* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter

- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
  the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
  layers partial packs (nested or flat dotted) over bundled/en via
  mergeTranslations; a string over a function-valued en entry becomes a
  positional {0}/{1} formatter; $appLocaleVersion bumps so translators
  re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
  registry as source 'backend' (method-not-found is silent); re-synced on
  socket open, display.language change and profile switch. A saved pack-only
  language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
  tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
  from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).

* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)

Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.

* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()

- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
  ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
  accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
  properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces

* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)

* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()

Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.

locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).

* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()

- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions

* i18n(platforms): route Google Chat and Teams user-facing text through t()

Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.

Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.

* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()

LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.

* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)

Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.

* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)

_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.

* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)

RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.

* i18n(cli): live-work dock, subagent monitor and render copy through t()

cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.

* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()

* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)

get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.

* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output

* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time

Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).

* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)

cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.

* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()

- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
  at call time; category labels via slash.category.*; help/alias/usage suffixes
  via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
  bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
  lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
  status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
  approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
  bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
  heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
  sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
  keeps its English constant; callers use history_unreadable() ->
  gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.

* i18n(telegram): route adapter chat copy through t()

Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.

Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.

* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()

- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor

* i18n(discord): route adapter chat copy through t()

Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().

* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)

* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()

- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
  existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
  via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
  _DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
  Updating/Generating) are one full template per variant; plurals use
  <key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
  translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).

* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English

* test(i18n): pin Telegram/Discord adapter catalog wiring

Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.

* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge

* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint

* i18n(tr): translate bundled catalog + tui pack

* i18n(ja): translate bundled catalog + tui pack

* i18n(ko): translate bundled catalog + tui pack

* i18n(zh): translate bundled catalog + tui pack

* i18n(fr): translate bundled catalog + tui pack

* i18n(af): translate bundled catalog + tui pack

* i18n(uk): translate bundled catalog + tui pack

* i18n(ar): translate bundled catalog + tui pack

* i18n(pt): translate bundled catalog + tui pack

* i18n(it): translate bundled catalog + tui pack

* i18n(es): translate bundled catalog + tui pack

* i18n(zh-hant): translate bundled catalog + tui pack

* i18n(ru): translate bundled catalog + tui pack

* i18n(hu): translate bundled catalog + tui pack

* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)

* i18n(de): translate bundled catalog + tui pack

* i18n(ga): translate bundled catalog + tui pack

* test(i18n): fixture matches _normalize_lang(lang, home) signature

* i18n(tui): scaffold userMessages/slashCmd en siblings

* i18n(tui): wire secure prompts + content tables

* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays

Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.

* i18n(tui): wire slash ops/wake replies

* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)

* i18n(tui): wire slash core/debug/setup replies

* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)

* i18n(tui): wire slash session/topup/subscription replies

* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)

* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys

* i18n(tui): wire userMessages copy through the userMessages namespace

* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json

* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)

Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).

* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)

* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)

* tui: i18n-export-en script (English templates for pack translators)

* docs(i18n): bundled TUI packs are the bottom layer of the tui surface

* i18n(ru): translate TUI pack

* i18n(ar): translate TUI pack

* i18n(es): translate TUI pack

* i18n(pt): translate TUI pack

* i18n(ko): translate TUI pack

* i18n(de): translate TUI pack

* i18n(ja): translate TUI pack

* i18n(fr): translate TUI pack

* i18n(tr): translate TUI pack

* i18n(it): translate TUI pack

* i18n(zh): translate TUI pack

* i18n(zh-hant): translate TUI pack

* i18n(hu): translate TUI pack

* i18n(uk): translate TUI pack

1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.

Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).

* i18n(ga): translate TUI pack

* i18n(af): translate TUI pack

* plugin_guard: locale catalogs in language packs step down the agent-config family

A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.

* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)

* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header

Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.

* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel

- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
  and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
  agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)

* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})

* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property

* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal

* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export

The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.

* test: unbreak two main-red timing tests the PR merge-ref inherits

- test_local_runtime racing fake publishes the modern state record (legacy pid-only
  records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
  a fixed 0.35s, which a loaded CI runner does not always meet

* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)

* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak

SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.

* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings

* chore(i18n): regenerate desktop key catalog after main sync

---------

Co-authored-by: Teknium <teknium@nousresearch.com>
2026-09-28 14:16:18 -07:00

348 lines
16 KiB
Python

"""Live session context-window breakdown for UI surfaces.
Estimates system prompt tiers, tool schemas, and conversation history for the
category breakdown. Overall occupancy retains its provider-usage or estimate
provenance; category estimates are not exact tokenizer counts or gate authority.
"""
from __future__ import annotations
import json
import re
from typing import Any, Dict, List, Optional, Sequence, Tuple
from agent.i18n import t
_SKILLS_BLOCK_RE = re.compile(r"<available_skills>.*?</available_skills>", re.DOTALL)
_SUBAGENT_TOOL_NAMES = frozenset({"delegate_task"})
# A category at zero tokens is dropped from the payload, which reads as "not
# configured" - true for MCP, memory and skills, and false for the
# conversation. Every session has one, so hiding the row at zero makes an empty
# transcript indistinguishable from a breakdown that never measured it (#87903).
#
# The membership rule, stated so a later addition argues from the same
# principle rather than from "this one felt important": a category belongs
# here when zero is a MEASUREMENT of something every session has, not the
# ABSENCE of something optional. "conversation" qualifies because a session
# cannot not have a transcript, so zero means "nothing said yet" and is worth
# showing. "mcp", "memory", "skills" and "subagent_definitions" do not: zero
# there means the user configured none, which is what dropping the row already
# communicates, and a permanent 0-token row would be noise on most hosts.
# "system_prompt" and "tool_definitions" are always present too but are never
# zero in practice, so adding them would buy nothing.
_ALWAYS_REPORTED = frozenset({"conversation"})
# id -> (dashboard color, /context glyph); declaration order is display order. The label is
# ``gateway.context.category.<id>`` resolved when the payload is built (never at import).
_CATEGORIES = {
"system_prompt": ("var(--context-usage-system)", "■"),
"tool_definitions": ("var(--context-usage-tools)", "▣"),
"rules": ("var(--context-usage-rules)", "▩"),
"skills": ("var(--context-usage-skills)", "▤"),
"mcp": ("var(--context-usage-mcp)", "▥"),
"subagent_definitions": ("var(--context-usage-subagents)", "▦"),
"memory": ("var(--context-usage-memory)", "▧"),
"conversation": ("var(--context-usage-conversation)", "▨"),
}
def _category_label(category_id: str) -> str:
return t(f"gateway.context.category.{category_id}")
_FREE_GLYPH = "·"
_CONTEXT_SOURCES = frozenset({"local_estimate", "provider_usage", "provider_usage_plus_estimate"})
_GRID_COLUMNS = 20
_GRID_ROWS = 5 # 100 cells → 1 cell per percent of the context window
_DETAILS_TABLE_LIMIT = 15 # display cap only; the underlying data keeps everything
def _chars_to_tokens(text: str) -> int:
from agent.model_metadata import estimate_tokens_rough
return estimate_tokens_rough(text)
def _json_tokens(value: Any) -> int:
return _chars_to_tokens(json.dumps(value, ensure_ascii=False)) if value else 0
def _bytes_to_tokens(size: Optional[int]) -> Optional[int]:
from agent.model_metadata import CHARS_PER_TOKEN
return None if size is None else (int(size) + 3) // CHARS_PER_TOKEN
def _skills_block(stable: str) -> str:
"""The live ``<available_skills>`` block inside the stable tier, or ''."""
m = _SKILLS_BLOCK_RE.search(stable)
return m.group(0) if m else ""
def _split_tools(tools: Sequence[dict]) -> Tuple[List[dict], List[dict], List[dict]]:
builtin: List[dict] = []
mcp: List[dict] = []
subagent: List[dict] = []
for tool in tools:
fn = tool.get("function") if isinstance(tool, dict) else None
name = str((fn if isinstance(fn, dict) else tool).get("name") or "")
bucket = mcp if name.startswith("mcp_") else subagent if name in _SUBAGENT_TOOL_NAMES else builtin
bucket.append(tool)
return builtin, mcp, subagent
def _memory_blocks(agent: Any) -> Tuple[str, str]:
memory_block = user_block = ""
store = getattr(agent, "_memory_store", None)
try:
if store is not None and getattr(agent, "_memory_enabled", True):
memory_block = store.format_for_system_prompt("memory") or ""
if store is not None and getattr(agent, "_user_profile_enabled", True):
user_block = store.format_for_system_prompt("user") or ""
except Exception:
pass
return memory_block, user_block
def _strip_blocks(text: str, *blocks: str) -> str:
for block in blocks:
if block:
text = text.replace(block, "")
return text.strip()
def _join(*parts: str) -> str:
return "\n\n".join(part for part in parts if part).strip()
def _glyph(cat: Dict[str, Any]) -> str:
return _CATEGORIES.get(str(cat.get("id") or ""), (None, "▪"))[1]
def context_display_source(compressor: Any) -> str:
"""Distinguish the built-in preflight display seed from a provider reading.
Engines without the built-in real-usage ledger own their occupancy figure.
A seed never updates that ledger, even if its number later matches real usage.
"""
real = getattr(compressor, "last_real_prompt_tokens", None)
shown = getattr(compressor, "last_prompt_tokens", 0) or 0
return "local_estimate" if isinstance(real, (int, float)) and shown > 0 and shown != real else "provider_usage"
def context_usage_fields(compressor: Any) -> Dict[str, Any]:
"""Current occupancy only; lifetime throughput is never a context fallback."""
used = max(0, getattr(compressor, "last_prompt_tokens", 0) or 0)
maximum = getattr(compressor, "context_length", 0) or 0
if not used or not maximum:
return {}
used = min(used, maximum)
source = context_display_source(compressor)
return {"context_used": used, "context_max": maximum,
"context_percent": max(0, min(100, round(used / maximum * 100))),
"context_source": source, "context_estimated": source != "provider_usage"}
def compute_session_context_breakdown(agent: Any, messages: Optional[List[dict]] = None) -> Dict[str, Any]:
"""Return a Cursor-style context usage breakdown for one live agent."""
from agent.model_metadata import estimate_messages_tokens_rough
from agent.usage_anchor import anchored_context_tokens
from agent.system_prompt import build_system_prompt_parts
messages = messages or []
parts = build_system_prompt_parts(agent)
stable = parts.get("stable", "") or ""
skills_index = _skills_block(stable)
memory_block, user_block = _memory_blocks(agent)
system_prompt_text = _join(
_strip_blocks(stable, skills_index), _strip_blocks(parts.get("volatile", "") or "", memory_block, user_block)
)
builtin_tools, mcp_tools, subagent_tools = _split_tools(list(getattr(agent, "tools", None) or []))
tokens_by_id = {
"system_prompt": _chars_to_tokens(system_prompt_text),
"tool_definitions": _json_tokens(builtin_tools),
"rules": _chars_to_tokens(parts.get("context", "") or ""),
"skills": _chars_to_tokens(skills_index),
"mcp": _json_tokens(mcp_tools),
"subagent_definitions": _json_tokens(subagent_tools),
"memory": _chars_to_tokens(_join(memory_block, user_block)),
"conversation": estimate_messages_tokens_rough(messages),
}
estimated_total = sum(tokens_by_id.values())
comp = getattr(agent, "context_compressor", None)
context_max = int(getattr(comp, "context_length", 0) or 0) if comp else 0
# Usage-anchored figure (provider-exact tokens of a response + delta of what was
# appended since) beats last_prompt_tokens (lags) and the heuristic. Prefer the
# turn-base anchor: on reasoning models later same-turn responses inflate
# prompt_tokens with replayed thinking that evaporates at the turn boundary, so
# anchoring on the LAST response makes the meter sawtooth. Fall back to the
# last-response anchor, then measured, then estimated.
anchor = getattr(agent, "_turn_base_usage_anchor", None)
context_used = anchored_context_tokens(messages, anchor, charge_stale_thinking=False)
if context_used is None:
anchor = getattr(agent, "_usage_anchor", None)
context_used = anchored_context_tokens(messages, anchor)
if context_used is None:
measured_used = int(getattr(comp, "last_prompt_tokens", 0) or 0) if comp else 0
context_used = measured_used if measured_used > 0 else estimated_total
source = context_display_source(comp) if measured_used > 0 else "local_estimate"
else:
delta = messages[int(anchor["base_count"]):]
if delta and delta[0].get("role") == "assistant":
delta = delta[1:]
source = "provider_usage_plus_estimate" if delta else "provider_usage"
# A single prompt can never exceed the model window; any excess is estimate drift.
if context_max:
context_used = min(context_used, context_max)
return {
"categories": [
{"color": color, "id": category_id, "label": _category_label(category_id), "tokens": tokens_by_id[category_id]}
for category_id, (color, _glyph_) in _CATEGORIES.items()
if tokens_by_id[category_id] > 0 or category_id in _ALWAYS_REPORTED
],
"context_max": context_max,
"context_percent": max(0, min(100, round(context_used / context_max * 100))) if context_max else 0,
"context_used": context_used,
"context_source": source,
"context_estimated": source != "provider_usage",
"estimated_total": estimated_total,
"model": getattr(agent, "model", "") or "",
}
def compute_context_details(agent: Any) -> Dict[str, Any]:
"""Expanded per-skill / per-toolset cost listing for ``/context all``.
Reuses the ``hermes prompt-size`` attribution (index-line bytes from the
live skills block; schema bytes via the registry's tool→toolset map).
"""
from hermes_cli.prompt_size import _compute_skills_breakdown, _compute_toolsets_breakdown
from agent.system_prompt import build_system_prompt_parts
skills_block = _skills_block(build_system_prompt_parts(agent).get("stable", "") or "")
tools = list(getattr(agent, "tools", None) or [])
return {
"skills": [
{
"name": entry.get("name", ""),
"index_tokens": _bytes_to_tokens(entry.get("index_line_bytes")) or 0,
"skill_md_tokens": _bytes_to_tokens(entry.get("skill_md_bytes")),
}
for entry in (_compute_skills_breakdown(skills_block) if skills_block else [])
],
"toolsets": [
{
"toolset": group.get("toolset", ""),
"tool_count": int(group.get("tool_count", 0) or 0),
"schema_tokens": _bytes_to_tokens(group.get("json_bytes")) or 0,
}
for group in (_compute_toolsets_breakdown(tools) if tools else [])
],
}
# ── /context rendering (CLI + gateway) ──────────────────────────────────────
# Pure text renderers over the payload above. The gateway skips the glyph grid
# (monospace is not guaranteed on messaging platforms).
def render_context_grid(payload: Dict[str, Any]) -> List[str]:
"""Glyph grid: 100 cells, one per percent of the context window; categories
fill in declaration order, the remainder is free space."""
context_max = int(payload.get("context_max") or 0)
total_cells = _GRID_COLUMNS * _GRID_ROWS
cells: List[str] = []
if context_max > 0:
for cat in payload.get("categories") or []:
tokens = int(cat.get("tokens") or 0)
# never render a nonzero category as invisible
n = round(tokens / context_max * total_cells) or (1 if tokens > 0 else 0)
cells.extend([_glyph(cat)] * n)
cells = cells[:total_cells]
cells.extend([_FREE_GLYPH] * (total_cells - len(cells)))
return [" ".join(cells[row * _GRID_COLUMNS:(row + 1) * _GRID_COLUMNS]) for row in range(_GRID_ROWS)]
def render_context_category_lines(payload: Dict[str, Any]) -> List[str]:
"""Render the 'Estimated usage by category' table as plain-text lines."""
categories = payload.get("categories") or []
context_max = int(payload.get("context_max") or 0)
estimated_total = int(payload.get("estimated_total") or 0)
denom = context_max or estimated_total
lines = [t("gateway.context.category_header")]
if not categories:
return [*lines, t("gateway.context.no_data_yet")]
free_label = t("gateway.context.free_space")
width = max(len(free_label), *(len(str(cat.get("label") or "")) for cat in categories))
for cat in categories:
tokens, label = int(cat.get("tokens") or 0), str(cat.get("label") or cat.get("id") or "")
lines.append(t("gateway.context.category_row", glyph=_glyph(cat), label=f"{label:<{width}}",
tokens=f"{tokens:>9,}", pct=f"{tokens / denom * 100 if denom else 0.0:>5.1f}"))
if context_max > 0:
free = max(0, context_max - estimated_total)
lines.append(t("gateway.context.category_row", glyph=_FREE_GLYPH, label=f"{free_label:<{width}}",
tokens=f"{free:>9,}", pct=f"{free / context_max * 100:>5.1f}"))
return lines
def _toolset_row(group: Dict[str, Any]) -> str:
return t("gateway.context.toolset_row", toolset=f"{group['toolset']:<24}", count=f"{group['tool_count']:>3}",
tokens=f"{group['schema_tokens']:>8,}")
def _skill_row(entry: Dict[str, Any]) -> str:
name = str(entry.get("name") or "")
if len(name) > 28:
name = name[:27] + "…"
md = entry.get("skill_md_tokens")
md_str = f"~{md:>8,}" if md is not None else f"{t('gateway.context.not_available'):>8}"
return t("gateway.context.skill_row", name=f"{name:<28}", index_tokens=f"{entry['index_tokens']:>6,}", md_tokens=md_str)
def _table(lines: List[str], title: str, rows: List[Dict[str, Any]], fmt) -> None:
"""Append a titled, display-capped table (blank-separated from a preceding one)."""
if not rows:
return
if lines:
lines.append("")
lines.append(title)
lines.extend(fmt(row) for row in rows[:_DETAILS_TABLE_LIMIT])
if len(rows) > _DETAILS_TABLE_LIMIT:
lines.append(t("gateway.context.and_more", count=len(rows) - _DETAILS_TABLE_LIMIT))
def render_context_details_lines(details: Dict[str, Any]) -> List[str]:
"""Render the expanded ``/context all`` per-skill / per-toolset tables."""
lines: List[str] = []
_table(lines, t("gateway.context.toolsets_title"), details.get("toolsets") or [], _toolset_row)
_table(lines, t("gateway.context.skills_title"), details.get("skills") or [], _skill_row)
return lines
def render_context_breakdown_lines(
payload: Dict[str, Any],
*,
details: Optional[Dict[str, Any]] = None,
grid: bool = True,
) -> List[str]:
"""Full /context view. ``grid`` prepends the glyph grid (CLI; the gateway
keeps its own gauge); ``details`` appends the expanded listings."""
lines: List[str] = [*render_context_grid(payload), ""] if grid else []
lines.extend(render_context_category_lines(payload))
context_max = int(payload.get("context_max") or 0)
if context_max > 0:
used, pct = int(payload.get("context_used") or 0), int(payload.get("context_percent") or 0)
mark = "~" if payload.get("context_estimated") else ""
lines.extend(["", t("gateway.context.window_line", mark=mark, used=f"{used:,}", max=f"{context_max:,}", pct=pct)])
source = payload.get("context_source")
if source:
source_label = t(f"gateway.context.source.{source}") if source in _CONTEXT_SOURCES else source
lines.append(t("gateway.context.source_line", source=source_label))
if details is None:
lines.extend(["", t("gateway.context.hint_all")])
elif detail_lines := render_context_details_lines(details):
lines.extend(["", *detail_lines])
return lines