Files
hermes-agent/tools/plugin_guard.py
Teknium 9bcbe7b5df feat(i18n): pluggable, layered language packs across core, Desktop and TUI (#126296)
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales

* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter

ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.

* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs

* chore(tui): split en catalog siblings by lane (slash sibling)

* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports

* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts

* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list

* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes

* feat(plugins): report language-pack layers in the mid-run activation summary

* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n

StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.

* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter

- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
  the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
  layers partial packs (nested or flat dotted) over bundled/en via
  mergeTranslations; a string over a function-valued en entry becomes a
  positional {0}/{1} formatter; $appLocaleVersion bumps so translators
  re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
  registry as source 'backend' (method-not-found is silent); re-synced on
  socket open, display.language change and profile switch. A saved pack-only
  language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
  tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
  from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).

* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)

Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.

* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()

- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
  ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
  accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
  properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces

* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)

* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()

Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.

locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).

* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()

- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions

* i18n(platforms): route Google Chat and Teams user-facing text through t()

Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.

Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.

* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()

LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.

* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)

Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.

* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)

_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.

* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)

RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.

* i18n(cli): live-work dock, subagent monitor and render copy through t()

cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.

* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()

* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)

get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.

* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output

* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time

Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).

* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)

cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.

* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()

- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
  at call time; category labels via slash.category.*; help/alias/usage suffixes
  via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
  bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
  lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
  status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
  approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
  bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
  heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
  sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
  keeps its English constant; callers use history_unreadable() ->
  gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.

* i18n(telegram): route adapter chat copy through t()

Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.

Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.

* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()

- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor

* i18n(discord): route adapter chat copy through t()

Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().

* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)

* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()

- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
  existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
  via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
  _DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
  Updating/Generating) are one full template per variant; plurals use
  <key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
  translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).

* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English

* test(i18n): pin Telegram/Discord adapter catalog wiring

Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.

* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge

* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint

* i18n(tr): translate bundled catalog + tui pack

* i18n(ja): translate bundled catalog + tui pack

* i18n(ko): translate bundled catalog + tui pack

* i18n(zh): translate bundled catalog + tui pack

* i18n(fr): translate bundled catalog + tui pack

* i18n(af): translate bundled catalog + tui pack

* i18n(uk): translate bundled catalog + tui pack

* i18n(ar): translate bundled catalog + tui pack

* i18n(pt): translate bundled catalog + tui pack

* i18n(it): translate bundled catalog + tui pack

* i18n(es): translate bundled catalog + tui pack

* i18n(zh-hant): translate bundled catalog + tui pack

* i18n(ru): translate bundled catalog + tui pack

* i18n(hu): translate bundled catalog + tui pack

* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)

* i18n(de): translate bundled catalog + tui pack

* i18n(ga): translate bundled catalog + tui pack

* test(i18n): fixture matches _normalize_lang(lang, home) signature

* i18n(tui): scaffold userMessages/slashCmd en siblings

* i18n(tui): wire secure prompts + content tables

* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays

Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.

* i18n(tui): wire slash ops/wake replies

* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)

* i18n(tui): wire slash core/debug/setup replies

* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)

* i18n(tui): wire slash session/topup/subscription replies

* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)

* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys

* i18n(tui): wire userMessages copy through the userMessages namespace

* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json

* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)

Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).

* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)

* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)

* tui: i18n-export-en script (English templates for pack translators)

* docs(i18n): bundled TUI packs are the bottom layer of the tui surface

* i18n(ru): translate TUI pack

* i18n(ar): translate TUI pack

* i18n(es): translate TUI pack

* i18n(pt): translate TUI pack

* i18n(ko): translate TUI pack

* i18n(de): translate TUI pack

* i18n(ja): translate TUI pack

* i18n(fr): translate TUI pack

* i18n(tr): translate TUI pack

* i18n(it): translate TUI pack

* i18n(zh): translate TUI pack

* i18n(zh-hant): translate TUI pack

* i18n(hu): translate TUI pack

* i18n(uk): translate TUI pack

1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.

Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).

* i18n(ga): translate TUI pack

* i18n(af): translate TUI pack

* plugin_guard: locale catalogs in language packs step down the agent-config family

A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.

* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)

* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header

Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.

* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel

- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
  and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
  agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)

* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})

* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property

* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal

* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export

The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.

* test: unbreak two main-red timing tests the PR merge-ref inherits

- test_local_runtime racing fake publishes the modern state record (legacy pid-only
  records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
  a fixed 0.35s, which a loaded CI runner does not always meet

* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)

* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak

SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.

* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings

* chore(i18n): regenerate desktop key catalog after main sync

---------

Co-authored-by: Teknium <teknium@nousresearch.com>
2026-09-28 14:16:18 -07:00

354 lines
17 KiB
Python

#!/usr/bin/env python3
"""Plugin Guard — ``skills_guard`` engine applied to ``hermes plugins install``/``update``.
Plugins run in-process but are *expected* to read their own env keys, call provider APIs
and spawn subprocesses, so: full pattern set on docs/config files (where prompt-injection
lives); the "reads own secret"/"HTTP call with key" family exempt on *code* files;
plugin-sized structural limits; VCS/venv noise skipped. ``safe`` installs, ``caution``
needs confirmation, ``dangerous`` is blocked and ``--force`` does NOT override.
"""
from __future__ import annotations
import ast
from datetime import datetime, timezone
from pathlib import Path
from typing import Iterator, List, Optional, Tuple
from tools.plugin_guard_context import (
STEP_DOWN, catalog_cap, is_agent_facing, is_base64_media, is_ci_workflow, is_data_decode, is_doc_prose,
is_inert_fixture_line, is_locale_catalog, is_loopback_only, is_pip_install_in_prose_literal,
is_regex_alternation_token, is_self_uninstall_doc, is_test_tree, prose_cap)
from tools.skills_guard import (
Finding, ScanResult, SUSPICIOUS_BINARY_EXTENSIONS, _determine_verdict, format_scan_report,
scan_file)
PLUGIN_SCANNER_VERSION = "plugin-guard-v8"
# Never scanned: VCS internals, caches, vendored envs.
EXCLUDED_DIRS = {
".git", "__pycache__", "node_modules", ".venv", "venv",
".mypy_cache", ".pytest_cache", ".ruff_cache", ".tox"}
# Test trees ARE scanned (``plugins_loader`` sets ``submodule_search_locations`` to the
# plugin root, so ``from .tests import evil`` runs whatever lives there), but findings under
# them step down one severity (``plugin_guard_context.is_test_tree``): fixtures deliberately
# hold hostile strings to prove the plugin rejects them, and an un-overridable ``dangerous``
# made such plugins uninstallable and taught authors to obfuscate their own tests (#89610).
# Code files, where "reads an env secret" / "HTTP call with a key" is normal (requires_env).
CODE_FILE_EXTENSIONS = {".py", ".js", ".ts", ".sh", ".bash", ".rb", ".pl", ".php"}
# Line-comment marker per code extension. Whole-line comments explain intent; hardening
# notes like "# a symlink could point at /etc/passwd" are prose *about* a defense.
COMMENT_PREFIXES_BY_EXTENSION = {
".py": "#", ".sh": "#", ".bash": "#", ".rb": "#", ".pl": "#", ".r": "#", ".jl": "#",
".js": "//", ".ts": "//", ".php": "//"}
# One severity step down from the pattern's default.
_COMMENT_SEVERITY_CAP = {"critical": "high", "high": "medium"}
# History, not an agent-facing instruction surface: a hardening entry mentioning the threat
# it fixed ("A symlink could point at /etc/passwd, so ...") is documentation, not the attack.
CHANGELOG_FILENAMES = {"changelog.md"}
# Pattern ids exempt on code files (every legitimate provider plugin trips them); still
# applied in full to docs/config files.
CODE_EXEMPT_PATTERN_IDS = {
"python_environ_get_secret", "python_getenv_secret", "python_os_environ", "node_process_env",
"ruby_env_secret", "env_exfil_httpx", "env_exfil_requests", "env_exfil_fetch",
"env_exfil_curl", "env_exfil_wget",
# Agent-facing instruction patterns are meaningless inside code (prompt docstrings trip them).
"context_exfil", "send_to_url", "fake_policy",
# Plugins legitimately write config.yaml in post_setup and base64 credentials (Basic auth).
"agent_config_mod", "agent_config_contract", "encoded_exfil"}
# Severity remaps: a bundled binary is warn-tier (repos occasionally vendor one); a mere
# ``~/.hermes/.env`` mention is how READMEs say where keys go (READING it still trips
# ``read_secrets_file``, critical); ``curl | sh`` in READMEs is caution, not a hard block.
SEVERITY_REMAP = {
"binary_file": "high", "hermes_env_access": "medium", "curl_pipe_shell": "high"}
# In JS/TS, these text matches cannot distinguish a UI label or DNS lookup
# template from a write or exfiltration operation. Keep them visible and require
# confirmation; do not silently allow them. Shell commands and instructions keep
# their critical severity, as do separate credential-read/exfiltration findings.
JS_CAPABILITY_REMAP = {"dns_exfil": "high", "ssh_backdoor": "high"}
# Plugin scans gate a HOST install: what matters is what executes on the host. Two critical
# families describe the author's own dev workflow when they appear in documentation files, so
# they land at high (caution) there instead of hard-blocking an otherwise auditable plugin; the
# same content in runtime code keeps its critical severity. The generic one-step prose cap for
# command/path-shaped findings lives in ``plugin_guard_context`` (``DOC_PROSE_EXTENSIONS``).
DOC_PROSE_DEMOTIONS = {
# Prose modification bullets ("- Modify: `CLAUDE.md`") in plan/design docs describe the
# repo's own files; only executable intent (shell writes, code) stays critical.
"agent_config_mod": "high",
# Example/demo credentials quoted in docs (placeholder hex, test tokens). Real token-shaped
# literals (sk-, ghp_, AKIA, glpat-, private keys) keep their own critical patterns.
"hardcoded_secret": "high",
}
# A root-level ``if __name__ == "__main__":`` block is the module's own self-test harness:
# ``plugins_loader`` imports plugins and never runs them as scripts, so a sample credential
# quoted there is a fixture, not a shipped secret — the test-tree reasoning applied where a
# root-level runtime file has no ``tests/`` to hold it (#112139). Narrower than the
# test-tree cap because the block is still directly executable code: only the generic
# sample-token pattern is demoted; destructive/persistence/exfil findings and the
# provider-signature patterns (``sk-``, ``AKIA``, ``ghp_`` ...) keep full severity there.
MAIN_GUARD_DEMOTIONS = {"hardcoded_secret": "high"}
# Structural limits — plugins are real codebases, far larger than skills.
MAX_PLUGIN_FILE_COUNT = 400
MAX_PLUGIN_TOTAL_SIZE_KB = 10 * 1024 # 10MB of scannable tree
MAX_PLUGIN_SINGLE_FILE_KB = 1024 # 1MB single file
def _walk(plugin_dir: Path) -> Iterator[Tuple[Path, str]]:
"""Yield (path, "a/b/c" relative path) for every non-excluded entry under plugin_dir."""
for f in plugin_dir.rglob("*"):
try:
rel_parts = f.relative_to(plugin_dir).parts
except ValueError:
continue
if not any(part in EXCLUDED_DIRS for part in rel_parts):
yield f, "/".join(rel_parts)
def _finding(pattern_id: str, severity: str, category: str, file: str, match: str, description: str) -> Finding:
return Finding(pattern_id, severity, category, file, 0, match, description)
def _is_main_guard(node: ast.If) -> bool:
"""Return whether an ``if`` node is the conventional module self-test guard."""
test = node.test
if not isinstance(test, ast.Compare) or len(test.ops) != 1 or not isinstance(test.ops[0], ast.Eq):
return False
if len(test.comparators) != 1:
return False
left, right = test.left, test.comparators[0]
return (
isinstance(left, ast.Name) and left.id == "__name__"
and isinstance(right, ast.Constant) and right.value == "__main__"
) or (
isinstance(right, ast.Name) and right.id == "__name__"
and isinstance(left, ast.Constant) and left.value == "__main__"
)
def _main_guard_body_lines(file_path: Path) -> set[int]:
"""Return lines executed only by ``if __name__ == '__main__'`` blocks.
Invalid Python deliberately returns no lines so its findings retain the
conservative severity.
"""
try:
tree = ast.parse(file_path.read_text(encoding="utf-8-sig"))
except (OSError, SyntaxError, ValueError): # ValueError: UnicodeDecodeError, NUL bytes
return set()
lines: set[int] = set()
for node in ast.walk(tree):
if not isinstance(node, ast.If) or not _is_main_guard(node):
continue
for statement in node.body:
lines.update(range(statement.lineno, getattr(statement, "end_lineno", statement.lineno) + 1))
return lines
def _filter_findings(findings: List[Finding], rel_path: str, file_path: Path) -> List[Finding]:
"""Apply plugin-specific exemptions and severity remaps to raw findings."""
is_code = Path(rel_path).suffix.lower() in CODE_FILE_EXTENSIONS
main_guard_lines = _main_guard_body_lines(file_path) if file_path.suffix.lower() == ".py" else set()
is_js = Path(rel_path).suffix.lower() in {".js", ".ts"}
# A CI workflow definition runs on the forge's runner, not the host: same cap as a README.
doc_prose = is_doc_prose(rel_path) or is_ci_workflow(rel_path)
locale_catalog = is_locale_catalog(rel_path)
lines = _file_lines(file_path) if findings else []
out: List[Finding] = []
for f in findings:
if is_code and f.pattern_id in CODE_EXEMPT_PATTERN_IDS:
continue
f.severity = (
(JS_CAPABILITY_REMAP.get(f.pattern_id) if is_js else None)
or SEVERITY_REMAP.get(f.pattern_id) or f.severity
)
if doc_prose and f.pattern_id in DOC_PROSE_DEMOTIONS:
f.severity = DOC_PROSE_DEMOTIONS[f.pattern_id]
line = lines[f.line - 1] if 0 < f.line <= len(lines) else f.match
f.severity = _context_severity(f, rel_path, line, doc_prose or locale_catalog, is_code, locale_catalog)
if _is_defensive_documentation(f, rel_path):
f.severity = _comment_severity(f)
# Last and critical-only: a one-step cap that can never re-raise a finding an
# earlier remap already lowered.
if (
f.pattern_id in MAIN_GUARD_DEMOTIONS
and f.severity == "critical"
and f.line in main_guard_lines
):
f.severity = MAIN_GUARD_DEMOTIONS[f.pattern_id]
out.append(f)
return out
_SEVERITY_RANK = {"low": 0, "medium": 1, "high": 2, "critical": 3}
def _at_most(severity: str, cap: str) -> str:
"""Lower *severity* to *cap*; never raise it."""
return cap if _SEVERITY_RANK.get(severity, 0) > _SEVERITY_RANK[cap] else severity
def _comment_severity(f: Finding) -> str:
"""A whole-line comment / changelog entry cannot execute: one step down for every finding,
a second for command/path shapes (a comment is prose); agent-facing shapes keep one step."""
sev = _COMMENT_SEVERITY_CAP.get(f.severity, f.severity)
return sev if is_agent_facing(f) else STEP_DOWN.get(sev, sev)
def _file_lines(file_path: Path) -> List[str]:
"""Full source lines (``Finding.match`` is truncated to 120 chars); unreadable → []."""
try:
return file_path.read_text(encoding="utf-8-sig").split("\n")
except (OSError, UnicodeDecodeError):
return []
def _context_severity(f: Finding, rel_path: str, line: str, doc_prose: bool, is_code: bool,
locale_catalog: bool = False) -> str:
"""Severity after the inert-context demotions (``plugin_guard_context``). Each rule only
ever lowers, and every finding stays in the report; the order runs from the broadest
context (where the text lives) to the narrowest (what the token sits inside)."""
sev = f.severity
if doc_prose:
sev = (catalog_cap(f) if locale_catalog else prose_cap(f)) or sev
if is_self_uninstall_doc(f, line):
sev = _at_most(sev, "medium")
if is_test_tree(rel_path):
# A key-shaped literal or quoted-only hostile string in a fixture is the corpus the
# plugin's own tests reject (#89610): a note. Executable test code steps down once.
inert = f.category == "credential_exposure" or is_inert_fixture_line(f, line, is_code)
sev = _at_most(sev, "medium") if inert else STEP_DOWN.get(sev, sev)
if f.pattern_id == "encoded_exfil" and is_base64_media(line):
sev = "low"
if is_code and is_regex_alternation_token(f, line):
sev = STEP_DOWN.get(sev, sev)
if f.pattern_id == "base64_decode_pipe" and is_data_decode(line):
sev = STEP_DOWN.get(sev, sev)
if is_loopback_only(f, line):
sev = "low" # 127.0.0.0/8 is a local service, not egress
if is_code and is_pip_install_in_prose_literal(f, line):
sev = "low" # "no pip install is needed" in a user-facing message
return sev
def _is_defensive_documentation(finding: Finding, rel_path: str) -> bool:
"""A whole-line code comment or a changelog entry *describes* threats (the attack a
defense rejects, the hardening a release shipped) instead of executing them, so its
findings cap one severity step lower — visible and reviewable, never un-overridable
``dangerous`` from prose alone. Runtime code and agent-facing docs keep full severity.
"""
if Path(rel_path).name.lower() in CHANGELOG_FILENAMES:
return True
prefix = COMMENT_PREFIXES_BY_EXTENSION.get(Path(rel_path).suffix.lower())
if prefix is None or not finding.match:
return False
stripped = finding.match.lstrip()
if not stripped.startswith(prefix):
return False
if prefix == "#" and stripped.startswith(("#!", "#:")):
return False
return True
def _dangerous_findings_summary(findings: List[Finding]) -> str:
"""Describe the critical findings that made a plugin install dangerous."""
critical = [finding for finding in findings if finding.severity == "critical"]
pattern_ids = sorted({finding.pattern_id for finding in critical})
names = f" ({', '.join(pattern_ids)})" if pattern_ids else ""
return f"{len(critical)} critical of {len(findings)} findings{names}"
def _check_plugin_structure(plugin_dir: Path) -> List[Finding]:
"""Structural checks sized for plugin repositories."""
findings: List[Finding] = []
file_count = 0
total_size = 0
resolved_root = plugin_dir.resolve()
for f, rel in _walk(plugin_dir):
if f.is_symlink():
file_count += 1
try:
resolved = f.resolve()
except OSError:
findings.append(_finding("broken_symlink", "medium", "traversal", rel,
"broken symlink", "broken or circular symlink"))
continue
if not resolved.is_relative_to(resolved_root):
findings.append(_finding("symlink_escape", "critical", "traversal", rel,
f"symlink -> {resolved}", "symlink points outside the plugin directory"))
continue
if not f.is_file():
continue
file_count += 1
try:
size = f.stat().st_size
except OSError:
continue
total_size += size
if size > MAX_PLUGIN_SINGLE_FILE_KB * 1024:
findings.append(_finding("oversized_file", "medium", "structural", rel, f"{size // 1024}KB",
f"file is {size // 1024}KB (limit: {MAX_PLUGIN_SINGLE_FILE_KB}KB)"))
ext = f.suffix.lower()
if ext in SUSPICIOUS_BINARY_EXTENSIONS:
findings.append(_finding("binary_file", SEVERITY_REMAP["binary_file"], "structural", rel,
f"binary: {ext}", f"binary/executable file ({ext}) bundled in plugin (cannot be scanned)"))
if file_count > MAX_PLUGIN_FILE_COUNT:
findings.append(_finding("too_many_files", "medium", "structural", "(directory)", f"{file_count} files",
f"plugin has {file_count} files (limit: {MAX_PLUGIN_FILE_COUNT})"))
if total_size > MAX_PLUGIN_TOTAL_SIZE_KB * 1024:
findings.append(_finding("oversized_bundle", "medium", "structural", "(directory)", f"{total_size // 1024}KB",
f"plugin is {total_size // 1024}KB total (limit: {MAX_PLUGIN_TOTAL_SIZE_KB}KB)"))
return findings
def scan_plugin(plugin_dir: Path, source: str = "") -> ScanResult:
"""Scan a plugin directory (typically the temp clone); every external plugin is ``community`` trust."""
all_findings: List[Finding] = []
if plugin_dir.is_dir():
all_findings.extend(_check_plugin_structure(plugin_dir))
for f, rel in sorted(_walk(plugin_dir)):
if f.is_file() and not f.is_symlink():
all_findings.extend(_filter_findings(scan_file(f, rel_path=rel), rel, f))
verdict = _determine_verdict(all_findings)
if all_findings:
categories = sorted({f.category for f in all_findings})
summary = f"{plugin_dir.name}: {verdict} — {len(all_findings)} finding(s) in {', '.join(categories)}"
else:
summary = f"{plugin_dir.name}: clean scan, no threats detected"
result = ScanResult(
skill_name=plugin_dir.name, source=source or plugin_dir.name, trust_level="community",
verdict=verdict, findings=all_findings, scanned_at=datetime.now(timezone.utc).isoformat(),
summary=summary)
result.scan_provenance = {
"scanner_version": PLUGIN_SCANNER_VERSION, "verdict": verdict, "source": result.source}
return result
def should_allow_plugin_install(
result: ScanResult, force: bool = False) -> Tuple[Optional[bool], str]:
"""Map a verdict to ``(allowed, reason)``: True installs, None asks to confirm, False blocks."""
n = len(result.findings)
if result.verdict == "safe":
return True, "Allowed (clean scan)"
if result.verdict == "caution":
if force:
return True, f"Force-installed despite caution verdict ({n} findings)"
return None, f"Requires confirmation (caution verdict, {n} findings)"
return False, (
f"Blocked (dangerous verdict, {_dangerous_findings_summary(result.findings)}). "
f"--force does not override a dangerous verdict.")
__all__ = [
"scan_plugin", "should_allow_plugin_install", "format_scan_report", "PLUGIN_SCANNER_VERSION"]