The generic hardcoded_secret pattern fired on constants whose value is the
NAME of the credential environment variable (an ENV_PASSWORD-style constant
holding the string "MYPLUGIN_" + "APP_PASSWORD"): a line-oriented regex
cannot tell a reference to a secret from the secret itself, so one critical
finding made such plugins uninstallable with no --force override (#116221).
Both copies of the generic pattern (the shared threat-pattern library and
the skill/plugin guard table the installer actually walks) now skip a value
that is itself a SHOUTY_SNAKE environment-variable name (at least two
underscore-separated segments). The carve-out is scoped case-sensitive
because both tables compile with IGNORECASE: a lowercase snake value is the
passphrase shape, and requiring an underscore segment keeps
underscore-free all-caps credentials (AWS access key IDs, base32 secrets)
matched. Prefixed provider tokens (sk-, ghp_, ...) and the dedicated
provider-signature patterns are unaffected.
Regression tests pin both directions: the env-var-name line no longer
flags, and the passphrase / AKIA / base32 / prefixed-token shapes still do.
Review of the write-verb gate found sed -i, chmod, truncate, curl -o, wget -O,
git clone and a scripted open(...) against ~/.ssh all slipping to no finding,
where the bare path regex on main caught them. Add those verbs and the open(
shape to the gate; read-only mentions stay clear.
Salvage trim of #89249:
- keep pattern id `ssh_access` (test_memory_tool and callers key on it; renaming buys nothing)
- word-bound the verb alternation: the unanchored form matched `add` inside "address" and
`dd` inside "middle", so read-only prose still fired; `\b` closes that leak
- fold the second "bare leading redirect" regex into the same alternation (`>>?` branch)
- tests: one parametrized "write shapes still fire" (echo, cat, cp, tee, mv, install,
printf, dd, scp, rsync, ln, leading redirect, option clusters) and one "read-only
mention does not fire"; the trade-off/change-detector tests are dropped
- test_memory_tool: the persistence fixture used a read-only mention; use a write shape
Extend the write-verb alternation with mv/install/printf/dd/scp/rsync/ln,
allow short option clusters between verb and path, and add a bare-leading-
redirect branch (a > ~/.ssh/... line carries no verb word at all). Prose
that names a write primitive before an SSH path still fails closed — a
false positive costs a review, a false negative is a backdoor.
(cherry picked from commit 4865b96ca3c8661c0ea6f5c479c198ccf68f86c4)
The bare \$HOME/.ssh|~/.ssh regex fired on ANY mention of SSH paths in
scanned content, so operational documentation (VPS recovery notes, SSH
configuration write-ups stored in memory or skills) was blocked as a
persistence threat. Require an echo/cat/cp/tee/append/add/write/>>
verb before the path so only backdoor-insertion shapes match.
(cherry picked from commit 8c9b4c8972ac849e4c82fa20d355733ca3e5a231)
Salvage the bounded target-slot design from #104617, using mandatory
word separators to avoid ambiguous repeated matches. Preserve long
payload detection and execution-verb boundaries. Replace the three
candidate tests with two context-loader invariants and document the
heuristic's limits.
Fixes#104609
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Same bug class as the salvaged terminal-scanner fix: skills_guard's
env_exfil_curl/wget/fetch used unanchored \w*(KEY|TOKEN|...|API)
alternations, so any var with API/KEY/TOKEN mid-name
($TRILLIUM_ETAPI_URL) scored a critical exfiltration finding. Applied
the same \b anchor + plural tolerance, dropped mid-name API (every real
secret it caught already ends in KEY/TOKEN), kept CREDENTIAL, and kept
the loopback exemption from #98246. httpx/requests patterns unchanged —
their (KEY|TOKEN|...) alternation is unanchored-by-design against
argument text, not var-name suffixes.
Anchor env var name matches with \b to avoid matching legitimate
env vars that contain KEY/TOKEN/API as substrings (e.g.,
$TRILLIUM_ETAPI_URL). The patterns now require KEY/TOKEN/SECRET/PASSWORD
to appear at the END of the env var name, reducing false positives on
common API-usage documentation in SOUL.md while still catching actual
exfiltration attempts.
Fixes#63977
Salvaged from PR #35130 (the safe subset of jnibarger01's security pass):
- threat_patterns.py: replace unbounded (?:\w+\s+)* filler with bounded
{0,8} + cap scan input at MAX_SCAN_CHARS (64KiB), and bound the .*
runs in the exfil/config-mod patterns. Kills catastrophic backtracking
on adversarial near-misses.
- hermes_state.py: cap FTS5 query length (MAX_FTS5_QUERY_CHARS) and
extract quoted phrases with a linear scan instead of a regex so
pathological quote runs can't induce backtracking.
- acp_adapter/edit_approval.py + agent/tool_dispatch_helpers.py: recognize
'*** Move File: src -> dst' V4A headers so patch-mode edits are
permissioned/traversal-checked (previously only Update/Add/Delete), and
surface a proposal for mode=patch V4A calls (previously replace-only).
Tests: +ReDoS-bound + FTS5-cap + Move-File-target + V4A-approval cases.
Three independent security-scanner hardenings, re-homed onto the current
shared threat-pattern architecture (tools/threat_patterns.py):
- approval.py: add bash/sh/zsh/ksh heredoc to DANGEROUS_PATTERNS. The
existing heredoc pattern only covered python/perl/ruby/node, so
`bash <<'EOF' ... EOF` ran arbitrary shell — including exfil pipelines
whose inner commands don't individually match a pattern — with no prompt.
- threat_patterns.py: apply unicodedata.normalize("NFKC", ...) before
pattern matching so full-width / compatibility homographs (e.g.
`cat ~/.hermes/.env`) are folded to ASCII and no longer bypass the
keyword scanners. Invisible-char detection still runs on the raw content
first (NFKC can strip those codepoints).
- code_execution_tool.py: add CREDS/BEARER/APIKEY to _SECRET_SUBSTRINGS so
vars like HERMES_LLM_CREDS, API_BEARER, MY_APIKEY are scrubbed from the
sandbox env. PASS was intentionally dropped from the original proposal —
it false-positives on BYPASS_CACHE / COMPASS_DIR / PASSENGER_HOST while
PASSWORD/PASSWD already cover the credential cases.
The original PR also proposed a 'synonym' injection pattern block
(overlook/forget/set aside/bypass/discard + developer-mode); dropped here
because it false-positives on ordinary AGENTS.md/SOUL.md prose ("don't
forget to follow the rules", "run in developer mode"), exactly the
bossy-English class threat_patterns.py is documented to avoid.
Salvaged from #9028.
Co-authored-by: Hermes Agent <agent@nousresearch.com>
The known_c2_framework threat pattern included 'praxis' in its
alternation alongside genuine offensive-security tool brands (Cobalt
Strike, Sliver, Havoc, Mythic, Metasploit, Brainworm). Unlike those
distinctive brand names, 'praxis' is a common English word (Greek for
practice/action) and a legitimate agent name, so any context file that
mentioned an agent named Praxis matched at 'context' scope and the whole
AGENTS.md / SOUL.md was replaced with a [BLOCKED] placeholder before it
reached the system prompt.
Remove 'praxis' from the alternation and add a guard comment: every
token in this list must be a distinctive tool brand, not a common word.
Real C2 brands still fire.
Hardens the context window against Brainworm-class promptware attacks
(see #496). Three changes:
1. tools/threat_patterns.py — single source of truth for injection/promptware
patterns. Replaces the duplicated pattern lists in prompt_builder.py and
memory_tool.py. Adds ~15 new Brainworm/C2 patterns (node registration,
heartbeat/beacon, pull tasking, anti-forensic disk avoidance, identity
override, known framework names). Three scopes — 'all' (narrow, classic
injection), 'context' (adds promptware/role-play, broader detection),
'strict' (adds persistence/SSH-backdoor patterns for user-mediated writes).
2. MemoryStore.load_from_disk() now scans entries at snapshot-build time.
Poisoned entries are replaced with [BLOCKED: ...] placeholders in the
frozen system-prompt snapshot. Live state keeps the original so the
user can still inspect + remove via memory(action=read/remove). Scan is
deterministic from disk bytes — prefix-cache invariant holds.
3. make_tool_result_message() wraps results from high-risk tools
(web_extract, web_search, browser_*, mcp_*) in
<untrusted_tool_result source="...">...</untrusted_tool_result>
delimiters with framing prose telling the model the content is data,
not instructions. Architectural defense against indirect injection
from poisoned web pages, GitHub issues, MCP responses — does NOT
regex-scan tool results (pattern arms race + per-iteration latency).
Multimodal content lists pass through unwrapped to preserve adapter
compatibility.
Pattern philosophy: anchor on C2-specific vocabulary or unambiguous attack
behavior, NOT on bossy English. Dropped patterns suggested in #496 that
would have tripped legitimate content: standalone 'you are obligated to',
'do not respond immediately', 'you must X' without a C2-verb anchor.
Validation:
- 257/257 targeted tests pass (test_threat_patterns + test_memory_tool +
test_tool_dispatch_helpers + test_prompt_builder)
- E2E run with real Brainworm payload: blocked from AGENTS.md context-file
path, blocked from MEMORY.md snapshot, wrapped in delimiters when
arriving via web_extract. Legitimate 'you must follow conventions'
phrasing not flagged.
Explicitly NOT in this PR (per #496 discussion):
- Per-tool-result regex scanning (pattern arms race)
- SessionBehaviorMonitor / polling-loop detection (wrong layer)
- Outbound network gating (Docker backend already covers this)
- security.context_scanning warn|block knob (current behavior is always
block-with-placeholder — there's no warn mode that makes sense)
Closes#496 for Phase 1 + the architectural delimiter piece of Phase 2.
Phase 3 stays in tracking issue territory.