feat(plugin-catalog): add desensitize (community, tools)
Chinese-context PII redaction. Adds one catalog entry pinning yubingz/hermes-desensitize at 8ef2d22. The gap this fills: the catalog has secret redaction (vaultknox: API keys, tokens, passwords — a 28-pattern English key-format registry) and, as open proposals, English/Western PII middleware (#102922) and document-scoped anonymization (#102049). None of them recognize Chinese-language entities, which are a different problem rather than a translation of the English one: - Person names have no whitespace or capitalization boundary; the name is identifiable only by part-of-speech, not token shape. - Company and institution names are marker-SUFFIXED (有限公司 / 研究院 / 分行), so recognition keys off the tail with a variable-length run before it. - Addresses are hierarchical and recursive across 省/市/区/县/街道. - Magnitudes are written with the unit attached (3.5 亿元), so masking digits alone still leaks the scale. Capabilities declared (verified against the pinned commit by loading it through the real plugin loader against a temp HERMES_HOME — all four hooks register, and provides_tools / provides_middleware are genuinely empty): provides_hooks: pre_llm_call, pre_api_request, transform_llm_output, post_llm_call provides_tools: [] requires_env: [] Design note for review: it uses the existing hook pair rather than register_middleware("llm_request", ...) as in #102922. Hooks were sufficient and need no new surface; if middleware is the preferred long-term seam for message-rewriting plugins, that is worth settling before more plugins are written against hooks. Raised in #118722. The entry's description carries a Disclosure paragraph, following the vaultknox precedent, covering the parts a reviewer should not have to discover by reading the code: inbound messages and outbound replies are both rewritten, placeholders are code-generated and can over-match, company names are matched greedily (a modifier is absorbed into the placeholder), magnitude masking drops the numeric value and fires only on configured trigger words, and the placeholder mapping is memory-only. Structural check: `python3 scripts/validate_plugin_catalog.py plugin-catalog/desensitize.yaml` -> OK, 1 file(s) valid. No contributors/emails file is needed: yflmq001@users.noreply.github.com is already mapped. Fixes #118722
This commit is contained in:
34
plugin-catalog/desensitize.yaml
Normal file
34
plugin-catalog/desensitize.yaml
Normal file
@@ -0,0 +1,34 @@
|
||||
name: desensitize
|
||||
repo: https://github.com/yubingz/hermes-desensitize
|
||||
sha: 8ef2d22bafff5754edb1bbafd969b67ab4a1c06d
|
||||
subdir: src/hermes_desensitize
|
||||
description: "Chinese-context desensitization for Hermes chats. Replaces PII and business-sensitive
|
||||
entities — person names, company names, addresses, file paths, and magnitude figures — with
|
||||
reversible placeholders before a message reaches the model, then restores them in the reply.
|
||||
Two layers: an optional local/OpenAI-compatible LLM for semantic entity judgment, plus a
|
||||
sub-millisecond regex fallback (phone, ID card, email, IP, path) that keeps the conversation
|
||||
working when the model is unavailable. Disclosure — rewrites inbound user messages and outbound
|
||||
assistant replies, so transcript content differs from what you typed; placeholders are
|
||||
code-generated and may over-match (18-digit ID-card-shaped strings, names that collide with
|
||||
the public-entity allowlist); company names are matched greedily and a modifier such as
|
||||
原北京某某有限公司 is absorbed into the placeholder; magnitude masking drops the numeric
|
||||
value and only fires on the configured trigger words, so a domain term absent from that list
|
||||
is left unmasked without warning; the placeholder mapping is held in memory only and never
|
||||
written to disk."
|
||||
maintainer: yubingz
|
||||
tier: community
|
||||
category: tools
|
||||
requires_hermes: ">=0.19"
|
||||
docs_url: https://github.com/yubingz/hermes-desensitize#readme
|
||||
version: "1.0.0"
|
||||
readme: true
|
||||
platforms: []
|
||||
capabilities:
|
||||
provides_tools: []
|
||||
provides_hooks:
|
||||
- pre_llm_call
|
||||
- pre_api_request
|
||||
- transform_llm_output
|
||||
- post_llm_call
|
||||
provides_middleware: []
|
||||
requires_env: []
|
||||
Reference in New Issue
Block a user