Commit Graph

109 Commits

Author SHA1 Message Date
Teknium
4b8c01f691 fix(multiplex): key tool-side and agent-side memos by profile home
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).

Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
2026-09-12 01:35:05 -07:00
Teknium
adf23550f5 fix(tools): profile-scoped checkpoint/snapshot paths, tool caches, TZ and schema paths under multiplex
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:

- tools/process_registry.py, tools/environments/{modal,singularity}.py:
  `_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
  call time (same seam as `tools/skills_tool._skills_dir`, so the existing
  `monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
  the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
  path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
  per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
  resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
  the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
  the `_X_resolved` flags and the lifecycle reset keep their shape.
  tools/file_tools.py drops its private `file_read_max_chars` memo and reads
  the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
  env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
  startup) is ignored in favour of the routed profile's config.yaml. Both
  sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
  the static schema text is profile-neutral and `dynamic_schema_overrides=`
  rebuilds the `display_hermes_home()` / create-dir hint per
  `get_definitions()`, so a routed profile's model sees its own paths.

Refs #95685.

Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
2026-09-11 15:44:00 -07:00
Teknium
0dcadf6f41 revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three
model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces.

Later non-Wisdom work on shared files (guest onboarding i18n, dashboard
startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call
sites and config were stripped from those files.
2026-09-11 11:54:49 -07:00
shannonsands
a6ee31f55a feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation

* feat(wisdom): add private contribution loop

* feat(wisdom): add managed consumption workflows

* fix(wisdom): close cross-repository safety gaps

* fix(wisdom): align local package and lifecycle policy

* fix(wisdom): require explicit profile setup

* docs(wisdom): repin reconciled gateway head

* fix(wisdom): fence content downloads and approval receipts

* docs(wisdom): record generation-fenced downloads

* docs(wisdom): record unified delivery PR

* fix(ci): stop passing invalid classifier inputs

* docs(wisdom): remove internal requirements ledger

* feat(wisdom): localize dashboard and desktop copy

* feat(wisdom): complete local contribution and consumption UX

* style(wisdom): satisfy desktop lint

* chore(wisdom): refresh requirements pin

* test(dashboard): allow formatted profile copy

* test(wisdom): stabilize desktop interaction coverage

* fix(wisdom): surface dashboard action failures

* fix(wisdom): add repeatable Portal demo login

* feat(wisdom): add actionable skill notifications

* feat(wisdom): add notification install and update actions

* fix(wisdom): make Telegram skill alerts actionable

* fix(wisdom): always refresh demo Agent login

* feat(wisdom): embed Telegram notification actions

* fix(wisdom): preserve Telegram notifications after actions

* fix(wisdom): keep Telegram notification cards readable

* feat(wisdom): add Telegram candidate approval flow

* feat(wisdom): explain Telegram qualification reasons

* fix(wisdom): reconcile cross-surface candidate actions

* feat(telegram): add Collective Wisdom management command

* chore(wisdom): refresh Gateway contract pin

* chore(wisdom): advance Gateway contract pin

* feat(wisdom): align command UX across clients

* feat(slack): add Collective Wisdom management parity

* feat(wisdom): add security and professionalism reviews

* feat(wisdom): add first-time qualification guidance

* feat(wisdom): simplify qualification sharing choices

* feat(skills): add optional editorial metadata

* feat(wisdom): enrich legacy skill presentation

* fix(wisdom): harden review and update boundaries

* fix(wisdom): emit canonical review timestamps

* fix(wisdom): align with merged gateway and main

* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)

- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
  7-day evidence builder that excludes bundled/hub/managed skills and
  dismissed/handled/recently-suggested content hashes, strict pydantic
  schemas for agent output with repair-or-reject, fixed copy templates
  (Share / Teammate / Published / Update / Mute), idempotent retried
  delivery ledger with stale-action resolution, weekly review job,
  resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.

* wisdom: agent-led renderers and button action dispatcher

- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
  is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
  ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
  packaging flow, Install/Update -> plan command. Never publishes/installs.

* wisdom: CLI verbs, agent_led config default, conversational catalog skill

- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
  verbs, share/install flows and fixed notification templates.

* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons

- gateway housekeeping tick calls maybe_run_weekly_review with a home
  channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
  duration keyboard, send_wisdom_agent_recommendation rich card + fallback.

* fix(wisdom): integrate local mediation and harden model and setup boundaries

* fix(wisdom): honor authoritative recommendation policy and defer on failure

* fix(wisdom): synchronize opaque suppression and recheck delivery preferences

* feat(wisdom): route weekly selection through the session-owned assessment queue

* fix(wisdom): prepare and submit the reviewed generated share package

* feat(wisdom): separate native Share preparation from publication consent

* feat(wisdom): sync native mute choices through a leased preference outbox

* feat(wisdom): bind native mute controls to durable preference choices

* feat(wisdom): add scoped desktop and dashboard notification settings

* fix(wisdom): revalidate feed recommendations before assessment and delivery

* fix(wisdom): persist validated delivery receipts before completing notices

* feat(wisdom): add private notification claim and receipt client

* Persist Wisdom send reservations and recover delivery acknowledgements

* Route legacy Wisdom controls through current native review

* Add typed private Wisdom operation outcome client

* fix(wisdom): make agent-led advice usable in the local demo

* fix(wisdom): keep requested consent outside proactive limits

* fix(wisdom): distinguish unavailable assessments and preserve digest text

* fix(wisdom): assess ongoing usefulness beyond the current task

* fix(wisdom): restore immediate qualification sharing controls

* fix(wisdom): separate qualification review from installation advice

* fix(wisdom): collapse review checklists and simplify sharing copy

* fix(wisdom): show compact sharing progress and publication receipts

* fix(wisdom): require credential prefixes rather than matching skill names

* fix(wisdom): finish package checks before presenting sharing consent

* fix(wisdom): scan local skills before qualification cards

* fix(wisdom): update moderation results on existing sharing cards

* fix(wisdom): keep sharing review accessible from receipt cards

* fix(wisdom): align mediated review cards and collapsible checks

* fix(wisdom): clarify clean security summary wording

* fix(wisdom): normalize consent plans and add explicit recheck

* fix(wisdom): keep install and update receipts concise

* fix(wisdom): collapse assessments and deduplicate operation cards

* fix(wisdom): restore private Portal review from native cards

* fix(wisdom): sync Portal publication to original consent card

* fix(wisdom): show local skill version on sharing cards

* fix(wisdom): skip agent recommendations for self-published versions

* fix(wisdom): simplify candidate notices and local-edit recovery copy

* feat(wisdom): submit locally reviewed packages with one confirmation

* feat(wisdom): expose safe receipt and outcome sync recovery

* wisdom: onboarding notice says detect and share, names the user's own skill

Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark

Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.

* wisdom: one opener, no approval line, ask to share after the skill is shown

Product owner review of the candidate card.

- The Hermes written card now opens with the same sentence as the fixed card
  ("Your organisation has enabled Collective Wisdom, a feature designed to
  automatically detect and share useful skills across all team members.")
  instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
  and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
  It is now the last line, after the skill name, description, why suggested
  and the checks, and reads "Would you like to share it?" (matching the
  agent led template wording).

Tests updated for the new order; proposalNotice removed from all desktop locales.

* wisdom: American spelling, organization

Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.

* wisdom: candidate card copy round 4 (owner review)

Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:

1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
   the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
   card (Telegram rich card and plain fallback, legacy agent-led share
   template).
3. The skill name and description are labelled: "Skill name: <name>" and
   "What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
   inappropriate content found)" with no per-check bullets and no "Pass";
   a failed review reads "Needs a look before sharing at work (possible
   inappropriate content)" and lists only the checks that flagged
   something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
   details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
   days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
   editorial_name, a simple one_line_description and a compelling
   why_coworkers_benefit under 300 characters; "Be concise and
   convincing." becomes "Be concise and compelling: the goal is that the
   user wants to share it."

Tests updated for the new strings; review_text() gains direct coverage.

* wisdom: re-apply owner copy after rebase

- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice

* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors

Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.

* fix(wisdom): reconcile optional SDK tests and frontend lint

* fix(wisdom): default to agent-written notification summaries

* fix(wisdom): restore deferred install review and browse controls

* feat(wisdom): inspect installed setup with exact package provenance

* feat(wisdom): run native-approved installed setup steps with durable evidence

* fix(wisdom): recover interrupted setup with explicit native consent

* feat(wisdom): hand native installs into guided setup review

* fix(wisdom): continue requested setup with fixed notification copy

* fix(wisdom): preserve setup while waiting for a session model

* fix(wisdom): expose canonical setup review controls on desktop

* fix(wisdom): resume setup after recorded automatic updates

* fix(wisdom): make missing setup prerequisites recheckable

* chore(wisdom): align Agent with verified Gateway contract

* fix(wisdom): stop guessing team slugs in portal links

* fix(wisdom): retire pending advice on account sign-out

* fix(wisdom): cancel advice after terminal account revocation

* fix(wisdom): fence feed responses across account sign-out

* fix(wisdom): checkpoint signed-out feed before reactivation

* fix(wisdom): link proactive advice to scoped notification settings

* fix(wisdom): coalesce queued publication recommendations by version

* fix(wisdom): keep package review navigation local and deferable

* fix(wisdom): reflect installed state in discovery controls

* fix(wisdom): show exact checks before command confirmation

* chore(wisdom): pin bounded analytics privacy contract

* chore(wisdom): pin retired legacy notification contract

* feat(wisdom): review publisher usage with exact sharing copy

* fix(wisdom): align discovery and review check summaries

* fix(wisdom): show expired consent before confirmation

* fix(wisdom): require fresh review for legacy install controls

* fix(wisdom): preserve review expiry across check toggles

* fix(wisdom): retain update policy in native install reviews

* fix(wisdom): surface failed native card edits

* fix(wisdom): persist local command approval reviews

* fix(wisdom): use saved approvals for messaging commands

* test(wisdom): provide scan result in setup handoff fixture

* test(wisdom): exercise Telegram approvals with saved review state

* fix(wisdom): retain suppression policy for offline deferral

* fix(wisdom): reconsider candidates after deferred suppression expires

* fix(wisdom): bind review checks and report verified readiness separately

* fix(wisdom): persist accepted publication intent and recover exact outcomes

* fix(sync): pin UTF-8 tree ordering across writers

* chore(wisdom): pin organisation-scoped Gateway authorization

* fix(wisdom): restrict consent delivery to user-facing sessions

* chore(wisdom): refresh reviewed Gateway contract pin

* fix(wisdom): preserve kept tools in Blank Slate exclusions

* test(auth): reset anonymous fixture with a profile-scoped cache

* fix(wisdom): gate local surfaces and work on current profile entitlement

* fix(wisdom): invalidate quiet tool cache on entitlement changes

* test(wisdom): authorize local consent gateway fixtures

* fix(wisdom): keep entitlement decoding free of native crypto imports

* test(wisdom): provide local entitlement to demo CLI subprocess

* ci: leave upstream workflow unchanged in Wisdom PR

* fix(wisdom): ship package and contracts in Nix wheels

---------

Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
2026-09-11 19:04:06 +10:00
Teknium
c240e65399 fix(skills): self-improvement writes lessons, not incident logs
The skill review fork, the combined memory+skill review, and the curator's
consolidation pass all described references/ as the place for
"session-specific detail", and the curator's demote step said to move a
sibling's file under the umbrella. Followed literally over months that
produced one dev skill with a 100k SKILL.md dense in PR numbers and 443
one-per-session reference files, plus five sibling skills restating the
same rules and the repo's AGENTS.md.

The three prompts now share one shape contract: an entry is an imperative
rule plus one clause of why, stated once; no PR/issue numbers, dates, or
quoted chat as content; references/ is a small topical set extended in
place, never a per-session file; skills do not restate always-loaded
context. Consolidation is defined as distilling, and copying a sibling
verbatim under references/ is named as the failure. skill_manage's schema
carries the one-sentence version.

Two advisory linter rules make the shape visible in the tool result the
moment it starts to drift: incident-log-shape (PR/issue-number density in
prose) and references-sprawl (>60 reference files), the latter also run
on references/ writes. On the real before/after: the old skill trips both,
the consolidated one trips neither.
2026-09-04 08:15:22 -07:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
eb8a30cc2f simplify(compat): skills_sync/skill_manager/kanban/computer_use — drop 37 re-exports/aliases, repoint 8 callers + 9 test files 2026-09-03 13:27:01 -07:00
Teknium
b1047e80d1 review-fix(comments): restore lost #NNNN rationale comments (sweep, remaining files) 2026-09-03 09:44:36 -07:00
Teknium
812943b3ed review-fix(suppress-audit): whatsapp adapter, skill_manager_tool, skill_usage — restore BASE exception semantics 2026-09-03 09:42:58 -07:00
Teknium
e6048ff6aa refactor(tools): skill_manage/ledger — shared identifier check, chainable result decorators, ledger read/filter fold 2026-09-03 01:30:26 -07:00
Teknium
370c6a2456 refactor(tools): skill_manage — dispatch default for unknown action, simpler ledger read, squeeze body blanks 2026-09-03 01:26:03 -07:00
Teknium
4ba77a6c13 refactor(tools): skill_manage — shared missing-arg predicates, compact category validation and comments 2026-09-03 01:23:09 -07:00
Teknium
eeb86e4b6f refactor(tools): skill_manage — fold gate/flat-op plumbing, shared root/pin/org helpers in guards, ledger tarball lookup inline; compact docstrings 2026-09-03 01:17:57 -07:00
Teknium
d9d25c89ba refactor(tools): unify skill_manage guard/ledger boilerplate; compact batch + ledger helpers 2026-09-03 01:06:07 -07:00
Teknium
fbd054e0ff refactor(tools): skill_manager_tool — walrus gates, compact profile lookup and file-path validation 2026-09-02 23:13:20 -07:00
Teknium
ce992963a2 refactor(tools): skill ledger/batch — inline one-use helpers, drop section banners 2026-09-02 23:10:14 -07:00
Teknium
a861976519 refactor(tools): skill guards/tool — is_relative_to collapses, string reflow 2026-09-02 23:03:24 -07:00
Teknium
b187d4e547 refactor(tools): skill linter checks as generators, skill_manage arg-shape table, suppress() collapses 2026-09-02 22:43:10 -07:00
Teknium
9832b0800d refactor(tools): skill_manager_tool — unify guarded writes, compact docstrings 2026-09-02 22:34:26 -07:00
Teknium
71dadfeb99 refactor(tools): trim skill_manage stack — dead linter CLI/format helpers, merged path resolvers, compact layout 2026-09-02 22:23:24 -07:00
Teknium
7c1ec19d4a refactor(tools/skills): split skills_hub into per-source modules; extract skill_manager guards/batch, skills_tool dedup/plugin/setup; compact skill_usage, skills_guard 2026-09-02 14:45:15 -07:00
giwaov
42c2838674 feat(skills): skills.create_dir routes agent-created skills to a configured directory
Salvaged from PR #13996 (@giwaov, issue #13963), modernized onto current
main: config key renamed to skills.create_dir (per PR #81002's naming),
resolution centralized in agent/skill_utils.get_skill_create_dir() with
~/${VAR} expansion and HERMES_HOME-relative paths, and the directory is
folded into get_all_skills_dirs() so created skills are discovered,
trusted, findable, and patchable like local ones. Out-of-root creations
report their absolute path instead of crashing relative_to().
2026-09-01 07:32:45 -07:00
zengzheqing
98e2f110cf fix(curator): restore complete skill packages on ledger rollback (#96962)
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.

The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.

Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
2026-08-31 10:08:13 -07:00
kshitijk4poor
5cc1369fa2 refactor(skills): clarify _find_skill docstring, narrow _local_root except
Simplify-code pass findings:
- The docstring claimed 'Matching bare-name suffix' but the fast path
  matches the exact directory name (parent.name) — reworded to say
  what the code actually does.
- _local_root() swallowed every Exception silently; narrowed to
  OSError (what resolve() raises) with a logger.debug breadcrumb so a
  recurring resolve failure is diagnosable instead of degrading every
  categorized lookup to a silent not-found.
2026-08-30 20:04:40 +05:30
kshitijk4poor
70220e529a refactor(skills): short-circuit bare-name match before resolve machinery
Review feedback (kokhlo): the categorized-name match ran
resolve().relative_to() for every SKILL.md in the walk even when the
bare-name branch already matched — 50+ resolve calls per invocation on
a bare-name lookup in a large profile.

Restructure so the bare directory-name check stays first and the
resolve/relative_to machinery only runs when the lookup name actually
contains a path separator. The skills root is resolved once, lazily,
only when a categorized lookup happens at all. Also compare the
relative path via as_posix() so 'category/skill' lookups work on
Windows, where str(Path) renders backslashes.
2026-08-30 20:04:40 +05:30
kshitijk4poor
70370e089c fix(skills): skill_view directory file_path + skill_manage categorized name resolution
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):

1. skill_view(name, file_path='references') returned a raw
   '[Errno 21] Is a directory' OS error. The local-skill branch gated
   on target_file.exists(); a directory passes exists(), fell through
   to read_text(), and raised. The plugin-skill sibling branch already
   gated on is_file() — this aligns the local branch so a directory
   request gets the same helpful not-found payload with
   available_files listing instead of an OS error.

2. skill_manage rejected categorized names ('category/skill-name')
   with 'not found in active profile'. _find_skill matched only the
   bare directory name, while skill_view's own ambiguity hint tells
   the caller to use exactly the categorized form — every call that
   followed the hint failed. _find_skill now also matches the full
   relative path of the skill dir, giving skill_manage resolution
   parity with skill_view across edit/patch/delete/write_file/
   remove_file.

Both fixes are covered by regression tests that fail on main.
2026-08-30 20:04:40 +05:30
kshitijk4poor
154fd10af0 fix: name the preserved snapshot path in the ROLLBACK FAILED payload
Folded from #97748 (the competing fix by @lEWFkRAD): when a rollback
restore fails, the error note now points the operator at the surviving
snapshot directory instead of leaving them to find it in tempdir.
2026-08-29 19:38:31 +05:30
Adolanium
1315e65a52 fix(skills): a failed rollback restore keeps the skill and the snapshots
Rollback removed the live skill directory before restoring its
snapshot. When copytree then failed (disk full, locked file, path too
long on Windows) the except only added a note, and the finally deleted
the snapshot directory too, so nothing survived: the skill was gone
with a success-shaped error payload.

The broken state is now renamed aside first and deleted only after the
snapshot is restored. If the restore still fails, the broken state is
renamed back, so the worst outcome is the half applied batch instead of
no skill at all. When rollback reports any failure the snapshots are
kept on disk and their location is logged, instead of being deleted by
the finally.

Follow-up to #97692, same batch executor.
2026-08-29 19:38:31 +05:30
Clifford Garwood
10e93c6ab9 fix(skills): drop redundant identical-strings guard and its vacuous tests
The guard and the tests around it pinned behavior that already existed.
fuzzy_find_and_replace rejects old_string == new_string at
tools/fuzzy_match.py:69-70, returning "old_string and new_string are
identical" — so main already answered success=False, and the three tests
asserting "identical" in the error passed with the guard deleted.

The earlier claim that this case "silently applies a no-op the model reports
as success" was wrong. Verified against main:

  {"success": false, "error": "old_string and new_string are identical",
   "file_preview": "..."}

The duplicate guard was also strictly worse: it fired before the skill
lookup and returned no file_preview, shadowing the richer message.

Removes the guard, the two identical-strings tests, and the third
parametrize case. What remains is the genuinely new behavior: an actionable
missing-old_string error, reachable through the public tool.
2026-08-28 23:17:35 -07:00
Clifford Garwood
4f4e778db8 fix(skills): make skill_manage patch failures recoverable instead of a dead end
`skill_manage(action='patch')` rejected a missing `old_string` with:

    old_string is required for 'patch'. Provide the text to find.

That is a dead end. The model cannot tell whether it omitted the argument or
supplied text that did not match, so it retries blindly — and then escapes to
the neighbour that always works: `action='write_file'`, which rewrites the
entire skill file and destroys unrelated content. `skill_manage`'s own action
enum puts that destructive path one token away from the failing one.

The error now names the recovery route: `old_string` must be the EXACT text
currently in the file, read the target first (the skill's SKILL.md, or the
file named by `file_path`), copy the snippet verbatim, and do not fall back
to `action='write_file'`.

Validation lives in `_patch_skill` rather than the dispatcher. `skill_manage()`
previously returned its own bare missing-argument error before ever calling
the helper, which would leave the new guidance unreachable through the public
tool. Removing that duplicate makes the helper the single source of truth; its
{"success": False, "error": ...} flows through the same json.dumps path, so
the serialized shape is unchanged, and validation still precedes the skill
lookup — a missing old_string on an unknown skill still reports the argument
error rather than "skill not found".

Fixes #33064
2026-08-28 23:17:35 -07:00
Teknium
62e8126c69 fix(skill_manage): batch failure results carry the failing op's teaching payload (file_preview, hints)
Live A/B eval (old flat vs operations[] on qwen3.8-27b / gpt-5.6-terra /
claude-sonnet-5) caught a real regression: on a fuzzy-match miss the flat
path returns file_preview so the model can self-correct, but the batch
wrapper rebuilt the error dict and dropped every field except error/
failed_index. Sonnet, recovering blind, probed the file by writing and
reverting placeholder patches for 8+ turns (50k tokens vs 15k on the flat
arm). Batch failures now merge through all non-error fields from the
failing op's result.
2026-08-28 19:38:21 -07:00
Teknium
72874b0675 feat(skill_manage): operations[] is the call — each op names its skill; atomic with cross-skill rollback (#97295)
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate

* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing

* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal

* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback

* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization

Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
2026-08-28 12:15:18 -07:00
Teknium
eff97a8a05 refactor(profiles): retire the cross-profile write guard — profiles are not isolated (maintainer decision); mirror lost-write guards (#32049) survive; patch/write_file schemas drop cross_profile (-83 tok/call) (#97165) 2026-08-28 06:36:22 -07:00
Teknium
9978706e93 refactor(skill_manage): schema diet — patch args defer to the patch tool's semantics, file_path states its skill-dir-relative shape, authoring curriculum compressed (517 -> 427, -17%) (#97152) 2026-08-28 05:48:15 -07:00
Teknium
3fd70ad1c9 refactor(skills): skill_manage 924 → 518 tok/call — dedup diet, retire 'edit', unadvertise curator-only absorbed_into (#95697)
* refactor(skills): diet skill_manage schema (924->567 tok/call) — drop prose duplicated in system prompt + validation errors

* refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)

* Revert "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"

This reverts commit 9988b33f177fa82a145d5243d556f113a7e2fd2b.

* Reapply "refactor(skills): retire 'edit' — patch takes content for a full rewrite (legacy alias kept)"

This reverts commit 49a0a8f18b478824398767d7ae681913cee00640.

* refactor(skills): unadvertise absorbed_into — curator-only vocabulary leaves the shared schema
2026-08-26 15:13:30 -07:00
Jony
335c60ecdd fix(skills): preserve review marks across contexts 2026-08-24 23:55:39 -07:00
Teknium
3733e4aff5 fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
2026-08-24 20:25:10 -07:00
Vadelma
968853c5b5 fix(skills): preserve document modes during atomic writes
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 12:52:53 -07:00
Teknium
efe41abde0 feat(curator): per-mutation audit ledger + single-edit rollback
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.

- tools/skill_ledger.py: append/list/get, blob store, actor derivation
  (curator|agent|user), single-entry rollback that takes a pre-rollback
  safety entry first and FAILS CLOSED when that capture fails (consistent
  with the whole-run tarball rollback hardening from #63366). Path
  containment check so a hand-edited ledger can't write outside
  HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
  delete intent recorded via absorbed_into/archived evidence),
  archive_skill()/restore_skill(), and curator auto-transitions (tagged
  actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
  Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
  hermes curator rollback <entry-id> (whole-tree snapshot rollback
  unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
  (default 0 = never) + explicit hermes curator purge, recorded in the
  ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
  archive TTL purge.

Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).

Closes #45778, #50875. Tests adapted from #50261 by @yu-xin-c.
2026-08-16 22:06:41 -07:00
kshitij
6061377bbf fix(tools): mirror must-differ guidance in skill_manage new_string schema
skill_manage's patch action uses the same fuzzy_find_and_replace engine
as the file patch tool and surfaces the identical-strings error verbatim
— and unlike the file path it has NO is_already_applied no-op rescue, so
identical old/new ALWAYS errors there. Mirror the new_string description
so the schema warns before the error fires (sibling-site parity with
tools/file_tools.py PATCH_SCHEMA).
2026-08-12 15:21:47 +05:30
Teknium
fce314eabd feat(skills): advisory SKILL.md convention linter on create
Adds tools/skill_linter.py — a soft companion to the hard frontmatter
validator. It encodes the CONTRIBUTING 'Skill authoring standards
(HARDLINE)' conventions that today only a human reviewer catches:

- shell-utility references in prose (`grep`/`sed`/`cat`...) that should
  name the native tool (search_files/patch/read_file)
- missing version/author/license/metadata.hermes block
- name != directory, invalid name format
- description over the 60-char prompt budget, marketing words
- dangling references/ links, forbidden scaffolding files
- POSIX-only script primitives without a platforms: gate

Findings are ADVISORY. skill_manage(create) attaches them as
lint_warnings + lint_hint in the success result; nothing is blocked
(the hard rejects already run in _validate_frontmatter). A CLI
(python -m tools.skill_linter <dir>) exits 1 only on ERROR-severity
findings so CI can gate on structural breakage without failing on nits.

Calibrated against the bundled skills/ tree: 76 advisory findings, exit 0,
no false positives after excluding repo-root scripts/ refs.

Inspired by MiniMax Code's skill-creator lint step; adapted to our
existing validator + skill_utils rather than a parallel system.
2026-08-08 11:12:27 -07:00
Alex Fournier
3d5fcab70f Merge updated tool metrics into skill metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>

# Conflicts:
#	hermes_cli/observability/schemas/hermes.shared_metrics.v1.schema.json
#	scripts/smoke_nemo_relay_shared_metrics.py
#	tests/agent/test_skill_commands.py
#	tests/hermes_cli/test_relay_shared_metrics.py
#	tests/hermes_cli/test_relay_shared_metrics_runtime.py
#	tests/tools/test_skill_manager_tool.py
#	tests/tools/test_skill_usage.py
#	tests/tools/test_skills_tool.py
2026-07-31 07:30:30 -07:00
Ben Barclay
f5b68ad58b Merge origin/main into feat/hsp-sync-client
Resolves the PR's conflict with main (2252 commits). Two conflicts, both
"each side added an independent block in the same place" — kept both:

- gateway/run.py — the housekeeping loop. This branch adds the Skill Sync
  pulls inside the CURATOR_EVERY branch (12-space indent); main adds a
  stale-session auto-archive as a sibling `if` at loop level (8-space).
  Different scopes, so the naive union would have mis-nested the archive
  block into the curator branch; kept each at its own indent level.
- tools/skill_manager_tool.py — the _edit_skill result dict. This branch
  appends the org auto-propose note; main appends
  _add_description_prompt_preview(). Independent, order-insensitive.

No behaviour dropped from either side.

Verified: 3552 passed / 0 failed across 63 suites (scope regenerated to
include main's new maybe_auto_archive / _add_description_prompt_preview
consumers) via scripts/run_tests.sh. `hermes sync` and `hermes sync status`
still work against a live token, resolving the production plane default.

The Pyright Optional-parameter warnings in skill_manager_tool.py are
pre-existing on main (`content: str = None` etc.), not introduced here.
2026-07-29 13:01:36 -07:00
Alex Fournier
f1fd678e44 feat(observability): add Relay skill metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-07-29 12:20:34 -07:00
teknium1
f57cb2e482 refactor: use utils atomic writes in cron/skill_manager 2026-07-29 10:13:50 -07:00
kshitij kapoor
22492f0c46 refactor: extract atomic_write_text to utils.py; fix write-failure error handling
Deduplicates the mkstemp→fsync→atomic_replace pattern that existed in
three places: agent_import.py (added by #72983), MemoryStore._write_file,
and skill_manager_tool._atomic_write_text. All three now call a single
utils.atomic_write_text helper.

Also wraps the atomic_write_text call in _merge_memory_entries with
try/except OSError so a write failure records a per-item error instead
of propagating uncaught and aborting the entire import with no record.

Follow-up to #72983.
2026-07-29 16:49:07 +05:30
Ben Barclay
f9c4d835f9 refactor(sync): name the access gate for what it is (Nous admin)
The client called its gate "the DEV-PHASE gate (tool_gateway_admin)", which
reads as though Skill Sync is gated on an unrelated service's admin right.
It isn't. NAS populates that claim from Permissions.ADMIN_ACCESS — the global
portal admin permission that guards /admin/* — so the gate is "is this user a
Nous admin?". The claim is simply named for its first consumer, the tool
gateway.

Renamed on this side to say what it means, while keeping the wire string
(other services read it):

- DEV_GATE_CLAIM -> NOUS_ADMIN_CLAIM (value unchanged: "tool_gateway_admin",
  with a comment recording why the wire name differs).
- identity/status key dev_gate_ok -> nous_admin, across the client, the CLI
  consumers, and the tests.
- The module docstring now states where the claim comes from, that the wire
  name is misleading, and that this gate is pre-launch containment rather
  than the shipping entitlement — admin status conflates "may administer
  Nous" with "has Skill Sync enabled" and has no middle setting for a beta
  cohort. Choosing the real entitlement is left as a separate decision.

Naming only — no behaviour change, and no change to which accounts can sync.
The user-facing messages stay deliberately vague ("not enabled for your
account yet") rather than telling users they need portal admin.

Verified: 2346 passed / 0 failed across 56 suites via scripts/run_tests.sh;
`hermes sync status` against a live token reports "nous_admin": true. Zero
stale dev_gate_ok / DEV_GATE_CLAIM references remain.
2026-07-29 07:57:25 +10:00
Ben Barclay
4f990ec09e refactor(sync): put every Skill Sync verb under hermes sync; drop HSP naming
Encapsulates the feature behind one command for launch, and adopts the
official product name.

One command:
- `propose` moves from `hermes skills propose` to `hermes sync propose`, so
  the whole feature is one command to learn and one to document. Its handler
  moves from cmd_skills to cmd_sync accordingly.
- The `hermes sync` parser now documents both halves plainly: personal sync
  across your devices, and sharing with your organisation. Added an examples
  epilog; rewrote the verb help in user language ("Include a skill in your
  sync" rather than "Opt a skill into sync").
- Every user-facing string that pointed at `hermes skills propose` now points
  at `hermes sync propose` (8 sites, including the agent-visible guidance
  returned by skill_manage and the org provenance header).

This also clears the way for #39343, which adds its own top-level `sync` for
git-repo profile backup — that feature nests under `skills`, this one owns
`sync`.

Naming:
- HSP / "Hermes Sync Protocol" is gone from prose, docstrings, and comments.
  The feature is "Skill Sync".
- Public identifiers renamed: HSPClient -> SyncClient, HSPError -> SyncError,
  HSPConflict -> SyncConflict, hsp_address -> wire_address, HSP_VERSION ->
  WIRE_VERSION.
- The WIRE names are deliberately NOT renamed: the `hsp_version` capability
  field and the `x-hsp-object-type` response header are set by the deployed
  gateway-gateway sync plane (verified in src/sync/syncRouter.ts), so
  renaming them client-side would break sync against a live server. A comment
  at the version constant records why they differ from the product name.
- The version-mismatch error is now actionable ("this server speaks sync
  version X, but this Hermes speaks Y — update Hermes to sync with it")
  instead of leaking the protocol acronym.

Also fixes a wiring gap found on the way: the gateway housekeeping tick
pulled personal skills but never org skills — the same defect already fixed
for the CLI. Org pull now runs there too, gated on real org membership.

Tests: the jargon guard now also fails on a bare "HSP". The two tests that
asserted the old cross-command structure are replaced by three asserting the
new one (propose IS under sync, propose is NOT under skills, sync usage
lists it). 2294 passed / 0 failed across all 51 suites that import the
changed modules, via scripts/run_tests.sh.

Verified by running the real CLI: `hermes sync --help` lists all eight verbs,
`hermes skills --help` no longer mentions propose, `hermes sync propose
--help` parses, and `hermes sync status` still reports live org state.
2026-07-29 07:47:06 +10:00
Ben Barclay
981feb6730 feat(skills): org skills are editable in place; local edits survive org updates
The read-only org mirror broke the learning loop precisely where it matters
most. The system prompt tells every agent to patch a skill the moment it
finds a gap, and shared skills are the ones the most people use — but every
write to _org/ was refused, and the curator was excluded from them outright.
So org skills froze while personal skills kept improving, and the offered
alternative ("fork it into a personal skill, then propose the fork") is not
something an agent does mid-task. The refusal WAS the feature; improvements
were simply lost, and manual forks would have fragmented the shared set.

Edit in place:
- skill_manage patch/edit/write_file now work on org skills. Only delete is
  still refused (the mirror is a view of org HEAD — a local delete returns on
  the next pull; removing a shared skill is an admin action).
- Org skills are curation-eligible again, so the curator can improve the
  highest-leverage skills in the system instead of skipping them.
- The load-time provenance header now says edits are allowed and kept,
  instead of instructing the agent not to edit.

Local edits are never overwritten:
- pull_org_skills previously rmtree'd each skill dir and re-materialized it,
  silently destroying local work on the next session start. It now records a
  content fingerprint per skill (.org-baseline.json) when it writes one, and
  SKIPS any skill whose local content diverges from that baseline.
- When upstream ALSO changed such a skill, it is reported in the pull
  result's "conflicted" list and left untouched for the user to resolve
  deliberately (propose the local version, or delete it and re-pull to take
  theirs). A missing baseline is treated as unmodified so pre-existing
  mirrors do not raise phantom conflicts.
- Fingerprints are content-based (path + bytes, sorted), so a touch/mtime
  change is not mistaken for an edit.

Sharing back:
- Default: the edit stays local and the tool result tells the user to run
  "hermes skills propose <skill>".
- Opt-in sync.org_auto_propose / HERMES_SYNC_ORG_AUTO_PROPOSE submits each
  edit immediately. Defaults OFF — pushing every agent edit to a whole
  organisation is not a safe default. A failed submission never fails the
  edit; the change is saved and can be proposed later.
- "hermes sync status" lists org skills with unshared local edits;
  "hermes sync pull" reports conflicts it declined to overwrite.

Tests: 25 in the namespace suite (was 15). The two that asserted the old
read-only behaviour now assert the opposite. New coverage for edit-applied,
share-back guidance, delete-still-refused, curation-allowed, edit detection,
missing-baseline tolerance, mtime-insensitivity, and the auto-propose
default. 428 passed / 0 failed via scripts/run_tests.sh.

Verified through the REAL pull path against a mock plane: pull v1 -> edit in
place -> upstream ships v2 -> pull leaves the local edit intact, reports the
conflict, and surfaces it in status.
2026-07-28 09:08:04 +10:00