The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Agent skill evolution harness for Claude Code and Codex
WikiSkill keeps a persistent wiki beside an evolving skill set, then runs agent tasks, distills failures into wiki pages, proposes skill patches, and gates them on held-out tasks. The repo is built around a real CLI, isolated agent profiles, and backend adapters for Hermes, Claude Code, Codex, and Copilot.
Builders who want reusable agent workflow rules that keep improving from real runs.
You can keep useful lessons across sessions and only accept skill changes that beat the current baseline.
What it does
Persistent wiki loop
Stores raw traces, pattern pages, and skill impact notes so knowledge compounds across iterations.
Maintainer and proposer agents
Uses one agent to turn traces into wiki patterns and another to propose skill changes from that wiki.
Strict validation gate
Accepts a skill change only when held-out validation is strictly better than the best score so far.
Isolated agent profiles
Runs each workspace in its own agent home directory so candidate skills are tested without touching the real profile.
Backend adapters
Supports Hermes, Claude Code, Codex, and GitHub Copilot CLI with the same workspace flow.
Installable skill packs
Ships SKILL.md packages under `skills/` that can be installed into Hermes as distilled patterns.
How to get it
- 1Run
pip install wikiskill # from PyPI (wheel + sdist, Python ≥3.10) wikiskill init demo # workspace + 22-task auto-graded bench (13 train / 9 val) wikiskill status wikiskill evolve --iters 3 # full Algorithm 1 loop with your default model
- 2This repo doubles as a Hermes skills tap — the patterns distilled from live evolution…
hermes skills tap add ashutoshsinghpr7/wikiskill hermes skills install ashutoshsinghpr7/wikiskill/skills/wikiskill-evolve hermes skills install ashutoshsinghpr7/wikiskill/skills/search-miss-binary # ...one `install` per skill you want
README
🧠 WikiSkill
Compile agent experience into a persistent wiki — and let skills evolve themselves.
📚 Docs site: ashutoshsinghpr7.github.io/wikiskill · arXiv: 2608.27454
A faithful, production-minded implementation of WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution (arXiv:2608.27454, Google Research). The loop is agent-agnostic — Hermes Agent is the reference backend (built natively), Claude Code, Codex and GitHub Copilot CLI ship in the box, and OpenCode is on the roadmap (issue #13). Your agent becomes both the student and the teacher.
What this is
Agents fail. They also learn — but the lessons usually die with the session. WikiSkill fixes that by keeping a persistent knowledge wiki alongside the skill set, and running a closed evolution loop:
- The agent runs training tasks with its current skills → raw execution traces
- A Wiki Maintainer agent distills the traces into pattern pages (root causes, fixes)
- A Skill Proposer agent reads the wiki + traces and proposes one skill change (create or patch)
- Gating: the change is validated on held-out tasks — strictly better than the best score so far → kept; otherwise rolled back. The wiki is never rolled back.
Over iterations, knowledge compounds in the wiki while only proven improvements touch the skills.
┌──────────────────────────────────────────────────────┐
│ EVOLUTION LOOP (Algorithm 1) │
│ │
tasks ─────► │ Inference Agent ──► raw/traces/ (immutable) │
│ │ │
│ ▼ │
│ Wiki Maintainer ──► wiki/patterns/, index, log │
│ │ │
│ ▼ │
│ Skill Proposer ──► proposal (create/patch skill) │
│ │ │
│ ▼ │
│ GATE: val score > R_best? ──yes──► keep, R_best=R │
│ │ no │
│ ▼ │
│ rollback skills; wiki retained forever │
└──────────────────────────────────────────────────────┘
Why Hermes?
This is not a toy simulator. Every component is a real Hermes agent turn:
| WikiSkill (paper) | This repo |
|---|---|
| Inference Agent | hermes chat --oneshot in an isolated HERMES_HOME profile |
| Raw Layer | Full session JSONL transcripts, exported via hermes sessions export |
| Wiki Layer | wiki/ — git-tracked, maintained by a real agent, never rolled back |
| Skill Layer | Real SKILL.md packages (frontmatter + instructions), git-managed |
| Wiki Maintainer | Agent turn with the paper's Appendix E.2 prompt (extracted verbatim) |
| Skill Proposer | Agent turn with the paper's Appendix E.3 ReAct prompt |
| Gating | Strict R_val > R_best; git reset --hard on reject |
Why the isolated profile matters: gating is only meaningful if the agent sees exactly the candidate skill set. Each evolution workspace gets its own HERMES_HOME (bundled skills opted out, empty memory, skills symlinked per stage) — your real profile is never touched.
Quickstart (60 seconds)
pip install wikiskill # from PyPI (wheel + sdist, Python ≥3.10)
wikiskill init demo # workspace + 22-task auto-graded bench (13 train / 9 val)
wikiskill status
wikiskill evolve --iters 3 # full Algorithm 1 loop with your default model
Or from source: pip install -e . (installs the same wikiskill CLI).
That's it. Each evolution workspace lives at workspaces/<domain>/:
workspaces/demo/
├── raw/traces/iter-01/{train,val}/<task>.jsonl # immutable execution traces
├── wiki/ # persistent knowledge (never rolled back)
│ ├── index.md · log.md · skill-impact.md · patterns/*.md
├── skills/active/ # git-managed evolving skill set (S₀ = ∅)
├── skills/framework/ # maintainer + proposer agent skills
├── bench/tasks/<id>/ # task sandboxes (inputs + grader)
└── runs/ # per-run stdout, proposals, state
CLI
| Command | What it does |
|---|---|
wikiskill init <domain> [--backend claude] | Create workspace + demo bench (pins the agent backend) |
wikiskill bench --reset | Regenerate tasks (deterministic, seed=42) |
wikiskill status | Workspace state: scores, skills, wiki, history |
wikiskill evolve --iters N [--model M] [--provider P] [--max-turns N] [--no-early-stop] | The full loop (--model/--provider patch the isolated profile's default model, e.g. google/gemini-2.5-flash-lite + openrouter) |
wikiskill run-task <id> | Single inference rollout (debug) |
wikiskill compare <wsA> <wsB> [--iters N] | Paired statistical comparison: per-task win/loss/tie + two-sided exact-binomial p-value (answers "did the skill actually help?" — see docs/COMPARING.md) |
Bring your own tasks
Tasks are plain JSON (tasks.json); anything auto-gradable works:
{
"id": "spec-format1-1", "split": "train",
"title": "Format products according to spec",
"prompt": "Read spec.md and products.json...",
"sandbox": {"spec.md": "...", "products.json": "..."},
"grader": {"type": "exact", "file": "output.txt", "expected": "alpha|35|active\n..."}
}
Graders: exact, contains, json_field, code_stdout (runs the produced script). Missing deliverables score 0, never crash.
Live results so far
Honest numbers from real agent runs on the bundled bench:
| Setup | Baseline (S₀) | What happened |
|---|---|---|
| deepseek-v4-flash, 15 turns | 1.0 | Algorithm 1 early-stop — nothing to evolve |
| deepseek-v4-flash, 8 turns | 1.0 | same |
| deepseek-v4-flash, forced | 1.0 | proposer created spec_literal_transform → R_val=1.0, not > R_best → rejected |
| deepseek-v4-flash, forced | 1.0 | maintainer distilled 4 pattern pages (incl. execute_code blocked in sandbox, ripgrep binary misses); proposer created exact-match-sandbox-task → R_val=0.8889 (skill hurt) → rejected |
| gemma-3-4b (free, OpenRouter) | — | invalid run, thrown out — dead agent sessions were phantom-graded against stale sandboxes. The maintainer's pattern page caught the framework's own bug; fixed + regression-tested (see docs/RUNS.md Run 4) |
| gemini-2.5-flash-lite (free, OpenRouter), 8 turns | 0.6667 (real) | small model fails at S₀ → maintainer distilled 5 patterns → proposer created find-secret → R_val=0.4444, the skill hurt (2 regressions) → rejected. Full loop live on a genuinely weak model, ~$0.09/iteration |
| gemini-2.5-flash-lite, 3 iterations (issue #5) | 0.4444 (real) | compounding run: train as low as 0.2308, 6/48 launch failures (detected + honest 0.0s), maintainer distilled 1 pattern, proposer declined (no_action) — nothing to gate, r_best preserved. Honest negative: accumulation needs a stronger model (see docs/RUNS.md Run 6) |
The gating mechanism has caught both a neutral and a harmful proposal live. Full logs in docs/RUNS.md.
Design decisions worth knowing
--indoesn't pin the agent's CWD in single-query runs → every inference prompt embeds an absoluteWORKING DIRECTORYand forbids exploring outside it.- Sessions live in
state.db, not loose files → transcripts are materialized viahermes sessions export --format jsonl. - Rejected proposals are never lost — their full content is embedded in
wiki/skill-impact.mdso future proposers don't repeat them (per Appendix E.3). - The demo bench has traps: subtle-spec tasks and multi-bug debug scripts whose bugs don't compensate (verified at generation time).
How this compares to other community implementations
We audited the three repos that appeared alongside the paper (see docs/RUNS.md). This is the only one that: runs on a real agent stack (Hermes), gates skills through a fully isolated profile, ships verbatim Appendix E prompts, and has a live-verified end-to-end loop (maintainer → proposer → gate → rollback).
Agent backends
The loop runs on any supported agent CLI — the raw/wiki/skill layers are backend-agnostic (issue #13).
| Backend | Pin a workspace | Notes |
|---|---|---|
hermes (default) | wikiskill init demo --backend hermes | reference implementation; isolated HERMES_HOME per workspace |
claude | wikiskill init demo --backend claude | Claude Code 2.x (claude -p), isolated CLAUDE_CONFIG_DIR, transcripts normalized from the stream-json output; claude auth login required once |
codex | wikiskill init demo --backend codex | OpenAI Codex (codex exec --json --full-auto), isolated CODEX_HOME, transcripts from session JSONL; codex login once + binary on PATH |
copilot | wikiskill init demo --backend copilot | GitHub Copilot CLI (copilot -p -s), isolated COPILOT_HOME, skills symlinked into the sandbox's .github/skills/; copilot login once (or GH_TOKEN) |
Each workspace pins its backend in workspaces/<domain>/workspace.json; switch
anytime with wikiskill evolve <domain> --backend claude. Skills evolved on
one backend transfer to another via wikiskill transfer (same SKILL.md format).
Skills tap — install the distilled patterns in Hermes
This repo doubles as a Hermes skills tap — the patterns distilled from live evolution runs, installable in Hermes:
hermes skills tap add ashutoshsinghpr7/wikiskill
hermes skills install ashutoshsinghpr7/wikiskill/skills/wikiskill-evolve
hermes skills install ashutoshsinghpr7/wikiskill/skills/search-miss-binary
# ...one `install` per skill you want
| Skill | What it teaches | Paper layer |
|---|---|---|
wikiskill-evolve | Run the full evolution loop from inside Hermes | framework meta-skill (tooling doc, not a paper artifact) |
search-miss-binary | ripgrep silently skips binary files — verify empty results | maintainer wiki pattern (distilled from live run) |
script-exec-blocked | Sandbox approval policy: use file tools, not python3 -c | maintainer wiki pattern |
spec-literal-execution | Apply only the spec's literal clauses — no hidden transforms | maintainer wiki pattern |
trace-harness-launch-failure | Empty traces = launch failure, not agent behavior | maintainer wiki pattern |
verify-output-readback | Re-read the deliverable before finishing | maintainer wiki pattern |
Paper alignment, stated honestly: the five operational skills are the
maintainer's distilled patterns from real graded runs (docs/RUNS.md) — the
paper's wiki layer. None has been accepted by the validation gate yet (every
live gate so far was a rejection or no_action), so treat them as
well-evidenced raw material for the proposer/gate pipeline, not as
gate-approved skills. All ship as standard SKILL.md (agentskills.io-compatible).
Roadmap
Done:
-
comparecommand — paired exact-binomial run comparison (#6) - Skill transfer across workspaces/models (#7)
- Cron-driven overnight evolution (
hermes cron, 01:00 IST nightly) + docs/CRON.md - Multi-agent backends: Hermes (reference) + Claude Code, Codex, GitHub Copilot CLI (#13, #15, #24)
- GitHub Pages — custom animated docs site (#18)
- PyPI package
wikiskillvia tokenless trusted publishing (#20) - Multi-iteration compounding run — honest negative documented (Run 6)
Planned:
- Codex backend — done (#15)
- OpenCode backend (#16)
- Cross-agent transfer demo — evolve on one agent, gate on another (#17)
- A live acceptance gate (proposal beats baseline on a real model) — still the open scientific question (see Run 6)
- Real-task domains — your recurring workflows as graded task packs (#2)
License
MIT — see LICENSE. Based on arXiv:2608.27454 (Google Research); all prompts in skills/ are adapted from the paper's Appendix E. Inspired by Karpathy's LLM Wiki.
Files in the repo
- .github
- assets
- docs
- skills
- tests
- web
- wikiskill
- .gitignore
- CONTRIBUTING.md
- LICENSE
- mkdocs.yml
- pyproject.toml
- README.md
Discussion (0)
Ask about usage, or say what you built with itSign in to join the discussion.
No comments yet. Be the first to say what this is good for.
More harnesses
The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
from vibe coding to agentic engineering - practice makes claude perfect
🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani and Boris Cherny). Includes loop-audit, loop-init, loop-cost.
Git. Ship. Done - Core