Sandbox
@gaasher/Agent-Loop-Skills

Agent Skills for verification-gated loops

This repo packages reusable workflows as open-standard Agent Skills. Each skill defines a loop that proposes a change, runs it in your environment, checks a real signal, keeps the change only if it improves the result, and repeats until a stopping condition is met.

168 stars19 forksPythonUpdated 2mo ago
Who it's for

Builders who want their agent to iterate on research, writing, analysis, optimization, or security tasks with a measurable gate.

What it delivers

You can turn one-off prompting into repeatable loops that keep only changes that pass your own checks.

What it does

Verification-gated loops

Each loop runs against a real signal such as tests, metrics, hashes, or a judged score before keeping a change.

Multiple workflow types

Includes loops for autoresearch, literature review, scientific writing, data analysis, code and SQL optimization, and red-team style testing.

Portable Agent Skills

The skills follow the open Agent Skills standard and are meant to work across hosts, with Claude Code verified and Codex and Cursor supported in skill form.

Claude Code plugin install

Ships `.claude-plugin/plugin.json` and `.claude-plugin/marketplace.json` so the collection can be installed as a Claude Code plugin.

Showcase runs and docs

Includes `showcase/` ledgers and charts plus `docs/` for compatibility, authoring rules, and API key setup.

How to get it

  1. 1Claude Code — plugin marketplace (add once, then install)
    /plugin marketplace add gaasher/agent-loop-skills
    /plugin install agent-loops@agent-loop-skills
  2. 2Any Agent-Skills host — the standard installers
    npx skills add gaasher/agent-loop-skills                   # auto-detects host, installs to the right dir
    gh skill install gaasher/agent-loop-skills --agent <host>  # claude-code | codex | cursor | …  (--pin, gh skill update)
  3. 3Manual — clone, then copy the loops into your host's skills dir (pick the line for your…
    git clone https://github.com/gaasher/agent-loop-skills
    
    cp -r agent-loop-skills/loops/* ~/.agents/skills/   # cross-tool: Codex, Cursor, Pi, OpenClaw, …
    cp -r agent-loop-skills/loops/* ~/.claude/skills/   # Claude Code
    # Hermes: hermes skills tap add gaasher/agent-loop-skills

README

agent-loop-skills

Loop until it's better — drop-in agentic loops, packaged as open-standard Agent Skills.

Autoresearch · scientific writing · data analysis · code/SQL/prompt optimization · red/blue/purple-teaming — each a generic, reusable loop you bind to your own task at invocation time, that iterates against a real signal until the work is actually better.

Agent Skills: open standard Works in Claude Code status: experimental PRs welcome License: MIT Stars


tournament-autoresearch improving a CIFAR-10 model from 0.734 to 0.798 val_acc over 11 iterations

A real run. The tournament-autoresearch loop on a CIFAR-10 model under a fixed 5-epoch budget — competing agents propose a change each step, a self-calibrating judge keeps the winners (green) and discards the regressions (gray): 0.734 → 0.798 val_acc, hands-off, 7 of 11 kept. Full ledger: showcase/tournament-autoresearch.
Far from SOTA by design — a deliberately tiny CNN at 5 epochs on a laptop GPU (Apple MPS). The demo is the loop's decision-making, not the absolute accuracy.


Why loops-as-skills

Two ideas collided in late 2025, and this repo lives in the overlap:

  • Skills became the portable unit. An Agent Skill is just Markdown + a little YAML that an agent loads only when relevant — "maybe a bigger deal than MCP … throw in some text and let the model figure it out" (Simon Willison). One SKILL.md now runs across ~30 hosts (Claude Code, Codex, Cursor, …).
  • The loop became the program. Karpathy ran ~700 autoresearch experiments in 2 days from one markdown prompt; Geoffrey Huntley's Ralph is, "in its purest form, a Bash loop." Agents get most of their power not from one clever prompt but from iterating against feedback.

This repo makes the loop be the skill. Instead of task-specific skills, each entry is a generic loop — program · artifact · feedback signal · run ledger · termination — that you bind to your task at invocation time. Paste your goal; the loop proposes a change, runs it in your environment, scores it on a real signal (tests, latency, a metric, a calibrated judge), keeps it only if it's better, logs it, and repeats.

The honest part: unsupervised agent loops are famous for spinning forever and confidently shipping garbage — at 90% per-step accuracy, a 5-step chain fails ~40% of the time. Every loop here is verification-gated: an objective feedback signal decides each step and an explicit termination condition ends it. That discipline — not autonomy for its own sake — is the point. (See Limitations.)

How a loop works

flowchart LR
  T["bind your task<br/>(artifact + signal + budget)"] --> P["propose<br/>one change"]
  P --> R["run it in<br/>your env"]
  R --> S{"score<br/>tests · metric · judge"}
  S -->|better| K["keep + log"]
  S -->|worse| X["revert"]
  K --> G{stop?}
  X --> G
  G -->|"plateau · budget · threshold"| B(["best artifact"])
  G -->|no| P

Every loop decomposes into the same five ingredients — program (SKILL.md), artifact slot (what's improved), feedback signal (what drives the next step), run ledger (append-only log), and termination (when to stop). Skills ship zero heavy dependencies: your code (a torch trainer, a SQL database, a dataset) runs in your environment via a bound run command; the skill shells out and reads the result. Multi-role loops use spawn-or-degrade — real isolated subagents on Claude Code, the same roles inline elsewhere.

Install

Any one of these installs all the loops:

Claude Code — plugin marketplace (add once, then install):

/plugin marketplace add gaasher/agent-loop-skills
/plugin install agent-loops@agent-loop-skills

Loops install namespaced as agent-loops:<name> (e.g. agent-loops:karpathy).

Any Agent-Skills host — the standard installers:

npx skills add gaasher/agent-loop-skills                   # auto-detects host, installs to the right dir
gh skill install gaasher/agent-loop-skills --agent <host>  # claude-code | codex | cursor | …  (--pin, gh skill update)

Manual — clone, then copy the loops into your host's skills dir (pick the line for your host):

git clone https://github.com/gaasher/agent-loop-skills

cp -r agent-loop-skills/loops/* ~/.agents/skills/   # cross-tool: Codex, Cursor, Pi, OpenClaw, …
cp -r agent-loop-skills/loops/* ~/.claude/skills/   # Claude Code
# Hermes: hermes skills tap add gaasher/agent-loop-skills

Then just describe your task — the host loads the matching loop. Research loops also call the shared literature-search skill; installing everything puts it alongside them, and any loop degrades gracefully (to WebSearch) if it's absent.

Loops in action

Most skill repos tell you what a skill is. Here's what these loops actually do — real Sonnet runs, full ledgers in showcase/.

🧑‍⚖️ tournament-autoresearch — competing ideas, a self-calibrating judge

<n> agents pitch competing changes each step; a judge critiques them, picks one, runs it, and recalibrates by comparing its predicted vs realized gain. On a CIFAR-10 SmallCNN under a fixed 5-epoch budget it climbed 0.734 → 0.798 val_acc, keeping 7 of 11 changes and reverting all 4 that regressed — escaping the plateaus a single-thread loop gets stuck on. (That's the run charted up top.)showcase/tournament-autoresearch

🔬 ml-autoresearch — analysis-first, every change traced to a cause

This loop reads inside each run — gradient flow, dead neurons, the loss curve — and grounds the next change in that evidence rather than guessing: "FC grad 57% vs first conv 3.3% — severe imbalance; 54% dead neurons" → add BatchNorm; "cosine schedule fixed the epoch-3 dip entirely (monotonic!), +0.033". It also reverts what hurts (augmentation, over-aggressive LR). The point isn't a leaderboard number — it's that every accepted change has a measured reason behind it. → showcase/ml-autoresearch

📊 data-analysis — findings with a number behind every one

Hypothesis → verify, stdlib-only. On a planted dataset it surfaced 3 real findings and correctly refuted 2, with effect sizes matching ground truth and no hallucinations: enterprise vs consumer order value 184.90 vs 109.16 (Cohen's d = 2.13), mobile return rate 32.8% vs 8.2% (RR 4.0) — and it reversed a plausible-but-wrong claim once it spotted a mobile confound. → showcase/data-analysis

More real runs — optimize-loop, research-proposal, red-team, power-analysis…
LoopWhat the run did
optimize-loopCorrectness-gated speedup: a SQLite query 1,131.75 ms → 1.055 ms (~1,073×), result-set hash matching baseline on every kept iteration; in code mode cut cyclomatic complexity 23 → 15 (nesting 7 → 3) with 13/13 tests green.
research-proposalScholarEval graded a proposal against the literature; Judge + Reviser iterated grade 45 → 84 (soundness 2→4, contribution 1→4) over 5 rounds.
scientific-figureSame ImageNet top-1-accuracy bar-chart brief, with vs without the loop: a single call truncated the y-axis at 50% and used non-paper numbers; the loop verified every value against the arXiv papers, flagged GoogLeNet's borrowed top-1, and iterated 80 → 96 (PASS).
red-teamAgainst a naive content filter, surfaced all 5 planted weaknesses (case bypass, leetspeak, spacing, synonyms, over-block) — 39 bypasses + 6 over-blocks — with a one-line root-cause fix each.
power-analysisSolved n = 100/group for 80% power via Monte-Carlo, fixed all 6 validity flaws, and emitted a full pre-registration.
research-questionSharpened 5 vague drafts → 3 strong questions (≥75), with real web novelty checks pivoting already-answered questions toward the open sub-problem.

The loops

= multi-role (real subagents on Claude Code, inline elsewhere). Browse any folder for its SKILL.md.

Autoresearch — iterate on an ML artifact against a metric
LoopWhy you'd reach for it
karpathyThe minimal baseline — propose, train, keep-if-better, loop. A faithful nod to Karpathy's autoresearch.
ml-autoresearchAnalysis-first: diagnoses each run and grounds the next change in evidence. A literature dial adds paper-grounded changes.
exploratory-autoresearchForces broad exploration via a temperature/swing scheduler — escapes hill-climbing one idea forever.
tournament-autoresearchCompeting changes judged each step by a self-calibrating judge.
dueling-autoresearchTwo approaches race the same metric in parallel and borrow ideas across lanes.
alpha-evolvePopulation-based evolution (MAP-Elites + islands, diff-mutate, cascade-eval).
Literature & writing · Data · Code & optimization · Security · Other (click to expand)

Literature & writing

LoopWhy you'd reach for it
literature-searchShared toolchain (not a loop): paper discovery, snippets, citation-graph, full-text over Semantic Scholar + arXiv.
literature-surveyBuilds a saturating evidence/contradiction matrix of sources × claims.
research-questionSharpens a vague topic into strong, novel, feasible research questions.
hypothesis-genGenerates and literature-vets a pool of research hypotheses.
research-proposalGrades a proposal against the literature (ScholarEval) and revises until it passes.
scientific-writerSpecialist judges + an independent peer-reviewer critique a draft; a writer revises until the score clears a bar.
scientific-figureA generator drafts a publication figure from your data/brief; an adversarial critic grades it against a fixed rubric — both can consult the literature (S2/arXiv) — and the two iterate until it's paper-ready.

Data

LoopWhy you'd reach for it
data-analysisHypothesis → verify discovery; every finding backed by a reproduced number at a meaningful effect size.
anomaly-investigationDiagnoses the cause of a known anomaly by forming, testing, and eliminating candidates.
claim-verifyAdversarially verifies a results draft's claims against the underlying data.
tabular-cleanupCleans a messy table to an inferred data contract with deterministic checks.

Code & optimization

LoopWhy you'd reach for it
optimize-loopEvaluator-optimizer with a pluggable correctness gate + minimized metric — refactor code (tests green + complexity↓) or speed up SQL (identical results + latency↓).
prompt-optimizeEvolves a prompt against a user-supplied scoring command (a black-box oracle).
plan-loopRefines a prompt into an executable plan — first-principles decomposition → PR-sized tasks (deps, tests, subtasks) → a principal-engineer critique loop; emits plan.md + a validated tasks.json.
swe-loopExecutes plan-loop's tasks.json task-by-task — an Engineer subagent writes the code (source only), a QA subagent authors the tests + grades a strict simplicity/readability rubric (tests only); they loop until the task's tests pass, regression stays green, and quality holds, then commit.

Security — authorized testing only

LoopWhy you'd reach for it
red-teamAdversarial loop-until-dry that surfaces distinct failure classes of a system you own.
blue-teamThe defensive fixer — closes a failure catalogue's classes under a regression gate, then opens a PR.
purple-teamOrchestrates red → blue → re-verify until a fresh attack pass stays dry; opens a PR with the patch set.

Other

LoopWhy you'd reach for it
power-analysisSizes and pre-registers a two-arm experiment to hit a target statistical power.

Compatibility

Skills (a SKILL.md the model invokes) work broadly across the open standard. A loop dispatching a subagent from a role file at runtime is confirmed only on Claude Code — elsewhere multi-role loops () run their roles inline (still correct, just serial). Single-agent loops run fully everywhere.

HostSkillsLoop-dispatched subagents
Claude Code✅ real, isolated, parallel — verified
Claude Agent SDK
Codex CLI · Cursor➖ inline
Hermes · Antigravity · Pi · OpenClaw(reported)➖ inline

Full, citation-backed matrix and the precise "why subagents are Claude-Code-only" reasoning: docs/compatibility.md. Off Claude Code, don't rely on parallel subagent isolation.

Compose with other skill collections

These are self-contained, open-standard skills, so they coexist with any other collection — install both into the same skills dir and use them together. For example alongside K-Dense scientific-agent-skills:

npx skills add gaasher/agent-loop-skills              # these loops
npx skills add K-Dense-AI/scientific-agent-skills     # + a domain-skill library

They install as sibling folders (Claude Code namespaces each plugin; other hosts load all and pick by description). A loop's analysis step can invoke any other installed skill — the same mechanism the research loops use to call literature-search.

Roadmap

Loops are most useful when they're honest about what's next. PRs on any of these are very welcome:

  • Blue-teaming + communication interface — shipped as blue-team (the defensive fixer) and purple-team (find → fix → re-verify), with the pull request as the communication interface to the target's owner.
  • Per-loop sandbox/ eval cases committed for every loop, so anyone can reproduce a run end-to-end.
  • Host-specific subagent adapters (Cursor subagents, Hermes delegate_task) so multi-role loops get real isolation beyond Claude Code.
  • More domains — eval-harness optimization, refactoring-at-scale, agent-trace debugging.

See open issues and grab a good first issue.

Status & limitations

status: experimental Experimental — expect breaking changes. Pin a version if you need stability.

  • Non-deterministic. Treat every output as a draft to verify, not a result to trust. Workaround: seed where supported; verify against your own oracle.
  • No correctness guarantee. A loop can be confidently wrong. Workaround: every loop gates on a signal — keep a human in the loop and point it at checks you own.
  • Host-dependent. Real subagent isolation is verified only on Claude Code. Workaround: see Compatibility.
  • Cost & latency. Loops make many model calls. Workaround: start with a low iteration budget.

Non-goals (deliberate scope, not missing features): these are human-supervised loops, not fully-autonomous agents; not a model or runtime (bring your own host); not domain-exhaustive — for a very custom workflow, fork a loop, that's what they're for.

Contributing — and a note on open source 💜

I love open source, and this repo is built to be added to. You don't need to be an expert and you don't need to write code — a sharper description, a new loop, a bug report, or a pasted run transcript all make it better. New loops are welcome, and so are wild ideas.

Start with CONTRIBUTING.md and the authoring rubric in docs/skill-authoring-rules.md; grab a good first issue. Be kind, have fun, and if a loop helped you, a ⭐ genuinely helps others find it.

Provenance / credits

Repo layout

agent-loop-skills/
├── loops/            # one self-contained, installable skill per folder (SKILL.md + tools/roles/schemas/rubrics/examples)
├── showcase/         # real archived runs (ledgers + results) behind the examples above
├── assets/           # generated progress charts
├── docs/             # authoring rules, compatibility, api-keys, authoring quickstart
└── .claude-plugin/   # plugin.json + marketplace.json (Claude Code plugin / marketplace)

⭐ If looping until it's better is your kind of fun, star it and send a loop.

Star History

Files in the repo

Repository payload12 top-level entries
  • .claude-plugin
  • .github
  • assets
  • docs
  • loops
  • showcase
  • .gitignore
  • CONTRIBUTING.md
  • LICENSE
  • pyproject.toml
  • README.md
  • uv.lock

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More skills

obra/
superpowers

An agentic skills framework & software development methodology that works.

285k
1 add

Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.

117k
1 add
Vincentwei1021/
anything2explainer

Topic in, narrated explainer video out. A Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar. Chinese or English; every frame drawn in code with Remotion.

666

Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)

71k