Sandbox
48 repos for evals · Claude CodeClear
evalstate/fast-agentFrameworks & SDKs

Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support

3.9k
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
rpamis/cometHarnesses

Comet: agent skill harness for turning ideas into evaluated workflows

3k
skillberry-ai/
cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

56
MCPJam/
inspector

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

2.2k
alibaba/
skill-up

An evaluation and evolution tool for Agent Skills.

880

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

2.6k
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

2.8k

239 evaluated academic Claude/agent skills across 17 research domains (bioinformatics, data science, clinical, social-science methods, Turkish academia & more). Executable eval per skill, deterministic citation verifier, research→write→review→publish pipeline, and a skill-finder front door. Claude Code, Cursor, Codex, Gemini CLI & Copilot.

66
Evol-ai/
SkillCompass

Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.

216
ai-boost/
awesome-harness-engineering

Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.

4.1k
Sahir619/
fable-method

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

2.3k
tripleyak/
SkillForge

A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

890
Q00/ouroborosHarnesses

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

5.8k
thiientv/
godmode

Production-grade Agent Skills for AI coding agents—composable workflows for planning, TDD, debugging, review, UI/UX, releases, incidents, and evals.

94
RichSchefren/
atlas

Open-source local-first cognitive memory. AGM-compliant belief revision (49/49 postulates). When a fact changes, downstream beliefs are automatically re-evaluated, not just flagged.

80
SNL-UCSB/
paper-writing-skill

A Claude Code skill that encodes battle-tested editorial principles, section-specific rhetorical moves, and a structured writing pipeline for research papers. Brainstorm → Draft 0 → Evaluate → Write → Compress.

191

The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.

1.3k
tjboudreaux/
cc-thinking-skills

28 eval-informed mental models and critical-thinking skills for Claude Code, GitHub Copilot, Codex, Cursor, and other Agent Skills-compatible tools

1.3k