Sandbox
29 repos for evaluation · Claude Code · CodingClear
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

2.6k
rpamis/cometHarnesses

Comet: agent skill harness for turning ideas into evaluated workflows

3k
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

2.8k
Evol-ai/
SkillCompass

Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.

216
alibaba/
skill-up

An evaluation and evolution tool for Agent Skills.

880
evalstate/fast-agentFrameworks & SDKs

Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support

3.9k
MCPJam/
inspector

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

2.2k
Q00/ouroborosHarnesses

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

5.8k
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73
RichSchefren/
atlas

Open-source local-first cognitive memory. AGM-compliant belief revision (49/49 postulates). When a fact changes, downstream beliefs are automatically re-evaluated, not just flagged.

80

The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.

1.3k
oxbshw/
LLM-Agents-Ecosystem-Handbook

One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.

546
comet-ml/
opik-mcp

Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude Code, Cursor, or VS Code.

217

Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)

71k
MemTensor/
skills-vote

SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution

301

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
Ibrahim-3d/
orchestrator-supaconductor

Multi-agent orchestration system for Claude Code with parallel execution, automated quality gates, Board of Directors, and bundled Superpowers skills

377
MrZoyo/
deslop-GPT

Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

125
Sahir619/
fable-method

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

2.3k
greyhaven-ai/
autocontext

a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task

1.3k
skillberry-ai/
cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

56