Sandbox
51 repos for eval · CodingClear
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424
evalstate/fast-agentFrameworks & SDKs

Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support

3.9k

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

5.1k
agentscope-ai/
OpenJudge
agentscope-ai/OpenJudgeFrameworks & SDKs

OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

826
rpamis/cometHarnesses

Comet: agent skill harness for turning ideas into evaluated workflows

3k
skillberry-ai/
cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

56
MCPJam/
inspector

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

2.2k
alibaba/
skill-up

An evaluation and evolution tool for Agent Skills.

880

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

2.6k
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

2.8k
Evol-ai/
SkillCompass

Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.

216

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

646
google/adk-goFrameworks & SDKs

An open-source, code-first Go toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.

8.8k
ai-boost/
awesome-harness-engineering

Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.

4.1k
Sahir619/
fable-method

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

2.3k
tripleyak/
SkillForge

A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

890
Q00/ouroborosHarnesses

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

5.8k
trpc-group/trpc-agent-goFrameworks & SDKs

A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

1.8k