Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
Agent Skills Evaluation Framework

Comet: agent skill harness for turning ideas into evaluated workflows
[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
An evaluation and evolution tool for Agent Skills.
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.
A local MCP runtime that attacks what you own and only reports what it proved. 17 CVEs across 9 projects came out of this repo. Install: npx -y hacker-bob@latest install /path/to/project, then run /bob-evaluate target.com
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Multi-agent orchestration system for Claude Code with parallel execution, automated quality gates, Board of Directors, and bundled Superpowers skills
Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
Production-grade Agent Skills for AI coding agents—composable workflows for planning, TDD, debugging, review, UI/UX, releases, incidents, and evals.
AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.