Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
UiPath/coder_evalHarnesses
127
tirth8205/
code-review-graph
tirth8205/code-review-graphConnectors
Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.
31k
aisa-group/
skill-inject
aisa-group/skill-injectHarnesses
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
96
lmwilki/
civ6-mcp
lmwilki/civ6-mcpConnectors
An MCP server that lets LLM agents play Civilization VI.
172
pinecone-io/
cultivar
pinecone-io/cultivarTools
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.
37
morluto/flameoxConnectors
Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
121