Sandbox
6 repos for benchmark · Gemini CLI · TestingClear

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
aisa-group/
skill-inject

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

96
lmwilki/
civ6-mcp

An MCP server that lets LLM agents play Civilization VI.

172
pinecone-io/
cultivar

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

37
morluto/flameoxConnectors

Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.

121