Sandbox
4 repos for eval · Gemini CLI · SecurityClear

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
MCPJam/
inspector

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

2.2k
Evol-ai/
SkillCompass

Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.

216
FrancyJGLisboa/
agent-skills-platform

Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

2.4k