Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
UiPath/coder_evalHarnesses
127
jcaromiq/gokuTools
Goku is an HTTP load testing application written in Rust
148
Tiger3807861189/
GLM-5.3-Flash-J-Space-Capability-Realization-Report
GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite
1k
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
646
agentvitals/
checkup
agentvitals/checkupSkills
AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.
95