Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
UiPath/coder_evalHarnesses
127
hoangsonww/
Claude-Code-Agent-Monitor
🚀 A real-time monitoring dashboard for Claude Code & Codex, built with SQLite3, Node.js, Express, React, Vite, TailwindCSS, & WebSockets. It tracks sessions, agent activity, tool usage, and subagent orchestration, providing live analytics, a Kanban status board, status notifications, a cute buddy, & an interactive web UI/MacOS/Windows native app.
989
LGDiMaggio/predictive-maintenance-mcpConnectors
MCP server for predictive maintenance and machinery fault diagnosis. Gives AI assistants evidence-based vibration analysis - FFT, envelope, bearing fault detection, ISO 20816-3 severity - with a measured, blind CWRU benchmark. Local-first: raw signals never leave your machine. Includes a Claude Code plugin.
83