frontier-harness-eval/
eval
frontier-harness-eval/evalHarnesses
Public results and task definitions for FrontierHarness Eval
183
Public results and task definitions for FrontierHarness Eval
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.