8ddieHu0314/
Skill-Lab
Agent Skills Evaluation Framework
56
Agent Skills Evaluation Framework
Public results and task definitions for FrontierHarness Eval
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.