Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
Agent Skills Evaluation Framework
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
A local MCP runtime that attacks what you own and only reports what it proved. 17 CVEs across 9 projects came out of this repo. Install: npx -y hacker-bob@latest install /path/to/project, then run /bob-evaluate target.com
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.