Sandbox
5 repos for benchmarking · SecurityClear
uber/
ADR

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

1.5k

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
aisa-group/
skill-inject

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

96
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424