Sandbox
8 repos for benchmark · Tools · Any agentClear
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83
oneal2000/
SR-Agents

SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.

103
uber/
ADR

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

1.5k

Goku is an HTTP load testing application written in Rust

148
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

646