Sandbox
8 repos for benchmarks · Tools · TestingClear
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83
oneal2000/
SR-Agents

SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.

103
uber/
ADR

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

1.5k

Goku is an HTTP load testing application written in Rust

148
pinecone-io/
cultivar

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

37
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424