Sandbox
5 repos for benchmarks · Tools · Claude CodeClear
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83
uber/
ADR

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

1.5k
MemPalace/
mempalace

The best-benchmarked open-source AI memory system. And it's free.

59k
pinecone-io/
cultivar

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

37
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424