[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.
ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
The best-benchmarked open-source AI memory system. And it's free.
CLI tool for querying Apache Spark History Server REST API
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.