A test runner for agentskills.io-style AI agent skills
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Agent Skills Evaluation Framework
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
An evaluation and evolution tool for Agent Skills.
[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
The evaluation benchmark on MCP servers
Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
Help your agents create better skills
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.