Sandbox
15 repos for benchmarks · Any agent · ResearchClear
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83
oneal2000/
SR-Agents

SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.

103
Gen-Verse/
Skill-Entropy-RL

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

38
eli-labz/
Cognitive-Core-Skills

A universal, industry-neutral taxonomy of cognitive core skills (perception, memory, reasoning, planning, action, verification, learning, governance) for LLMs, SLMs, AI agents, and world models — with schemas, 159 skill cards, benchmarks, and CI.

165
XMUDeepLIT/
Awesome-Self-Evolving-Agents

A Survey of Self-Evolving Agents | A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Self-Evolving Agents.

425

A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.

172

MCP server for predictive maintenance and machinery fault diagnosis. Gives AI assistants evidence-based vibration analysis - FFT, envelope, bearing fault detection, ISO 20816-3 severity - with a measured, blind CWRU benchmark. Local-first: raw signals never leave your machine. Includes a Claude Code plugin.

83
naderelewa/
Product-to-Prod

AI product management skills and plugin for Claude Code, Cowork, Codex & other AI agents: evidence-tagged PRDs, specs, requirements, RICE prioritization, backlog and roadmap scoring, product strategy, GTM launch plans, release verification, benchmark packs, UX/UI design prompts for web + mobile apps. Every claim sourced or labelled unsourced.

42
lmwilki/
civ6-mcp

An MCP server that lets LLM agents play Civilization VI.

172
Ladbaby/PyOmniTSFrameworks & SDKs

🔬 A Researcher&Agent-Friendly Framework for Time Series Analysis. Train Any Model on Any Dataset!

98

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

646
agentvitals/
checkup

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.

95