Sandbox
36 repos for benchmarkingClear
naderelewa/
Product-to-Prod

AI product management skills and plugin for Claude Code, Cowork, Codex & other AI agents: evidence-tagged PRDs, specs, requirements, RICE prioritization, backlog and roadmap scoring, product strategy, GTM launch plans, release verification, benchmark packs, UX/UI design prompts for web + mobile apps. Every claim sourced or labelled unsourced.

42
aisa-group/
skill-inject

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

96
lmwilki/
civ6-mcp

An MCP server that lets LLM agents play Civilization VI.

172
pinecone-io/
cultivar

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

37
Ladbaby/PyOmniTSFrameworks & SDKs

🔬 A Researcher&Agent-Friendly Framework for Time Series Analysis. Train Any Model on Any Dataset!

98
NVIDIA/
SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

424

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

646
morluto/flameoxConnectors

Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.

121
pbshgthm/
arc-skill

An agent skill that plays ARC-AGI-3. One rule: say what an action will do before you spend it. Claude Code on Opus 5 finished all 25 public games at 100.00 RHAE in 7,645 actions.

89
agentvitals/
checkup

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.

95