[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.
Curated Claude Fable 5 use cases, tutorials, integrations, demos, and benchmark evidence with source links and multilingual README files.
Cross-platform .NET performance engineering skill for coding agents, covering CPU, memory, GC, benchmarking, concurrency, startup, native profiling, GPU rendering, and production diagnostics on macOS, Windows, and Linux.
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
An Agent Skill helping you to optimize Xcode incremental and clean builds by running benchmarks and optimizing build settings.
ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
A universal, industry-neutral taxonomy of cognitive core skills (perception, memory, reasoning, planning, action, verification, learning, governance) for LLMs, SLMs, AI agents, and world models — with schemas, 159 skill cards, benchmarks, and CI.
GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite
A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.
Public results and task definitions for FrontierHarness Eval
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
CLI tool for querying Apache Spark History Server REST API
An MCP server that lets LLM agents play Civilization VI.
🔬 A Researcher&Agent-Friendly Framework for Time Series Analysis. Train Any Model on Any Dataset!
250+ real-world TypeScript AI projects: workflows, agents, and multi-agent systems with production-ready architecture, not chatbot demos.
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.