[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
The best-benchmarked open-source AI memory system. And it's free.
A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.
Context engine for large codebases, exposed through MCP. Gives AI coding agents precise repository context; benchmarked at frontier-agent quality with ~25x lower model cost and 45% fewer tokens with semantic search.
AI-powered contract review skill with CUAD risk detection, market benchmarks, and lawyer-ready redlines. Works with Claude Code, Codex, Cursor, and 26+ tools.
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
MCP server for predictive maintenance and machinery fault diagnosis. Gives AI assistants evidence-based vibration analysis - FFT, envelope, bearing fault detection, ISO 20816-3 severity - with a measured, blind CWRU benchmark. Local-first: raw signals never leave your machine. Includes a Claude Code plugin.
AI product management skills and plugin for Claude Code, Cowork, Codex & other AI agents: evidence-tagged PRDs, specs, requirements, RICE prioritization, backlog and roadmap scoring, product strategy, GTM launch plans, release verification, benchmark packs, UX/UI design prompts for web + mobile apps. Every claim sourced or labelled unsourced.
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
An MCP server that lets LLM agents play Civilization VI.
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.
🔬 A Researcher&Agent-Friendly Framework for Time Series Analysis. Train Any Model on Any Dataset!
Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
An agent skill that plays ARC-AGI-3. One rule: say what an action will do before you spend it. Claude Code on Opus 5 finished all 25 public games at 100.00 RHAE in 7,645 actions.
AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.