[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/
SkillLearnBench
83
waybarrios/
vllm-mlx
waybarrios/vllm-mlxTools
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
1.6k
joe960913/
Jixu
joe960913/JixuHarnesses
Durable single-Agent Harness for TypeScript: recoverable Threads, context continuity, explicit side effects, and a native TUI.
100
Tencent/
SkillHone
Tencent/SkillHoneSkills
Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue / PR / wiki. Runs on any agentskills.io runtime — Claude Code, Codex, OpenClaw, Hermes.
146
ciouskeila-hue/
cybercode-cli
Local AI agent web app and CLI for running tools, code, web scans, and media generation.
44