Sandbox
27 repos for learn · TestingClear
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83
haddock-development/
claude-reflect-system

Continual Learning & Self-improving skills system for Claude Code - learn from corrections, never repeat mistakes

376
liwala/
sheal

self-healing and self-learning loop for coding agents

87
rusinikita/
acid

SQL transactions learning tool (AI ready)

30
Xiangyue-Zhang/
auto-deep-researcher-24x7

🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring, Leader-Worker architecture, constant-size memory.

1.3k
appsecco/
vulnerable-mcp-servers-lab

A collection of servers which are deliberately vulnerable to learn Pentesting MCP Servers.

277

The first AI plugin that speaks first. Code-enforced learning + active forgetting + PAC (Proactive Accountability Challenge). Works with Claude Code, Gemini CLI, Hermes, OpenClaw.

88

Full computer-use for AI agents. Self-learning workflows. Native macOS. No screenshots required.

1.7k

Self-evolving browser automation for Codex, Claude Code, and Cursor—learn reusable site knowledge and safely replay verified routines with agent-browser.

102
azalio/
map-framework

Plan-then-build AI coding for Claude Code & Codex CLI — you approve the plan before the model writes a line of code. SPEC → PLAN → TEST → CODE → REVIEW → LEARN

156

A native Python agent CLI built on DeepAgents CLI, featuring an independent memory Agent that captures learnings after each task and delivers efficient AI coding assistance through hierarchical memory management.

155
notque/
vexjoy-agent

VexJoy AI Agent with Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop.

419
rohitg00/
pro-workflow

Claude Code learns from your corrections: self-correcting memory that compounds over 50+ sessions. Context engineering, parallel worktrees, agent teams, and 17 battle-tested skills.

2.9k
ruvnet/rufloHarnesses

🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated

72k
itsmostafa/
gskill

CLI that learns repository-specific Claude skills with evolutionary search.

33

Open-source self-improving QA agent for software teams. A test harness with memory. Write tests in natural language for web and mobile. agent-qa learns from every run, adapts to UI changes, and catches regressions before you ship.

909
lllllllama/
RigorPilot-Skills

README-first research reproduction skills with bounded execution, auditable evidence, and byte-preserving README annotations.

486
kayba-ai/
Kyoko

🔨 Kyoko is the all-in-one, fully local tool for debugging and improving your AI agents.

97
oneal2000/
SR-Agents

SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.

103
FrankS-IntelLab/
agentic-kaggle-skill

🤖 AI Agent-driven Kaggle competition workflow. Battle-tested patterns for score stabilization, submission troubleshooting, kernel workflows, and spec-driven development.

183
EndymionLee/
PilotBrowseMCP

A browser runtime that lets AI agents control your real Chrome browser via MCP. Agents can explore websites, generate operation manuals, and reuse them to save tokens. AI操控你的真实浏览器。

99
dstackai/
dstack

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

2.2k