Sandbox
19 repos for lear · Any agent · TestingClear
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83
liwala/
sheal

self-healing and self-learning loop for coding agents

87
rusinikita/
acid

SQL transactions learning tool (AI ready)

30
appsecco/
vulnerable-mcp-servers-lab

A collection of servers which are deliberately vulnerable to learn Pentesting MCP Servers.

277

The first AI plugin that speaks first. Code-enforced learning + active forgetting + PAC (Proactive Accountability Challenge). Works with Claude Code, Gemini CLI, Hermes, OpenClaw.

88

Full computer-use for AI agents. Self-learning workflows. Native macOS. No screenshots required.

1.7k

A native Python agent CLI built on DeepAgents CLI, featuring an independent memory Agent that captures learnings after each task and delivers efficient AI coding assistance through hierarchical memory management.

155
notque/
vexjoy-agent

VexJoy AI Agent with Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop.

419

Open-source self-improving QA agent for software teams. A test harness with memory. Write tests in natural language for web and mobile. agent-qa learns from every run, adapts to UI changes, and catches regressions before you ship.

909
lllllllama/
RigorPilot-Skills

README-first research reproduction skills with bounded execution, auditable evidence, and byte-preserving README annotations.

486
kayba-ai/
Kyoko

🔨 Kyoko is the all-in-one, fully local tool for debugging and improving your AI agents.

97
oneal2000/
SR-Agents

SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.

103
FrankS-IntelLab/
agentic-kaggle-skill

🤖 AI Agent-driven Kaggle competition workflow. Battle-tested patterns for score stabilization, submission troubleshooting, kernel workflows, and spec-driven development.

183
EndymionLee/
PilotBrowseMCP

A browser runtime that lets AI agents control your real Chrome browser via MCP. Agents can explore websites, generate operation manuals, and reuse them to save tokens. AI操控你的真实浏览器。

99
dstackai/
dstack

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

2.2k

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

5.1k

High-performance AI pipeline engine with a C++ core and 50+ Python-extensible nodes. Build, debug, and scale LLM workflows with 13+ model providers, 8+ vector databases, and agent orchestration, all from your IDE. Includes VS Code extension, TypeScript/Python SDKs, and Docker deployment.

8.4k
FlyFission/
nuclear-grade-context-engineering

AI agents now operate with authority. Authority without discipline is how complex systems fail. Nuclear’s control loop, ported to AI-assisted software engineering.

33