Sandbox
111 repos for running · Any agent · TestingClear
openclaw/
crabbox

Crabbox: warm a box, sync the diff, run the suite.

1.4k

Runtime intelligence system that makes MCP servers debuggable, testable, and safe to run in production.

48
the-void-ia/
void-box
the-void-ia/void-boxFrameworks & SDKs

Composable agent runtime with enforced isolation boundaries

88
Muvon/
octomind

Open-source AI coding agent and agent runtime: one binary, any model, MCP-native. Runs in terminal, CI, or as a daemon.

135
dcouple/
Pane

Terminal-first, open-source AI agent manager for any CLI agent (agent agnostic), any OS (mac, windows, linux). The Open-Source Agentic Development Environment for running multiple coding agents in parallel. Run locally or self-host Remote Pane to manage agents from desktop or phone. Simplify multi-agent orchestration with the runpane CLI.

454
Tencent/
SkillHone

Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue / PR / wiki. Runs on any agentskills.io runtime — Claude Code, Codex, OpenClaw, Hermes.

146
ashutoshsinghpr7/
wikiskill

WikiSkill (arXiv:2608.27454) for Hermes Agent — self-evolving agent skills via a persistent knowledge wiki. Faithful Algorithm 1 implementation with real agent runs, isolated skill gating, and a documented live run log.

144
peakmojo/
agentic-mcp-client

A standalone agent runner that executes tasks using MCP (Model Context Protocol) tools via Anthropic Claude, AWS BedRock and OpenAI APIs. It enables AI agents to run autonomously in cloud environments and interact with various systems securely.

41

A local MCP runtime that attacks what you own and only reports what it proved. 17 CVEs across 9 projects came out of this repo. Install: npx -y hacker-bob@latest install /path/to/project, then run /bob-evaluate target.com

97
remorses/
playwriter

Chrome extension & CLI to let agents control your browser. Runs Playwright snippets in a stateful sandbox. Available as CLI or MCP

3.9k

Self-hosted control plane for AI agents: dispatch tasks, review runs, track spend, and operate OpenClaw, Claude Code, Codex, and other runtimes.

6.2k
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73
NVIDIA/
NeMo-Relay
NVIDIA/NeMo-RelayFrameworks & SDKs

Multi-language agent runtime and library for execution scope management, lifecycle events, and middleware on tool and LLM calls.

169
Happenmass/
omux

Orchestrate AI coding agents (Claude Code, Codex) as parallel subagents over tmux — a loop-engineering runtime with auto-continue, execute-then-review, and cross-session memory.

95

TypeScript multi-agent framework that runs in your own environment: consequential actions wait for approval and every run leaves a verifiable record. Describe the goal, not the graph. 13 built-in providers (Claude, OpenAI, Gemini, DeepSeek and more) plus any OpenAI-compatible endpoint, local models included.

6.9k

A Model Context Protocol (MCP) server that enables LLMs to run ANY code safely in isolated Docker containers.

121
risingwavelabs/
box0

Open-Source Platform for Subagents and Agent Teams. Long-running, collaborative, proactive.

81
SponsioLabs/SponsioFrameworks & SDKs

Deterministic safety solutions for probabilistic AI agents

440
uvwt/
agentdock
uvwt/agentdockConnectors

Secure MCP runtime for AI agents to operate local machines, servers, and containers with multi-device orchestration.

784

The multi-agent harness that checks the work: verifies agent runs by artifacts (stop-hook gates, independent judges, append-only event logs) across Claude Code, Codex, Cursor, and 10+ runtimes.

1.3k

Local-first, self-hosted AI agent runtime and MCP bridge with sandboxed sessions, memory, credentials, audit/replay, and a local Console.

642
mvschwarz/
openrig

Multi-agent harness that runs Claude Code and Codex together as one system

66
engasnm111/
lnwjud

lnwjud — local AI-agent runtime & MCP gateway

152