Sandbox
682 repos for re · TestingClear
AaronZ345/
codebase-argus

Multi-agent codebase review for PRs, CI, and downstream fork syncs.

58

Full computer-use for AI agents. Self-learning workflows. Native macOS. No screenshots required.

1.7k
starbaser/
ccproxy

Build mods for Claude Code: Hook any request, modify any response, /model "with-your-custom-model", intelligent model routing using your logic or ours

346

The harness layer for Claude Code — a reference implementation of harness engineering with hook-enforced dual review, state-machine gates that survive context compaction, and fail-closed safety where it counts. Quality gates that AI can't skip.

188
athola/
claude-night-market

23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context optimization, research, and multi-LLM delegation. 186 skills, 128 commands, 54 agents.

337
seekrays/
mcp-monitor

A system monitoring tool that exposes system metrics via the Model Context Protocol (MCP). This tool allows LLMs to retrieve real-time system information through an MCP-compatible interface.

91
vdaubry/
bottega

Coding agent orchestration for engineering teams — shipped as a spec plus a working reference implementation.

84

AI code reviews grounded in 12 classic engineering books — decay risk diagnostics with book citations, severity labels, and 6 analysis modes including full-sweep auto-fix

1.5k
sentrux/
sentrux

Real-time architectural sensor that helps AI agents close the feedback loop, enabling recursive self-improvement of code quality. Pure Rust.

3.2k
Xopoko/
build-swift-apps

Build, debug, profile, test, refactor, and release Swift apps across iOS, macOS, Xcode, SwiftUI, SwiftPM, Tuist, and App Store Connect.

45
Zhen-Bo/
smell-check

Agent Skill for code and test smell audits. Evidence-ranked findings from Refactoring, Clean Code, and the test-smell literature. Formerly pragmatic-code-review.

237

Vestige enhances agents by deterministic root-cause retrieval that reaches backward through time to find the quiet change, decision, or service that caused today’s failure, not the lookalike.

617
agent-room-alkl/
agent-room

Put Claude Code, Cursor, Codex, Gemini & Antigravity in one room — real-time multi-agent AI collaboration over MCP for distributed dev, code review, PR handoff & frontend↔backend integration. Free, self-hostable, MCP-native.

48
openclaw/
crabbox

Crabbox: warm a box, sync the diff, run the suite.

1.4k
bjcoombs/
ai-native-toolkit

Claude Code plugin & Agent Skills for AI-native development: codebase readiness scoring (/assess), Six Thinking Hats deliberation (/huddle), AI-slop removal (/deslop), skill hardening (/skill-forge), and more.

30
clarity-digital-development/
tworkflow

A practical, no-hype workflow for AI coding agents: context, plan, implement, review, QA, ship, retro. Templates, two Claude Code skills, and a 40% context rule - every claim traced to official docs.

47
stefan-jansen/
claude-code-toolkit

Superseded by coding-agent-toolkit — see README. Preserved for reference: Claude-Code-only patterns from earlier iteration.

86
FailproofAI/
failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement. 40 built-in policies, a local dashboard, no account required with a generous free cloud plan

2.7k
vje013/
darwin-agentic-cloud

Verifiable and free cloud compute for AI agents. webMCP + MCP native. Check out our sandboxed Beta + research in the README

36
ferrislucas/
iterm-mcp

A Model Context Protocol server that executes commands in the current iTerm session - useful for REPL and CLI assistance

567

Find and repair substance defects in AI-assisted prose, code, docs, and agent output. Reports defects, never authorship. Structural tests over model judgement, because LLM judges agree with human slop labels at chance.

48

humanizer, but for code — an agent skill that removes AI-generated code slop: duplicated helpers, try-import fallbacks, broad excepts, speculative abstractions. Test-gated, behavior-preserving.

48

Real-time visualization of Claude Code agent orchestration — see your agents think, branch, and coordinate as they work.

1.6k