Sandbox
124 repos for testing · Harnesses · CodingClear
striderZA/
OpenCodeGameStudios

Agentic game development taken to the next level: plan, build and test across major engines with any coding agent🤖

85

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
tripleyak/
SkillForge

A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

890
FrancyJGLisboa/
agent-skills-platform

Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

2.4k
TheAstrelo/
Claude-Pipeline

Tool-agnostic 13-phase AI development pipeline — turns a task description into reviewed, committed code through automated design, adversarial review, security, test, and code-review gates. One bash engine, balanced Opus/Sonnet routing, self-healing commit review.

45
azalio/
map-framework

Plan-then-build AI coding for Claude Code & Codex CLI — you approve the plan before the model writes a line of code. SPEC → PLAN → TEST → CODE → REVIEW → LEARN

156
shinpr/
agentic-code

Agentic coding framework powered by AGENTS.md: systematic, test-first workflows with quality gates for Cursor, Codex, Gemini CLI, and AI coding agents.

49

Self-hosted AI agent harness in a single Go binary — writes, sandbox-tests and repairs its own tools, and lets Claude Code, Codex and any MCP client build and share them.

508
BaseInfinity/
claude-sdlc-harness

Self-evolving SDLC enforcement for AI coding agents — hooks, skills, and one-command setup for Claude Code. Plan before coding, test before shipping, escalate when uncertain. Measures itself getting better over time.

49
notque/
vexjoy-agent

VexJoy AI Agent with Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop.

419
mcp-security-standard/
mcp-server-security-standard

MCP Server Security Standard (MSSS): an open, testable security control standard for certifying MCP servers, with levels, evidence requirements, and reporting schemas.

74

Comprehensive sets of standards and practices designed to elevate the capabilities of AI coding agents.

156
AlekseiUL/
agentforge-openclaw

Create production-ready skills and agents for OpenClaw. 4-level memory, auto-improvement, 22 battle-tested pitfalls.

43
wasintoh/
toh-framework

Type one sentence, approve once, walk away — Toh Framework installs an AI build department (14 commands, 8 agents, 23 skills) into your project and keeps building, testing and fixing until it is verified done. Works in Claude Code, Cursor, Antigravity, Codex and ZCode.

96
KbWen/
agentic-os

Governance framework for AI coding agents. It runs them through a five-step workflow (plan, build, review, test, ship) where no step counts as done without evidence. Drop-in rules and guardrails for Claude Code, Codex, Cursor, Copilot, and Antigravity, via AGENTS.md.

164
addyosmani/
factory

A reference software factory for Claude Code and Codex

179

Professional context and harness engineering for Claude Code and OpenAI Codex. Build production-grade software with spec-driven development, TDD, persistent memory, quality gates, code intelligence, human oversight, and end-to-end verification.

2.1k
github/
gh-aw
github/gh-awHarnesses

GitHub Agentic Workflows

5.1k
LarsCowe/
bmalph

Unified AI Development Framework - BMAD phases with Ralph execution loop

406
owainlewis/
machinist

Open source software factory infrastructure for advanced AI coding workflows

398