Sandbox
35 repos for verification · Any agent · TestingClear
clawplays/
ospec

Spec-driven, agentic workflow framework for AI coding agents. Turn a request into a verifiable goal loop — plan, act, verify — with durable specs and evidence in your repo. Works with Claude Code, Codex, Gemini, OpenCode, and plain CLI.

485
sv-number/
mcp-server

MCP server for AI agents that need a phone number: order a private number in 200+ countries, read the SMS verification code, hand it back. The widest country coverage in the category, and you can check it with one API call.

555

Claude Autoresearch Skill — Autonomous goal-directed iteration for Claude Code. Inspired by Karpathy's autoresearch. Modify → Verify → Keep/Discard → Repeat forever.

6.3k

Containment for AI agents - user isolation, sandboxed execution, network controls, backup/rollback. TLA+ verified.

173
OpenLAIR/
OpenSkill
OpenLAIR/OpenSkillFrameworks & SDKs

Open-World Self-Evolution for LLM Agents — agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)

88
hongnoul/hwatuConnectors

Fast, interruptible verification browser for AI coding agents: 35 ms checks, pixel diffs, live human hand-off

80
Agile-V/
agile_v_skills

Agent Skills for traceable requirements, independent verification, human approval gates, and auditable AI-assisted engineering. Supports Claude Code, Cursor, VS Code, and GitHub Copilot; ISO 9001/27001 aligned (design phase), GxP-aware.

53
vje013/
darwin-agentic-cloud

Verifiable and free cloud compute for AI agents. webMCP + MCP native. Check out our sandboxed Beta + research in the README

36
AbyssCN/
oh-my-dag

The orchestration layer under your coding agent. Turns work into a typed graph, runs one model per node, and takes the verdict from outside the model — exit codes, write-set checks, and a cross-family verifier. MCP server, 50 tools, bring your own models.

39

Headless product design for AI coding agents, backed by a transactional product graph | Design how it works, verify what you ship.

114

HAR: open agent harness (CLI + MCP) for coding agents. Isolated worktrees, deterministic verify, software factory workflows for Claude Code, Cursor, and Codex.

88
antibrow/
anti-detect-browser-skills

Launch and manage anti-detect browsers with unique real-device fingerprints for multi-account operations, web scraping, ad verification, and AI agent automation.

218
arian-gogani/
nobulex
arian-gogani/nobulexFrameworks & SDKs

Prior direction, kept rather than deleted. Signed, offline-verifiable receipts for AI agent actions, and a reference implementation of the OWASP Agentic Skills Top 10 AST09 receipt pattern. Nobulex is now the independent reliability registry for agent tools: github.com/arian-gogani/nobulex-registry

39
HenryZ838978/
deepseek-harness

Protocol-layer harness for DeepSeek: Python witness stack — posterior verification that keeps the protocol honest. dsh doctor --node probes included.

49
tathagat22/
plumb-mcp

Local Figma MCP server with no REST rate limits, no metered tool-call quotas, and a verification loop. Drop-in alternative to Figma's Dev Mode MCP and Framelink for Claude Code, Cursor, Windsurf — works on every plan including Free.

78
Vuk97/
forward-implementation-first

Stop your coding agent from stalling real work on self-invented bookkeeping - receipts, hashes, locks, certification rituals. Ship first, then verify. Skill for Claude Code, Codex, and other agents.

166
amplifthq/
opentag

Mention any ACP coding agent from Slack, GitHub, GitLab, Linear, or Lark. OpenTag runs Claude Code, Codex, Cursor and more on your own machine, then replies in-thread with verified, evidence-backed results.

1.4k
karanb192/
awesome-claude-skills

🎯 The definitive collection of 50+ verified Awesome Claude Skills for Claude Code, Claude.ai, and API. Boost productivity with TDD, debugging, git workflows, document processing, and more. Community-driven, actively maintained.

508
hookdeck/
webhook-skills

Webhook integration skills for AI coding agents (Claude Code, Cursor, Copilot). Step-by-step guidance for setting up webhook receivers, signature verification, and event handling for Stripe, Shopify, GitHub, and more. Built on the Agent Skills specification.

84

The multi-agent harness that checks the work: verifies agent runs by artifacts (stop-hook gates, independent judges, append-only event logs) across Claude Code, Codex, Cursor, and 10+ runtimes.

1.3k
AMAP-ML/
LongHorizon-Harness

The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.

1.5k

TypeScript multi-agent framework that runs in your own environment: consequential actions wait for approval and every run leaves a verifiable record. Describe the goal, not the graph. 13 built-in providers (Claude, OpenAI, Gemini, DeepSeek and more) plus any OpenAI-compatible endpoint, local models included.

6.9k
aimasteracc/
tree-sitter-analyzer

Cross-language-safe code-intelligence MCP for AI agents -- 13 languages, family-gated call graph (CodeGraph 745 vs TSA 6 cross-language mis-wires, v1.21.0 same-session -- ~124x by count, ~390x by rate; re-verifying vs current develop). Run miswire-audit on your repo. 8 facade tools, TOON output, 100% local. Python.

49