Sandbox
53 repos for verification · Any agent · CodingClear
sv-number/
skills

Give your AI agent a phone number: order a private number in 200+ countries over the API, read the SMS verification code, hand the number back. The widest country coverage in the category, checkable with one API call.

274
clawplays/
ospec

Spec-driven, agentic workflow framework for AI coding agents. Turn a request into a verifiable goal loop — plan, act, verify — with durable specs and evidence in your repo. Works with Claude Code, Codex, Gemini, OpenCode, and plain CLI.

485

Claude Autoresearch Skill — Autonomous goal-directed iteration for Claude Code. Inspired by Karpathy's autoresearch. Modify → Verify → Keep/Discard → Repeat forever.

6.3k

Containment for AI agents - user isolation, sandboxed execution, network controls, backup/rollback. TLA+ verified.

173
OpenLAIR/
OpenSkill
OpenLAIR/OpenSkillFrameworks & SDKs

Open-World Self-Evolution for LLM Agents — agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)

88
hongnoul/hwatuConnectors

Fast, interruptible verification browser for AI coding agents: 35 ms checks, pixel diffs, live human hand-off

80
Agile-V/
agile_v_skills

Agent Skills for traceable requirements, independent verification, human approval gates, and auditable AI-assisted engineering. Supports Claude Code, Cursor, VS Code, and GitHub Copilot; ISO 9001/27001 aligned (design phase), GxP-aware.

53
vje013/
darwin-agentic-cloud

Verifiable and free cloud compute for AI agents. webMCP + MCP native. Check out our sandboxed Beta + research in the README

36
AbyssCN/
oh-my-dag

The orchestration layer under your coding agent. Turns work into a typed graph, runs one model per node, and takes the verdict from outside the model — exit codes, write-set checks, and a cross-family verifier. MCP server, 50 tools, bring your own models.

39

Headless product design for AI coding agents, backed by a transactional product graph | Design how it works, verify what you ship.

114

HAR: open agent harness (CLI + MCP) for coding agents. Isolated worktrees, deterministic verify, software factory workflows for Claude Code, Cursor, and Codex.

88
nagisanzenin/
engram

Evidence-based learning engine — first-principles curricula, free-recall verification with receipts, FSRS-scheduled memory, and explorable artifacts. Learn anything; keep it.

1.4k
antibrow/
anti-detect-browser-skills

Launch and manage anti-detect browsers with unique real-device fingerprints for multi-account operations, web scraping, ad verification, and AI agent automation.

218
zhangqi444/
open-forge

Let your AI coding agent self-host any open-source app for you. 2,200+ verified recipes — provisioning, DNS, TLS, hardening. Works with Claude Code, Codex, Cursor, Aider, OpenClaw, Hermes.

98
aka-luan/
doc-cleanup

Agent skill that audits and cleans agent-facing markdown docs — archives completed-work history, fixes stale facts and contradictions, keeps only live verified facts.

42
arian-gogani/
nobulex
arian-gogani/nobulexFrameworks & SDKs

Prior direction, kept rather than deleted. Signed, offline-verifiable receipts for AI agent actions, and a reference implementation of the OWASP Agentic Skills Top 10 AST09 receipt pattern. Nobulex is now the independent reliability registry for agent tools: github.com/arian-gogani/nobulex-registry

39
HenryZ838978/
deepseek-harness

Protocol-layer harness for DeepSeek: Python witness stack — posterior verification that keeps the protocol honest. dsh doctor --node probes included.

49
gaasher/
Agent-Loop-Skills

Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills. Verification-gated; native on Claude Code, portable across Codex, Cursor & other Skills hosts.

168
tathagat22/
plumb-mcp

Local Figma MCP server with no REST rate limits, no metered tool-call quotas, and a verification loop. Drop-in alternative to Figma's Dev Mode MCP and Framelink for Claude Code, Cursor, Windsurf — works on every plan including Free.

78
wquguru/
harness-books
wquguru/harness-booksTutorials & Guides

📚 Two books on harness engineering — the design philosophies behind Claude Code & Codex: constraints, query loops, context governance, multi-agent verification. harness-books.agentway.dev

3.1k
solarch-dev/
solarch

Diagram→code through a deterministic rules gate: the AI proposes, 50 rules verify, only valid architecture lands. Try it: app.solarch.dev

47
Vuk97/
forward-implementation-first

Stop your coding agent from stalling real work on self-invented bookkeeping - receipts, hashes, locks, certification rituals. Ship first, then verify. Skill for Claude Code, Codex, and other agents.

166
amplifthq/
opentag

Mention any ACP coding agent from Slack, GitHub, GitLab, Linear, or Lark. OpenTag runs Claude Code, Codex, Cursor and more on your own machine, then replies in-thread with verified, evidence-backed results.

1.4k