Sandbox
72 repos for verification · TestingClear
gokeshenzhen/
awesome-formal-verification-skill

Open-source, agent-agnostic RTL formal verification skill library for AI coding agents: JasperGold FPV, SVA, proof optimization, and TCL workflows for Claude Code, Codex, Gemini CLI, and Cursor.

33
clawplays/
ospec

Spec-driven, agentic workflow framework for AI coding agents. Turn a request into a verifiable goal loop — plan, act, verify — with durable specs and evidence in your repo. Works with Claude Code, Codex, Gemini, OpenCode, and plain CLI.

485
sv-number/
mcp-server

MCP server for AI agents that need a phone number: order a private number in 200+ countries, read the SMS verification code, hand it back. The widest country coverage in the category, and you can check it with one API call.

555
Da7-Tech/
SureForge

Agent Skill for complex work: research before asking, ask before planning, plan before building, verify before delivering, independent review before calling it done. Plain text, no runtime.

84

Claude Autoresearch Skill — Autonomous goal-directed iteration for Claude Code. Inspired by Karpathy's autoresearch. Modify → Verify → Keep/Discard → Repeat forever.

6.3k
code-yeongyu/
lazycodex

The one and only agent harness for complex codebases. Project memory, planning, execution, and verified completion inside Codex.

3.4k

Containment for AI agents - user isolation, sandboxed execution, network controls, backup/rollback. TLA+ verified.

173
VibeCodingWithPhil/
agentwise

Multi-agent orchestration for Claude Code with 15-30% token optimization, self-improving agents, and automatic verification

46
OpenLAIR/
OpenSkill
OpenLAIR/OpenSkillFrameworks & SDKs

Open-World Self-Evolution for LLM Agents — agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)

88
fivetaku/
fablize

A Claude Code plugin that makes Opus behave like Fable — completion, evidence, and verification enforced as procedure. Ships only what a Fable-vs-Opus comparison proved transferable.

895

Self-evolving browser automation for Codex, Claude Code, and Cursor—learn reusable site knowledge and safely replay verified routines with agent-browser.

102
hongnoul/hwatuConnectors

Fast, interruptible verification browser for AI coding agents: 35 ms checks, pixel diffs, live human hand-off

80
Agile-V/
agile_v_skills

Agent Skills for traceable requirements, independent verification, human approval gates, and auditable AI-assisted engineering. Supports Claude Code, Cursor, VS Code, and GitHub Copilot; ISO 9001/27001 aligned (design phase), GxP-aware.

53
vje013/
darwin-agentic-cloud

Verifiable and free cloud compute for AI agents. webMCP + MCP native. Check out our sandboxed Beta + research in the README

36
AbyssCN/
oh-my-dag

The orchestration layer under your coding agent. Turns work into a typed graph, runs one model per node, and takes the verdict from outside the model — exit codes, write-set checks, and a cross-family verifier. MCP server, 50 tools, bring your own models.

39
boringmarketer/
kimi-first

Claude Code skill: Kimi types, Claude thinks & verifies — route implementation work-orders to the Kimi Code CLI (k3). A port of @steipete's codex-first (github.com/steipete/agent-scripts). By @boringmarketer · boringmarketing.com

46

Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic contracts, not model opinion. DeepWatch is the agent workspace built on DeepSeek Harness. Python + npm, MCP, CLI, REST, Web.

362

Headless product design for AI coding agents, backed by a transactional product graph | Design how it works, verify what you ship.

114
povvo/
claudikins-kernel

SRE thinking applied to Claude Code, based on Boris Cherny's Q&A. It enforces a strict 4-stage pipeline with gates between each step. You literally cannot skip verification. You cannot ship without approval.

126
fubak/
ultraswarm

Multi-CLI agent swarm orchestrated by Claude Code: external AI CLIs code in isolated worktrees, Claude verifies and merges

83

HAR: open agent harness (CLI + MCP) for coding agents. Isolated worktrees, deterministic verify, software factory workflows for Claude Code, Cursor, and Codex.

88
Hanyuyuan6/
remote-gpu-trainer

An Agent Skill for the DL experiment lifecycle: RUN (a GPU you own or rent) → VERIFY the number is real → DELIVER reproducible, single-source figures and tables.

63
callstack/
agent-device

Mobile app automation and verification for AI coding agents. CLI, MCP server, and typed Node.js API for iOS, Android, HarmonyOS, TV, web, macOS, and Linux.

4.5k
duthaho/
claudekit

A verification-first engineering toolkit for Claude Code. Built for senior ICs and tech leads who already know how to ship production code — and want a workflow that keeps the discipline tight without getting in the way.

97