Sandbox
33 repos for verification · Claude Code · Code reviewClear
Aurite-ai/
agent-verifier

Agent Verifier is a coding agent skill that verifies code against organizational policies, code quality patterns, security requirements, and framework best practices — before code ships. Works with Claude Code, Cursor, Windsurf, and 30+ agents.

44
Da7-Tech/
SureForge

Agent Skill for complex work: research before asking, ask before planning, plan before building, verify before delivering, independent review before calling it done. Plain text, no runtime.

84
fivetaku/
fablize

A Claude Code plugin that makes Opus behave like Fable — completion, evidence, and verification enforced as procedure. Ships only what a Fable-vs-Opus comparison proved transferable.

895

Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic contracts, not model opinion. DeepWatch is the agent workspace built on DeepSeek Harness. Python + npm, MCP, CLI, REST, Web.

362
povvo/
claudikins-kernel

SRE thinking applied to Claude Code, based on Boris Cherny's Q&A. It enforces a strict 4-stage pipeline with gates between each step. You literally cannot skip verification. You cannot ship without approval.

126
fubak/
ultraswarm

Multi-CLI agent swarm orchestrated by Claude Code: external AI CLIs code in isolated worktrees, Claude verifies and merges

83

HAR: open agent harness (CLI + MCP) for coding agents. Isolated worktrees, deterministic verify, software factory workflows for Claude Code, Cursor, and Codex.

88
callstack/
agent-device

Mobile app automation and verification for AI coding agents. CLI, MCP server, and typed Node.js API for iOS, Android, HarmonyOS, TV, web, macOS, and Linux.

4.5k
duthaho/
claudekit

A verification-first engineering toolkit for Claude Code. Built for senior ICs and tech leads who already know how to ship production code — and want a workflow that keeps the discipline tight without getting in the way.

97
yogthos/
chiasmus

Chiasmus is an MCP server that gives language models access to formal verification

211
MrZoyo/
deslop-GPT

Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

125

Professional context and harness engineering for Claude Code and OpenAI Codex. Build production-grade software with spec-driven development, TDD, persistent memory, quality gates, code intelligence, human oversight, and end-to-end verification.

2.1k

The design layer for agentic AI — design context, interface checks, and verification for coding agents.

41
GanyuanRan/
Aegis

Make AI coding agents architecture-aware: baseline-first, evidence-verified, drift-checked, and safe across long tasks.

1.2k

MCP server that orchestrates language servers into agent-native workflows. 65 tools, 30 CI-verified languages.

124
amplifthq/
opentag

Mention any ACP coding agent from Slack, GitHub, GitLab, Linear, or Lark. OpenTag runs Claude Code, Codex, Cursor and more on your own machine, then replies in-thread with verified, evidence-backed results.

1.4k

Compile, verify, and run multi-agent DAGs across Pi, Codex, Claude Code, OpenCode, and Grok—with resume, replay, and incremental recomputation.

71

Tools for AI agents to test, fix and optimise your codebase

46

The multi-agent harness that checks the work: verifies agent runs by artifacts (stop-hook gates, independent judges, append-only event logs) across Claude Code, Codex, Cursor, and 10+ runtimes.

1.3k
alicicek/
tale-mode

A discipline plugin for Claude Code + OpenAI Codex — plan first, verify against the real code, receipts on every decision, independent adversarial review, and a self-armed loop that keeps going until a real check passes. Built by Fable to make any Claude work like a careful senior engineer.

36

AI agents can generate code, but they still struggle to understand what they build. Reticle gives them runtime perception of web & desktop applications.

458

TypeScript multi-agent framework that runs in your own environment: consequential actions wait for approval and every run leaves a verifiable record. Describe the goal, not the graph. 13 built-in providers (Claude, OpenAI, Gemini, DeepSeek and more) plus any OpenAI-compatible endpoint, local models included.

6.9k
aimasteracc/
tree-sitter-analyzer

Cross-language-safe code-intelligence MCP for AI agents -- 13 languages, family-gated call graph (CodeGraph 745 vs TSA 6 cross-language mis-wires, v1.21.0 same-session -- ~124x by count, ~390x by rate; re-verifying vs current develop). Run miswire-audit on your repo. 8 facade tools, TOON output, 100% local. Python.

49
workersio/
skills

Agent skills to find and fix software bugs

166