Sandbox
44 repos for evidence · Any agentClear

Evidence-first reading for AI agents — turn articles, books and PDFs into traceable claims, evidence, source locations and knowledge maps.

49
lllllllama/
RigorPilot-Skills

README-first research reproduction skills with bounded execution, auditable evidence, and byte-preserving README annotations.

486
yaojingang/
GEOHub

GEOHub: open, evidence-bounded GEO and SEO agent skills for AI Search, with research-grounded discovery, diagnosis, content, measurement, and one-line SEO planning.

156

Opinionated Oxlint rules for rejecting low-evidence TypeScript and JavaScript patterns

4.3k
millwright-labs/
minto-pyramid-skill

Agent Skill: make Claude write in Barbara Minto's Pyramid Principle - answer first, grouped reasons, evidence under each.

70
Zhen-Bo/
smell-check

Agent Skill for code and test smell audits. Evidence-ranked findings from Refactoring, Clean Code, and the test-smell literature. Formerly pragmatic-code-review.

237

Tamper-evident integrity monitor for the MCP config & server files your local AI agents load.

46

Remote approvals, policy checks, and execution evidence for unattended AI agents.

301
AaravKashyap12/
advise-project-approach

A portable project-planning skill for Codex, Claude Code, pi, Hermes, and Agent Skills-compatible harnesses. Evidence before build advice.

300
DY-2026/
GameDesignOS

Local-first game design OS for AI agents: turn sessions into evidence, experiments, reviewable decisions, and durable project memory—Human Gates and rollback.

385

Cut context bloat in your AI-agent stack: find and safely prune unused skills, MCP servers and subagents from real transcript evidence

56
Gentleman-Programming/
gentle-pi

Turn Pi into el Gentleman: a senior-architect development harness with SDD/OpenSpec, subagents, strict TDD evidence, review guardrails, and skill discovery.

715
amplifthq/
opentag

Mention any ACP coding agent from Slack, GitHub, GitLab, Linear, or Lark. OpenTag runs Claude Code, Codex, Cursor and more on your own machine, then replies in-thread with verified, evidence-backed results.

1.4k
mcp-security-standard/
mcp-server-security-standard

MCP Server Security Standard (MSSS): an open, testable security control standard for certifying MCP servers, with levels, evidence requirements, and reporting schemas.

74
AmazingAng/
old-coder

An old coder's strategy for the agent era: don't read the code — make it run the gauntlet. Evidence-first development skill for coding agents, inspired by Uncle Bob.

720
k-telux/
OpticalModeler

Evidence-gated Agent Skill for reconstructing 2D photonics schematics as physically auditable Blender optical tables with CAD, beam-path, mechanics, and render proof.

216
Tranz007/
ux-skills

Agent Skills for working UX designers. Challenge ideas, find blind spots, use the real design system, preserve evidence and decisions, and carry UX intent into engineering.

41
nagisanzenin/
engram

Evidence-based learning engine — first-principles curricula, free-recall verification with receipts, FSRS-scheduled memory, and explorable artifacts. Learn anything; keep it.

1.4k
Cranot/
roam-code

Local codebase intelligence CLI + MCP server for AI coding agents: SQLite code graph, 28 languages, 287 commands, 246 MCP tools, change-safety gates, audit evidence, zero API keys.

517
lna-lab/
distill-kura

蒸留蔵 — distilled long-term memory for agents: recall by meaning, writing gated by evidence, one kura per agent mode. Ships as a DeepSeek Harness plugin and an MCP server.

48
AIPentest/
CyberStrikeAI

The system of action for AI-native cybersecurity—where intent becomes governed execution, evidence becomes operational memory, and every operation improves the next.

6.5k
DenisSergeevitch/
game-sensitivity-coach

An agent skill for evidence-based mouse sensitivity tuning, cross-game conversion, and gameplay review. Works with Codex, Claude Code, and Agent Skills hosts.

34
clawplays/
ospec

Spec-driven, agentic workflow framework for AI coding agents. Turn a request into a verifiable goal loop — plan, act, verify — with durable specs and evidence in your repo. Works with Claude Code, Codex, Gemini, OpenCode, and plain CLI.

485