Sandbox
132 repos for testing · ResearchClear
pinecone-io/
cultivar

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

37
RamKansal/
pentestMCP

pentestMCP: AI-Powered Penetration Testing via MCP, an MCP designed for penetration testers.

96
Epistates/
turbomcpstudio

A native desktop application for developing, testing, and debugging Model Context Protocol servers.

37

Touhou-inspired Agent Skills: distinct, testable, composable problem-solving workflows.

21

Security testing toolkit for AI Agent: curated SecLists wordlists, injection payloads, and expert agents for authorized pentesting, CTFs, and bug bounties

380

A self-hosted sandbox for red teams to test payloads against modern detection before deployment. MCP integration lets an LLM agent drive analysis end to end.

1.5k

BitDive Model Context Protocol (MCP) server. The Autonomous Quality Loop for AI agents. Provides real runtime context, before/after trace comparison, and integration testing workflows.

75

Self-hosted AI agent harness in a single Go binary — writes, sandbox-tests and repairs its own tools, and lets Claude Code, Codex and any MCP client build and share them.

508
aaif-goose/
goose

an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM

54k

MCP server giving AI agents full access to Julia's runtime via a live Gate — code execution, introspection, debugging, testing, and semantic search

41
sgharlow/
claude-code-recipes

100 field-tested Claude Code recipes for knowledge workers — prompts, steps, and 6 installable graded skills.

386
deanpeters/
Product-Manager-Skills

Product Management skills framework built on battle-tested methods for Claude Code, Cowork, Codex, and AI agents.

6.9k
aisa-group/
skill-inject

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

96
SNL-UCSB/
paper-writing-skill

A Claude Code skill that encodes battle-tested editorial principles, section-specific rhetorical moves, and a structured writing pipeline for research papers. Brainstorm → Draft 0 → Evaluate → Write → Compress.

191
moonlight-lupin/
agent-skills

A collection of AI agent skills for Hermes Agent — research, creative, productivity, devops, and more. Each skill is self-contained and tested.

62
cypress-io/
ai-toolkit

Fast, flexible, and open tooling for building intelligent workflows with Cypress.

40
ykdojo/
safeclaw

The easiest way to run multiple Claude Code sessions, each in its own container, with a dashboard to manage them all. Quick setup with battle-tested sensible defaults and skills.

183
KeWang0622/
agent-zero-to-hero

Build a Claude-Code-shaped agent harness from scratch. 7-week course, 20 chapters, ~5,000 lines of Python, 42 tests, 3 LLM providers, no frameworks.

49
kgraph57/
paper-writer-skill

Claude Code skill for academic manuscript writing: IMRAD workflows, literature matrices, tables/figures, and tested utilities.

55
XuanRanL/
loamwright-SEO-Skill

Production-grade SEO + GEO content factory for Claude Code — research, write, fact-check, optimize, publish to WordPress, monitor. Battle-tested by Loamwright 沃匠 SEO agency.

49
DMontgomery40/
pentest-mcp

NOT for educational purposes: An MCP server for professional penetration testers including STDIO/HTTP/SSE support, nmap, go/dirbuster, nikto, JtR, hashcat, wordlist building, and more.

145
taielab/
awesome-hacking-lists

A curated collection of top-tier penetration testing tools and productivity utilities across multiple domains. Join us to explore, contribute, and enhance your hacking toolkit!

1.4k
AIPentest/
CyberStrikeAI

The system of action for AI-native cybersecurity—where intent becomes governed execution, evidence becomes operational memory, and every operation improves the next.

6.5k