Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.
A native desktop application for developing, testing, and debugging Model Context Protocol servers.

A self-hosted sandbox for red teams to test payloads against modern detection before deployment. MCP integration lets an LLM agent drive analysis end to end.
an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM
The easiest way to run multiple Claude Code sessions, each in its own container, with a dashboard to manage them all. Quick setup with battle-tested sensible defaults and skills.
The evaluation benchmark on MCP servers
Experiment task scheduling made easy.
Run claude code in somewhat safe and isolated yolo mode
A local sandbox for your AI agents
[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
A2A protocol client for your terminal

Delegate tasks to DeepSeek right inside your Claude Code / Codex sessions.
Real-time visualization of Claude Code agent orchestration — see your agents think, branch, and coordinate as they work.
Compounding Context for AI Coding Assistants — MCP graph engine for Claude Code, Cursor, Copilot, Gemini, OpenCode
A fast, keyboard-driven HTTP intercepting proxy and hacking & pentesting toolkit for the terminal.
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.

ColecoVision emulator, debugger and embedded MCP server for macOS, Windows, Linux, BSD and RetroArch.
SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.
Unified execution environment for Python code, shell commands, and programmatic MCP tool calls.
Anti-detect browser automation CLI & Skills for AI agents — Camoufox-powered fingerprint spoofing, no bot-detectable Playwright leaks

PC Engine / TurboGrafx-16 / SuperGrafx / PCE CD-ROM² emulator, debugger, and embedded MCP server for macOS, Windows, Linux, BSD and RetroArch.
Mission control for Claude Code: run many sessions in parallel with multi-account isolation, transcript viewer, cost tracking, and memory dashboards. Windows + macOS (Apple Silicon).
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Build mods for Claude Code: Hook any request, modify any response, /model "with-your-custom-model", intelligent model routing using your logic or ours