Verifiable and free cloud compute for AI agents. webMCP + MCP native. Check out our sandboxed Beta + research in the README
The evaluation benchmark on MCP servers
Experiment task scheduling made easy.
A native desktop application for developing, testing, and debugging Model Context Protocol servers.
Run claude code in somewhat safe and isolated yolo mode
A local sandbox for your AI agents
[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
A2A protocol client for your terminal

Delegate tasks to DeepSeek right inside your Claude Code / Codex sessions.
Real-time visualization of Claude Code agent orchestration — see your agents think, branch, and coordinate as they work.
Compounding Context for AI Coding Assistants — MCP graph engine for Claude Code, Cursor, Copilot, Gemini, OpenCode
A fast, keyboard-driven HTTP intercepting proxy and hacking & pentesting toolkit for the terminal.
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.

ColecoVision emulator, debugger and embedded MCP server for macOS, Windows, Linux, BSD and RetroArch.
SRA-Bench and SR-Agents: a benchmark and toolkit for skill-retrieval-augmented LLM agents.
Unified execution environment for Python code, shell commands, and programmatic MCP tool calls.
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

A self-hosted sandbox for red teams to test payloads against modern detection before deployment. MCP integration lets an LLM agent drive analysis end to end.
Anti-detect browser automation CLI & Skills for AI agents — Camoufox-powered fingerprint spoofing, no bot-detectable Playwright leaks

PC Engine / TurboGrafx-16 / SuperGrafx / PCE CD-ROM² emulator, debugger, and embedded MCP server for macOS, Windows, Linux, BSD and RetroArch.
Mission control for Claude Code: run many sessions in parallel with multi-account isolation, transcript viewer, cost tracking, and memory dashboards. Windows + macOS (Apple Silicon).
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Build mods for Claude Code: Hook any request, modify any response, /model "with-your-custom-model", intelligent model routing using your logic or ours
Claude Code session log viewer for JSONL files in ~/.claude/projects. Browse conversations, tool calls, tokens, and live tail sessions on desktop, web, and TUI.