
Comet: agent skill harness for turning ideas into evaluated workflows

Comet: agent skill harness for turning ideas into evaluated workflows
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.