Sandbox
6 repos for eval · Harnesses · Any agentClear
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
skillberry-ai/
cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

56

The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.

42k
greyhaven-ai/
autocontext

a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task

1.3k