YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
Public results and task definitions for FrontierHarness Eval

Comet: agent skill harness for turning ideas into evaluated workflows
Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.