File-backed workflow harness for reliable Claude Code and Codex sessions.
My personal directory of skills.
Skills for controlling Android, iOS and cloud phones
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

The harness layer for Claude Code — a reference implementation of harness engineering with hook-enforced dual review, state-machine gates that survive context compaction, and fail-closed safety where it counts. Quality gates that AI can't skip.
A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.
Protocol-layer harness for DeepSeek: Python witness stack — posterior verification that keeps the protocol honest. dsh doctor --node probes included.
A kit for building with AI agents and also the engineering patterns around it.

A provider-agnostic scaffolding kit for running structured multi-agent workflows in your codebase.
Public results and task definitions for FrontierHarness Eval
Turn any repo into an agent-ready workspace for Claude Code, Codex, Cursor, and other coding agents.
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.
Local-first, self-hosted AI agent runtime and MCP bridge with sandboxed sessions, memory, credentials, audit/replay, and a local Console.
Self-evolving SDLC enforcement for AI coding agents — hooks, skills, and one-command setup for Claude Code. Plan before coding, test before shipping, escalate when uncertain. Measures itself getting better over time.
Desktop workbench for Claude Code — live subagent topology, token cost breakdown with drill-down, hook sandbox, config audit. Reads your session files locally. Electron, MIT.
Enterprise-grade AI Agent Skills for software development, DevOps, SRE, security, and product teams. Compatible with Claude Code, Cursor, Windsurf, Gemini CLI, GitHub Copilot, and 30+ AI coding agents.
HAR: open agent harness (CLI + MCP) for coding agents. Isolated worktrees, deterministic verify, software factory workflows for Claude Code, Cursor, and Codex.
Portable AI agent orchestration with mechanical protocol enforcement. 186 agents, zero runtime dependencies.
Open-source, desktop client/UI build to harness Claude Code, Codex and any other Agent accepting Agent Client Protocol. Run multiple AI coding agents side by side with rich tool visualization, MCP integrations, built-in terminal, git, browser and just about anything else you may need.
A lightweight agent harness you bolt onto your app so an LLM can operate it — safely, and cheaply.

Multi-Agent Harness for Production AI
Codebase harness + loop engineer
The Agent Harness for AI-Human Collaboration, inspired by the AI-DLC (AI-Driven Development Lifecycle)