Skills for controlling Android, iOS and cloud phones
File-backed workflow harness for reliable Claude Code and Codex sessions.
A workflow framework for statistical package development
The open-source trading harness for Claude Code / Codex agents.

Multi-Agent Harness for Production AI
Threat hunting command system for agentic IDEs
Containment for AI agents - user isolation, sandboxed execution, network controls, backup/rollback. TLA+ verified.
From thought to skill. From signal to structure.
Local-first workspace for Claude Code, Codex CLI, and Gemini CLI with sessions, analytics, workflows, and tools
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols.
Run a task with AI as a flow of steps you keep, reuse, and refine, not a one-off chat.
Governance standard and reference toolset for LLM-maintained knowledge corpora
Universal, model-agnostic operating harness for AI agents (Claude, Codex, Gemini, …) — a lean core + work-type profiles assembled by one setup script.
HAR: open agent harness (CLI + MCP) for coding agents. Isolated worktrees, deterministic verify, software factory workflows for Claude Code, Cursor, and Codex.
Local-first, self-hosted AI agent runtime and MCP bridge with sandboxed sessions, memory, credentials, audit/replay, and a local Console.
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.