Sandbox
80 repos for skill · Harnesses · CodexClear
tripleyak/
SkillForge

A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

890
aisa-group/
skill-inject

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

96
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

2.6k
FrancyJGLisboa/
agent-skills-platform

Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

2.4k

Co-creation infrastructure for humans and code agents — visual environment, skills, continuous learning, and distribution.

160

Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic contracts, not model opinion. DeepWatch is the agent workspace built on DeepSeek Harness. Python + npm, MCP, CLI, REST, Web.

362
appautomaton/
latex-arxiv-SKILL

A highly customizable agentic harness for arXiv-ready ML/AI review papers (and beyond). It drives agentic AI like Codex CLI and Claude Code through a gated LaTeX workflow with verified BibTeX citations.

427
ashutoshsinghpr7/
wikiskill

WikiSkill (arXiv:2608.27454) for Hermes Agent — self-evolving agent skills via a persistent knowledge wiki. Faithful Algorithm 1 implementation with real agent runs, isolated skill gating, and a documented live run log.

144

Plinth is an AI-native engineering toolkit for modern Java enterprise SDLC, built around reusable Commands, Agents, Skills, and MCP Servers.

438
rpamis/cometHarnesses

Comet: agent skill harness for turning ideas into evaluated workflows

3k
CodelyTV/
agent-harness

Our agent harness: Skills, plugins, hooks, and utilities to improve the quality of your agent.

261
JasonxzWen/
harness-hub

Repository-first deterministic migration and atomic Skill source for Claude Code and Codex.

72
Utopai-Research/
pai-pro

Local AI filmmaking studio — skills, canvas, timeline — driven from your coding agent.

340
yzhao062/
anywhere-agents

One config to rule all your AI agents: portable (every project, every session), effective (curated writing, routing, skills), and safer (destructive-command guard).

244
drvoss/
everything-copilot-cli

The definitive guide & configuration system for GitHub Copilot CLI — agents, skills, rules, multi-AI orchestration, and more

46
AnastasiyaW/
codex-claude-code-config

Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development

150
huytieu/
COG-second-brain

Self-evolving second brain with 33 AI skills, 10 agents, and people CRM. Closed-loop harness: a V-model verification lifecycle where the worker never grades its own homework. Plus paired anti-slop design skills for marketing and product UI. Works with Claude Code, Cursor, Kiro, Gemini CLI, Codex.

1.2k
notque/
vexjoy-agent

VexJoy AI Agent with Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop.

419

Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.

127
molefrog/moiHarnesses

💾 MOI: agent-agnostic generative UI workspace. Give your agent a skill and let it build software around itself.

158
MiaoDX/
intuitive-flow

Shared config, skills, and update scripts for Claude Code, Codex, and Gemini CLI

48
getnao/
sylph
getnao/sylphHarnesses

The open-source company brain. Run your entire company with AI agents, skills, and a self-improving context.

196