adewale/
skill-eval-harness
adewale/skill-eval-harnessHarnesses
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
73
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
An evaluation and evolution tool for Agent Skills.

Comet: agent skill harness for turning ideas into evaluated workflows
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
Open-source AI skills for SEO, AI search visibility, conversion copy, marketing strategy, and business operations. Reusable workflows for Claude Code and Codex, plus code review, app delivery, and creator tools.