Sandbox
8 repos for evaluation · Codex · DocsClear

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

2.6k
rpamis/cometHarnesses

Comet: agent skill harness for turning ideas into evaluated workflows

3k
alibaba/
skill-up

An evaluation and evolution tool for Agent Skills.

880
Epsilon617/
Codex-Academic-Skills

A curated list of research-oriented skills usable in OpenAI Codex, covering writing, literature review, evaluation, and research workflows.

180
oxbshw/
LLM-Agents-Ecosystem-Handbook

One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.

546
comet-ml/
opik-mcp

Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude Code, Cursor, or VS Code.

217

239 evaluated academic Claude/agent skills across 17 research domains (bioinformatics, data science, clinical, social-science methods, Turkish academia & more). Executable eval per skill, deterministic citation verifier, research→write→review→publish pipeline, and a skill-finder front door. Claude Code, Cursor, Codex, Gemini CLI & Copilot.

66
ggozad/
haiku.rag

Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server

606