Sandbox
19 repos for evals · Any agent · ResearchClear
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

5.1k
agentscope-ai/
OpenJudge
agentscope-ai/OpenJudgeFrameworks & SDKs

OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

826
cxcscmu/
SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

646
keyuchen21/
agentic-engineering-handbook

The definitive OpenAI, Claude, MCP, Harness, Evals, and Production Agent Systems learning roadmap.

187
RichSchefren/
atlas

Open-source local-first cognitive memory. AGM-compliant belief revision (49/49 postulates). When a fact changes, downstream beliefs are automatically re-evaluated, not just flagged.

80

🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

240
superjack2050/
1688-cli

1688 CLI is an AI-agent-friendly command-line tool for 1688 sourcing, product research, supplier evaluation, procurement Inquire, and order management.helping dropshippers and Amazon sellers automate and streamline their sourcing workflow from 1688.

84
joshuaswarren/
remnic

Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.

198

The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.

42k

Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)

71k
MemTensor/
skills-vote

SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution

301
greyhaven-ai/
autocontext

a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task

1.3k
growthxai/outputFrameworks & SDKs

The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already in place.

435
agentvitals/
checkup

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.

95
ggozad/
haiku.rag

Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server

606