Sandbox
6 repos for evaluation · Any agent · DocsClear
alibaba/
skill-up

An evaluation and evolution tool for Agent Skills.

880
RichSchefren/
atlas

Open-source local-first cognitive memory. AGM-compliant belief revision (49/49 postulates). When a fact changes, downstream beliefs are automatically re-evaluated, not just flagged.

80
icip-cas/
PPTAgent
icip-cas/PPTAgentFrameworks & SDKs

An Agentic Framework for Reflective PowerPoint Generation

5k
lc198707/
anti-lie

Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

89
skillberry-ai/
cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

56
ggozad/
haiku.rag

Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server

606