Sandbox
24 repos for trace · Codex · CodingClear
saidake/
code-trace-tree-vscode

Code Trace Tree is a VS Code/JetBrains plugin that traces code in a tree structure with AI support. Build and display code workflows as trace points. Add traces from the editor menu or the Agent Skill; double-click to jump to the source.

37
nikolai-vysotskyi/
trace-mcp

Framework-aware code intelligence MCP server for Claude Code and Codex — 70.5% fewer input tokens to review a pull request, median over 60 merged PRs in repos we don't own, comprehension at parity. 81 languages, 87 frameworks. Your code and index never leave the machine; an anonymous usage ping is on by default and opt-out.

176
nablo-io/
lerim

Compiles AI agent traces and truns them into reusable context.

97
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73
avivsinai/
langfuse-mcp

A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

105

Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.

3.2k
WecoAI/
awesome-autoresearch

Curated list of AutoResearch use cases with optimization traces and open source implementations

1k
comet-ml/
opik-mcp

Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude Code, Cursor, or VS Code.

217

Evidence-first reading for AI agents — turn articles, books and PDFs into traceable claims, evidence, source locations and knowledge maps.

49
postmelee/
hyper-waterfall

A human-governed AI coding workflow that distills ephemeral session context into persistent project memory—making work traceable, reviewable, and resumable.

80
Fergana-Labs/
stash

Automatically create new skills based on past agent traces

332
clarity-digital-development/
tworkflow

A practical, no-hype workflow for AI coding agents: context, plan, implement, review, QA, ship, retro. Templates, two Claude Code skills, and a 40% context rule - every claim traced to official docs.

47

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

2.8k
alibaba/
loongsuite-pilot

Local-first telemetry collector for AI coding agents — unified OpenTelemetry events for Claude Code, Codex, Cursor and more. Token usage, cost, traces and security audit, exported anywhere.

177
morluto/flameoxConnectors

Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.

121

Tools for AI agents to test, fix and optimise your codebase

46
FailproofAI/
failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement. 40 built-in policies, a local dashboard, no account required with a generous free cloud plan

2.7k
AIScientists-Dev/
Flowtrace

Run a task with AI as a flow of steps you keep, reuse, and refine, not a one-off chat.

486
jhonsfran/
unprice
jhonsfran/unpriceFrameworks & SDKs

Open-source customer money path for usage-based SaaS — authorize customer spend before paid work runs.

36
kayba-ai/
Kyoko

🔨 Kyoko is the all-in-one, fully local tool for debugging and improving your AI agents.

97
MCPJam/
inspector

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

2.2k
tigerless-labs/
design-harness

Feed your agent papers and half-formed ideas — it links them into a system design you can defend. Markdown keeps the record; a visual canvas makes it readable. An Agent Skill for Claude Code & any SKILL.md-compatible agent.

217

The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.

1.3k

Structural memory for AI coding agents. Bi-temporal graph, MCP-native, zero LLM calls. Cursor · Claude Code · Codex · DeepSeek Harness · Hermes · VS Code · Windsurf.

469