Sandbox
29 repos for observability · Claude Code · TestingClear
KryptosAI/
mcp-observatory

CI-native security testing for MCP servers. Attack simulation, schema drift detection, and health scoring before agents depend on them.

146

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

2.8k
FailproofAI/
failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement. 40 built-in policies, a local dashboard, no account required with a generous free cloud plan

2.7k

See your agent think. Zero-config observability & governance for 30 AI agent runtimes: Claude Code, OpenAI Codex, Hermes, OpenClaw & 26 more. Live token costs, sessions, tool calls, crons.

411

Official Monte Carlo toolkit for AI coding agents. Skills and plugins that bring data and agent observability — monitoring, triaging, troubleshooting, health checks — into Claude Code, Cursor, and more.

91

eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

12k
uber/
ADR

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

1.5k
RyjoxTechnologies/
Octopoda-OS

The open-source memory and observability layer for AI agents — persistent memory, loop detection, hash-chained audit trails, and a live dashboard, automatic on pip install.

490
avivsinai/
langfuse-mcp

A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

105
vaquarkhan/
data-engineering-agent-skills

Production-grade Agent Skills for data engineering AI agents: 73 workflows, platform presets, safe backfill/replay, Kafka & Spark reliability, MCP observability, and VS Code/JetBrains installers.

45

What Claude Code is doing between your prompt and its answer, drawn live from the logs it already writes: every model call, every tool, each subagent on its own context window and its own model, and what the session produced. Local and read-only — your session content never leaves the machine.

45
lookfree/
cc-harness

Desktop workbench for Claude Code — live subagent topology, token cost breakdown with drill-down, hook sandbox, config audit. Reads your session files locally. Electron, MIT.

48

Self-hosted runtime control plane for AI agents. Observe or HITL approve or Block rogue tool calls before it executes: secret leaks, prompt injection, supply chain etc in a local dashboard. Agent agnostic (Claude, codex, langchain etc.)

342
simota/
agent-skills

90 specialist AI agents + 3 project-local extensions for Claude Code / Codex CLI / Antigravity CLI (agy). Anthropic Agent Skills spec-aligned, hub-spoke orchestration via Nexus with 49 Recipes and 11 Skill Packs. Covers development, security, design, testing, FinOps, compliance, observability, and AI/ML.

76
everr-labs/
everr

CLI and telemetry system for querying local, CI, and production runtime data.

45
jfrog/
boost

Save tokens. Maximize context, Safely

473
the-void-ia/
void-box
the-void-ia/void-boxFrameworks & SDKs

Composable agent runtime with enforced isolation boundaries

88
dash0hq/
agent-skills

OpenTelemetry skills and reference documentation for AI coding assistants - instrumentation patterns, telemetry quality guides, and Dash0 integration

89

Tools for AI agents to test, fix and optimise your codebase

46
NVIDIA/
NeMo-Relay
NVIDIA/NeMo-RelayFrameworks & SDKs

Multi-language agent runtime and library for execution scope management, lifecycle events, and middleware on tool and LLM calls.

169

BitDive Model Context Protocol (MCP) server. The Autonomous Quality Loop for AI agents. Provides real runtime context, before/after trace comparison, and integration testing workflows.

75

Self-hosted control plane for AI agents: dispatch tasks, review runs, track spend, and operate OpenClaw, Claude Code, Codex, and other runtimes.

6.2k