Sandbox
14 repos for obs · Claude Code · Code reviewClear
KryptosAI/
mcp-observatory

CI-native security testing for MCP servers. Attack simulation, schema drift detection, and health scoring before agents depend on them.

146
obra/
superpowers

An agentic skills framework & software development methodology that works.

285k
1 add

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

2.8k

See your agent think. Zero-config observability & governance for 30 AI agent runtimes: Claude Code, OpenAI Codex, Hermes, OpenClaw & 26 more. Live token costs, sessions, tool calls, crons.

411

eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

12k
leopiney/
linus-torvalds-skills

A single CLAUDE.md file to improve Claude Code behavior, derived from Linus Torvalds' observations on coding pitfalls.

283
forloopcodes/
contextplus

Semantic Intelligence for Large-Scale Engineering. Context+ is an MCP server designed for developers who demand 99% accuracy. By combining RAG, Tree-sitter AST, Spectral Clustering, and Obsidian-style linking, Context+ turns a massive codebase into a searchable, hierarchical feature graph.

2k
georgeguimaraes/
elixir-agent-tools

Elixir development skills with optional Mix checks and Expert language server integration

169
backnotprop/
plannotator

Annotate and review coding agent plans and code diffs visually, share with your team, send feedback to agents with one click.

8.6k
tigerless-labs/
cost-xray

See what Claude Code and Codex actually send to the API — and what each part costs.

1.5k

Tools for AI agents to test, fix and optimise your codebase

46

Find and repair substance defects in AI-assisted prose, code, docs, and agent output. Reports defects, never authorship. Structural tests over model judgement, because LLM judges agree with human slop labels at chance.

48

Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic contracts, not model opinion. DeepWatch is the agent workspace built on DeepSeek Harness. Python + npm, MCP, CLI, REST, Web.

362

TypeScript multi-agent framework that runs in your own environment: consequential actions wait for approval and every run leaves a verifiable record. Describe the goal, not the graph. 13 built-in providers (Claude, OpenAI, Gemini, DeepSeek and more) plus any OpenAI-compatible endpoint, local models included.

6.9k