Sandbox
23 repos for reliabilityClear
cloudwego/
abcoder
cloudwego/abcoderFrameworks & SDKs

deep, reliable and confidential coding-context

390
FailproofAI/
failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement. 40 built-in policies, a local dashboard, no account required with a generous free cloud plan

2.7k
Ancienttwo/
repo-harness

File-backed workflow harness for reliable Claude Code and Codex sessions.

431
AgiFlow/
aicode-toolkit

Toolkit for Coding Agents to work reliably with repo of any size.

162
microsoft/
flint-chart

🪄 Flint is a visualization language that lets AI agents reliably create expressive, good-looking charts from simple, human-editable chart specs.

4.2k
twelvedata/
mcp
twelvedata/mcpConnectors

Twelve Data MCP (Model Context Protocol) Server provides seamless, real-time access to financial market data via WebSocket, enabling reliable streaming of price quotes, market metrics, and events directly into your applications.

76
CelestoAI/agentorFrameworks & SDKs

Open source version of Claude Managed Agents. Fastest way to build and deploy reliable AI agents, MCP tools and agent-to-agent.

205
mondaycom/
mcp
mondaycom/mcpConnectors

Enable AI agents to work reliably - giving them secure access to structured data, tools to take action, and the context needed to make smart decisions.

424
Klavis-AI/
klavis

Klavis AI: MCP integration platforms that let AI agents use tools reliably at any scale

5.8k
smithersai/smithersFrameworks & SDKs

Smithers is an agentic workflow framework for defining workflows in simple TypeScript configuration files and executing them quickly, durably, and reliably

410

Synapse is a local-first personal memory operating system designed for AI agents. It provides durable, typed, and highly reliable long-term memory using only local SQLite — giving agents the ability to remember, update, and reason over facts across sessions with full user control.

63
vaquarkhan/
data-engineering-agent-skills

Production-grade Agent Skills for data engineering AI agents: 73 workflows, platform presets, safe backfill/replay, Kafka & Spark reliability, MCP observability, and VS Code/JetBrains installers.

45
arian-gogani/
nobulex
arian-gogani/nobulexFrameworks & SDKs

Prior direction, kept rather than deleted. Signed, offline-verifiable receipts for AI agent actions, and a reference implementation of the OWASP Agentic Skills Top 10 AST09 receipt pattern. Nobulex is now the independent reliability registry for agent tools: github.com/arian-gogani/nobulex-registry

39
LetsFG/LetsFGConnectors

Agent-native flight & hotel search and booking — MCP server, CLI, and Python/JS SDKs. Hundreds of airlines plus the major booking sites, with per-flight reliability history. Free-cancellation hotel rates: hold the room with a small upfront charge, then pay the balance later by link, up to the hotel's own deadline.

2k
AMAP-ML/
LongHorizon-Harness

The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.

1.5k
digipulse-engineering/
GAAI-framework

Turns AI coding tools into reliable software delivery systems. Drop a .gaai/ folder into any project — Discovery defines what to build, Delivery executes autonomously until criteria pass. Works with Claude Code, Codex CLI, Gemini CLI, Cursor, and more. No SDK. No package. Markdown + YAML + bash.

160

Open-source governed, local-first memory control plane for AI agents and teams. arXiv:2608.08253

224
growthxai/outputFrameworks & SDKs

The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already in place.

435

Keep Claude Code context clean. Open-source toolkit: drift detection, re-read dedup, integrity scoring, AST-aware reads, 15 MCP tools. 62.6% measured savings, reproducible.

46