Sandbox
27 repos for agent-testing · CodingClear
idavidov13/
agentic-playwright

Production-grade Playwright + TypeScript Scaffold for Agentic Testing. Harness for all major AI coding agents baked in.

163

Runtime intelligence system that makes MCP servers debuggable, testable, and safe to run in production.

48
Epistates/
turbomcpstudio

A native desktop application for developing, testing, and debugging Model Context Protocol servers.

37
browser-use/
vibetest-use

Vibetest MCP - automated QA testing using Browser-Use agents

830
adewale/
skill-eval-harness

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

73
agentscope-ai/
OpenJudge
agentscope-ai/OpenJudgeFrameworks & SDKs

OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

826
lmwilki/
civ6-mcp

An MCP server that lets LLM agents play Civilization VI.

172
budtmo/
docker-android

Android in docker solution with noVNC supported, video recording and mcp server

16k

Token-efficient, local-first CLI tools for coding agents - compact Maven, npm/Node, and Go test output plus reusable development helpers.

184
AvdLee/
Swift-Testing-Agent-Skill

An agent skill focused entirely on Swift Testing, helping you write better tests, migrate from XCTest, improve test architecture, and adopt modern Swift testing patterns with confidence.

447
ShaftHQ/SHAFT_ENGINEFrameworks & SDKs

Java test automation framework for web, mobile, API, CLI, database, and desktop E2E testing with a fluent API and built-in reporting.

409

BitDive Model Context Protocol (MCP) server. The Autonomous Quality Loop for AI agents. Provides real runtime context, before/after trace comparison, and integration testing workflows.

75

Self-hosted AI agent harness in a single Go binary — writes, sandbox-tests and repairs its own tools, and lets Claude Code, Codex and any MCP client build and share them.

508

AI agents can generate code, but they still struggle to understand what they build. Reticle gives them runtime perception of web & desktop applications.

458
OpenEvident/
vindicate

A local-first Playwright test automation toolkit for AI coding agents (Cursor, Claude Code, Copilot), with MCP tools for codegen, browser control, and recordings.

46
agentvitals/
checkup

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.

95