ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
uber/
ADR
uber/ADRTools
1.5k
UiPath/coder_evalHarnesses
Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
127
aisa-group/
skill-inject
aisa-group/skill-injectHarnesses
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
96
rajudandigam/
Ultimate-TypeScript-Real-World-AI-Projects
250+ real-world TypeScript AI projects: workflows, agents, and multi-agent systems with production-ready architecture, not chatbot demos.
139
NVIDIA/
SkillEvaluator
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
424