Agent Skills Evaluation Framework
See how you really use AI — X-ray your AI coding sessions locally

Local, read-only audit for stale AGENTS.md, CLAUDE.md, and generic SKILL.md instructions.

A self-hosted sandbox for red teams to test payloads against modern detection before deployment. MCP integration lets an LLM agent drive analysis end to end.
Secure multi-platform AI skill installer — scan before you install. 49 agents, 12 plugins, 41 expert skills.
More than a skill manager — manage skills, MCP servers, plugins, hooks, CLIs, configs, memory & rules across every AI coding agent. 🌟 Star if you like it!
Cover your Mac screen with a hotkey while AI agents keep running. Dog/cat mascot, native Settings, Touch ID unlock.
Offline security scanner for AI-agent repos, skills, plugins, and MCP servers.
Skill engineering methodology and publishing pipeline for AI agent skills. Validates structure, scans for security, audits entire projects, and publishes to GitHub. Skills are code — engineer them like it.
Local-only Go static analysis engine with a built-in MCP server. Gives AI coding agents deterministic structural awareness: call graphs, impact analysis, symbol search, and more.
Local-first telemetry collector for AI coding agents — unified OpenTelemetry events for Claude Code, Codex, Cursor and more. Token usage, cost, traces and security audit, exported anywhere.
Unified CLI for running AI coding agents in isolated containers. Includes built-in local metrics collection, HTTP traffic tracking, and an analytics dashboard to track agent actions.

The security-first skill manager for AI agents — every install runs a security scan. Manage skills & MCP servers across 87 agents. Zero-dependency CLI.
CI-native security testing for MCP servers. Attack simulation, schema drift detection, and health scoring before agents depend on them.
Research-first architecture engine for Python, TS, Go and Rust. Mines GitHub Issues for real production failures before scaffolding, then audits drift, cycles and fragility in code you already have. Zero import cycles, zero critical anti-patterns - measured against itself.