Sandbox
15 repos for ai-safetyClear
aisa-group/
skill-inject

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

96
assafkip/
research-mode

Anti-hallucination research mode for Claude Code. Toggle on/off to enforce citation requirements and source grounding.

151
lc198707/
anti-lie

Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

89
JKHeadley/instarFrameworks & SDKs

Persistent Claude Code agents with scheduling, sessions, memory, and Telegram.

79
FlyFission/
nuclear-grade-context-engineering

AI agents now operate with authority. Authority without discipline is how complex systems fail. Nuclear’s control loop, ported to AI-assisted software engineering.

33
Govcraft/
rust-docs-mcp-server

🦀 Prevents outdated Rust code suggestions from AI assistants. This MCP server fetches current crate docs, uses embeddings/LLMs, and provides accurate context via a tool call.

295
yzhao062/
anywhere-agents

One config to rule all your AI agents: portable (every project, every session), effective (curated writing, routing, skills), and safer (destructive-command guard).

244
eli-labz/
Cognitive-Core-Skills

A universal, industry-neutral taxonomy of cognitive core skills (perception, memory, reasoning, planning, action, verification, learning, governance) for LLMs, SLMs, AI agents, and world models — with schemas, 159 skill cards, benchmarks, and CI.

165
CodeAlive-AI/
ai-driven-development

Practices, protocols, and skills for AI-driven software development. Skills and safety hooks for Claude Code, Codex, OpenCode, Cursor, Antigravity, and any agent supporting the Agent Skills standard.

132
elliot35/
deterministic-agent-control-protocol

Governance gateway for AI agents — bounded, auditable, session-aware control with MCP proxy, shell proxy & HTTP API. Works with Cursor, Claude Code, Codex, and any MCP-compatible agent.

88
Agile-V/
agile_v_skills

Agent Skills for traceable requirements, independent verification, human approval gates, and auditable AI-assisted engineering. Supports Claude Code, Cursor, VS Code, and GitHub Copilot; ISO 9001/27001 aligned (design phase), GxP-aware.

53

Self-hosted runtime control plane for AI agents. Observe or HITL approve or Block rogue tool calls before it executes: secret leaks, prompt injection, supply chain etc in a local dashboard. Agent agnostic (Claude, codex, langchain etc.)

342

A pre-execution guard for AI coding agents. It blocks destructive Git and file system commands, plus common attempts to access sensitive files, before a tool call runs. Supports Amp Code, Antigravity CLI, Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot CLI, Grok Build, Hermes Agent, Kimi Code, OpenClaw, OpenCode, and Pi.

1.5k