SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
Open-World Self-Evolution for LLM Agents — agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
Fluent argument validation for fluent software development.
Open-source, self-hosted Claude Code - a terminal AI assistant and the Python framework behind it. Tool-calling, sandboxed execution, multi-agent teams, skills, checkpoints, unlimited context - on Pydantic AI, any model.
Poirot is a deep research agent kernel built for those who care about how agents are architected.
Computer-Use SDK for E2E QA Testing

Deterministic safety solutions for probabilistic AI agents
Prior direction, kept rather than deleted. Signed, offline-verifiable receipts for AI agent actions, and a reference implementation of the OWASP Agentic Skills Top 10 AST09 receipt pattern. Nobulex is now the independent reliability registry for agent tools: github.com/arian-gogani/nobulex-registry
OpenBrowser is a framework for intelligent browser automation. It combines direct CDP communication with a CodeAgent architecture, where the LLM writes Python code executed in a persistent namespace, to navigate, interact with, and extract information from web pages autonomously.
Agent-SDK without CLI dependencies, as an alternative to claude-agent-sdk, completely open source

Real time communication for agents. Wake on message, channels, DMs and actions. Useful for orchestrating agents.
A minimal yet powerful framework for creating AI agents with full control over tools, providers, and execution flow.