Open-World Self-Evolution for LLM Agents — agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)
Multi-agent research automation framework for LLM agents, with adversarial lab meetings, paper-review rounds, auditable Markdown workflows, an autonomous runtime watchdog, and a pixel-art web dashboard.
LLM agents as your hyperparameter optimizer.
An MCP server that lets LLM agents play Civilization VI.
The python library for research and development in NLP, multimodal LLMs, Agents, ML, Knowledge Graphs, and more.

MCP server for the Live Tennis API — give Claude, Cursor and other LLM agents real-time tennis scores, odds and model win-probability

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Durable, searchable memory of your past agent sessions.
Turn project work into reusable knowledge — an AI-agent skill for Claude Code & Codex
From thought to skill. From signal to structure.
Causal memory layer for AI agents — MCP server that records decision→outcome relationships. Survives compaction.
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
An Automated AI Agent Tool for Plotting Your Data in Any Paper's Figure Style.
Compiles AI agent traces and truns them into reusable context.
MCP server for spatial transcriptomics analysis through natural language interfaces.
A collection of AI agent skills for Clawdbot, Claude Code, Codex
Run a task with AI as a flow of steps you keep, reuse, and refine, not a one-off chat.
Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue / PR / wiki. Runs on any agentskills.io runtime — Claude Code, Codex, OpenClaw, Hermes.
Research-grade investment decision engine for AI agents: isolated multi-agent committee, auditable verdicts, backtests with lookahead protection, published negative results
A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring, Leader-Worker architecture, constant-size memory.
WikiSkill (arXiv:2608.27454) for Hermes Agent — self-evolving agent skills via a persistent knowledge wiki. Faithful Algorithm 1 implementation with real agent runs, isolated skill gating, and a documented live run log.
Distilly — Distill how they think into reusable Skills for any Agent or Bot. Formerly Colleague Skill(原同事 Skill).