Sandbox
@oxbshw/LLM-Agents-Ecosystem-Handbook

Handbook for LLM agent workflows and templates

This repo teaches the full stack behind modern LLM agents: identity, memory, skills, MCP, safety, observability, evals, providers, and deployment. It pairs the concepts with blueprints, checklists, templates, and many example agent skeletons so you can adapt the patterns to your own projects.

546 stars87 forksPythonUpdated 2mo ago
Who it's for

Builders who want reusable guidance for agent design, prompts, workflows, and production rollout.

What it delivers

You can build and ship agent systems with clearer structure, safer tool use, and reusable starting points.

What it does

Agent OS layers

Defines identity, memory, skills, MCP, safety, observability, and workspace layout in `agent_os/`.

Provider routing

Shows a shared `LLMProvider` interface and router patterns for 24+ model providers in `providers/` and `utilities/`.

Coding-agent workflows

Includes repo instructions, prompts, and review workflows for Claude Code, Cursor, Codex, Aider, and Cline in `coding_agents/`.

Skills catalog

Documents skill design, taxonomy, maturity, packaging, validation, and a curated skill catalog in `skills/`.

Design and rollout templates

Provides agent design docs, ADRs, rollout plans, and machine-readable design specs in `design_docs/` and `templates/`.

Safety and evals

Covers guardrails, prompt-injection defense, approval policies, and evaluation methods in `safety/` and `evals/`.

Curated agent examples

Collects 100+ agent skeletons and worked examples under `agents/`, `blueprints/`, and `examples/`.

README

LLM Agents Ecosystem Handbook

A practical operating manual for building, evaluating, securing, and shipping modern LLM agent systems.

Awesome License: MIT PRs Welcome LLM-Friendly Providers


Modern agents are not "a prompt + a tool." They are systems — with identity, memory, skills, tools, MCP integrations, guardrails, observability, evals, and a provider strategy. This handbook teaches the whole stack and ships templates, blueprints, runnable adapters, and curated examples you can adopt today.

What's in this repo

A curated, opinionated, production-oriented handbook in seven parts:

  1. Concepts — Agent OS, identity, memory, skills, MCP, safety, observability — every layer of the modern agent stack
  2. Provider ecosystem — adapters + docs for 24+ LLM providers (frontier APIs, fast inference, marketplaces, enterprise clouds, specialty, local runtimes), with a router for fallback chains
  3. Skills ecosystem — design guide, taxonomy, maturity model, security checklist, and a curated skill catalog
  4. Prompt engineering — agent prompt patterns, instruction hierarchy, context engineering, prompt-injection defense
  5. Coding-agent workflows — for Claude Code, Cursor, Codex, Aider, Cline, and custom runtimes — repo instructions, prompts, review checklist, safe refactoring
  6. Design docs — agent / technical design docs, ADR guide, design reviews, rollout plans, the DESIGN.md machine-readable spec
  7. Curated catalog — 100+ existing agent skeletons, framework comparisons, evaluation tools, tutorials — preserved and improved

Who this is for

You are…Start at
New to agentsdocs/beginners_guide.mdagent_os/README.md
Building a production agentblueprints/checklists/production_readiness_checklist.md
Picking / wiring providersproviders/README.mdproviders/provider_matrix.md
Comparing frameworksdocs/framework_comparison.md
Adding memory / RAGmemory/tutorials/rag_tutorials
Adding MCPmcp/mcp/mcp_security.md
Designing Skillsskills/skills/skill_design_guide.md
Working with coding agentscoding_agents/coding_agents/prompts/
Writing better promptsprompt_engineering/
Designing & rolling outdesign_docs/
Hardening safety/evalssafety/evals/
Coding agent reading this repollms.txtllm_wiki/index.md

Modern Agent Stack

LayerPurposeWhere in this repo
Model / ProviderLLM choice + abstraction + routingproviders/
OrchestrationAgent loops, planning, handoffsdocs/framework_comparison.md, blueprints/
ToolFunction calling and external actionsagent_os/mcp_layer.md
MCPStandardized external context and toolsmcp/
MemoryDurable user/project/semantic memorymemory/
SkillsReusable, progressive-loading workflowsskills/
IdentityPersonality, mission, refusal styleagent_os/agent_identity.md, templates/
PromptSystem prompt design, instruction hierarchy, defensesprompt_engineering/
SafetyGuardrails, approvals, policysafety/
ObservabilityTracing, spans, cost, latency, evalsobservability/, evals/
DeploymentShipping agents to productiondesign_docs/rollout_plan.md
Coding-agent harnessClaude Code, Cursor, Codex, Aider, Clinecoding_agents/

📖 Deep dive: agent_os/README.md


Provider ecosystem

The handbook ships an LLMProvider abstraction with 24+ providers across six families. Most providers go through a single OpenAI-compatible code path; specialty / local providers are first-class.

Provider typeExamplesBest for
Frontier APIsOpenAI, Anthropic, Google GeminiReasoning, tool use, production agents
Fast inferenceGroq, Cerebras, SambaNovaLow-latency workloads
MarketplacesOpenRouter, Together, Fireworks, DeepInfraModel choice and routing
Enterprise cloudsAzure OpenAI, AWS Bedrock, Vertex AICompliance, governance
SpecialtyxAI, Perplexity, Mistral, Cohere, DeepSeek, Hugging Face, Replicate, NVIDIA NIM, MiniMaxDomain-specific
Local runtimesOllama, LM Studio, vLLM, llama.cppPrivacy, cost control, offline dev

If you want a governed OpenAI-compatible control plane in front of those providers, Tuning Engines is a useful runtime option for policy enforcement, approval gates, MCP and agent tracing, and usage or cost visibility without changing the surrounding agent framework.

Quick start:

from utilities import get_provider
from utilities.provider_router import ProviderRouter

# Use any single provider
out = get_provider("groq").chat(
    [{"role": "user", "content": "Summarize MCP."}],
    model="llama-3.1-8b-instant",
)

# Or route by task class with fallback
router = ProviderRouter()
out = router.chat(messages, task_class="cheap")  # Groq → DeepSeek → Together → OpenRouter

📖 providers/README.mdproviders/provider_matrix.mdproviders/router_patterns.mdproviders/local_models.md


Repository map

.
├── README.md • llms.txt • llms-full.txt
├── agent_os/                ← the Agent OS concept, layers, workspace examples
├── providers/               ← 24+ provider docs + adapters + router patterns
├── templates/               ← AGENTS.md / SOUL.md / MEMORY.md / SKILL.md / DESIGN_DOC / ADR / …
├── skills/                  ← design guide + taxonomy + maturity model + curated catalog + 4 examples
├── memory/                  ← memory taxonomy, distillation, security, examples
├── mcp/                     ← MCP basics, architecture, security, server catalog, examples
├── prompt_engineering/      ← agent prompt patterns, instruction hierarchy, defenses
├── coding_agents/           ← Claude Code, Cursor, Codex, workflows, prompts, review
├── design_docs/             ← agent + technical design docs, ADR guide, design.md spec
├── safety/                  ← guardrails, approvals, prompt injection, secure checklist
├── observability/           ← tracing, spans, cost/latency, dashboards
├── evals/                   ← eval design, regression / tool / memory / MCP / safety / prompt
├── blueprints/              ← production architectures by use case
├── examples/                ← end-to-end runnable agent workspaces
├── checklists/              ← agent design, prod readiness, MCP security, …
├── llm_wiki/                ← LLM-friendly index, glossary, matrices, wiki pattern
├── docs/                    ← framework comparison, best practices, beginners' guide
├── tutorials/               ← RAG, memory, fine-tuning, chat-with-X
├── utilities/               ← LLMProvider + router + provider_config
├── agents/                  ← 100+ curated agent skeletons (preserved)
├── complete_apps/, web_apps/, notebooks/, datasets/, design/, resources/, scripts/, tests/, ecosystem/
└── .github/                 ← issue / PR templates

Skills ecosystem

A curated, in-repo catalog plus a clear taxonomy and maturity model:

Curated skills shipped: research-summarizer, repo-auditor, mcp-security-reviewer, agent-memory-curator, api-design-reviewer, pr-summarizer, adr-writer, incident-postmortem, sprint-planner, dataset-profiler.


Prompt engineering

A dedicated section, agent-focused:

Templates: SYSTEM_PROMPT, AGENT_PROMPT. Checklist: agent_prompt_checklist.


Use this repo with coding agents

The handbook is itself a great surface for coding agents. Drop your favorite tool (Claude Code, Cursor, Codex, Aider, Cline) into the repo:

The guidance is tool-neutral: same AGENTS.md, same workflows, regardless of harness.


Design docs

Agent + technical design docs, ADRs, reviews, rollouts, and the DESIGN.md machine-readable spec for design tokens:

Templates: DESIGN_DOC, ADR.


Frameworks at a glance

FrameworkBest forLangMCPTracing
OpenAI Agents SDKProduction agentsPy / JS✅ built-in
LangGraphStateful, branching graphsPy / JS✅ LangSmith
CrewAIRole-based teamsPy⚠️ via partners
AutoGen (AG2)Event-driven multi-agent + HITLPy⚠️ partial
LlamaIndex WorkflowsData-heavy / RAG-firstPy / TS
Pydantic AIType-safe, FastAPI-nativePy✅ Logfire
SmolagentsCode-execution mini-agentsPy⚠️basic
Semantic Kernel.NET / enterprise / AzureC# / Py / Java
DSPyProgrammatic prompt optimizationPy
Strands AgentsProvider-agnostic, OpenTelemetryPy✅ OTEL
Vercel AI SDKApp-layer agents in Next.jsTS / JS
Google ADKGemini / Vertex hierarchical toolsPy

📖 Full comparison + decision tree: docs/framework_comparison.md. Capability tags hedged: verify against current upstream docs.


Skills, MCP, and Memory in one minute

  • Skills are reusable, model-loaded workflows (SKILL.md + scripts + references). Use when a task is repeatable, multi-step, and benefits from progressive disclosure. → skills/
  • MCP (Model Context Protocol) is a standard for exposing tools/context to any agent. Use when integrations should be reusable (GitHub, filesystem, browser, internal APIs). → mcp/
  • Memory is durable state across runs (MEMORY.md, vector stores, decision logs). → memory/

A useful rule of thumb:

If the thing is…Use
A repeatable workflow with steps and referencesSkill
An external system with tools to callMCP server
State that should outlive the current runMemory
A single function the model needs oncePlain tool

📖 Decision matrix: skills/skill_vs_tool_vs_mcp.md


Guardrails & safety

Production agents need risk-tiered tool controls and human approval gates for high-impact actions.

Risk levelExamplesApproval
Lowread-only search, summarizationnone
Mediumdrafting files, creating ticketssometimes
Highsending email, modifying repos, running shellrequired
Criticaldeleting data, spending money, changing permissionsalways + audit

📖 safety/README.mdsafety/prompt_injection.mdsafety/secure_agent_checklist.md


Observability & evals

You cannot ship what you cannot measure. The handbook ships:


Templates (copy-paste ready)

FilePurpose
AGENTS.mdRepo-specific agent instructions
SOUL.mdIdentity, voice, values, refusal style
MEMORY.mdDurable project + user memory index
USER.mdUser profile and preferences
TOOLS.mdAllowed/restricted/approval-gated tools
SKILL.mdSkill spec with progressive loading
MCP_SERVER.mdDocumenting an MCP integration
SYSTEM_PROMPT.mdLong-lived system prompt
AGENT_PROMPT.mdPer-task / per-session prompt
DESIGN_DOC.mdAgent / technical design doc
ADR.mdArchitecture Decision Record
EVAL_PLAN.mdWhat you'll evaluate and how
GUARDRAILS.mdPolicy, refusals, escalation
HUMAN_APPROVAL_POLICY.mdWho approves what
CODING_AGENT_TASK.mdTask contract for coding agents
REPO_MODERNIZATION_PROMPT.mdMulti-phase modernization
AGENT_RELEASE_CHECKLIST.mdShip/no-ship gate

Merged knowledge areas (1.0.1)

This release merged seven external projects into the handbook. Each was adapted (not bulk-copied) into the structure above:

Source themeLives in
Skills catalog + taxonomy patternsskills/ — taxonomy, maturity, packaging, validation, awesome catalog
Personal-wiki / self-maintaining KBllm_wiki/wiki_pattern.md, docs/llm_readable_docs.md
Agent prompt research patternsprompt_engineering/
Production coding-agent prompts + workflowscoding_agents/ — prompts, workflows, review
Machine-readable design specsdesign_docs/design_md_spec.md, templates/DESIGN_DOC.md.template
ADRs + design reviewsdesign_docs/adr_guide.md, design_docs/design_review.md

📖 Full migration plan: MIGRATION_AND_PROVIDER_EXPANSION_PLAN.md


Supported LLM providers

The utilities/llm_provider.py module exposes a single LLMProvider interface (and a backwards-compatible complete() function). Switch via LLM_PROVIDER without touching agent code; route automatically with ProviderRouter.

24+ providers across frontier / fast / marketplace / enterprise / specialty / local. See:


Contributing

Contributions are very welcome — new examples, framework updates, fixes, and translations all help. Start with:

Roadmap & changelog

License

MIT — see LICENSE.

Maintainer

Curated & maintained by Sayed Allam (oxbshw). If this handbook helped you ship, please ⭐ the repo and open a PR with what you learned along the way.

Files in the repo

Repository payload46 top-level entries
  • .github
  • agent_os
  • agents
  • blueprints
  • checklists
  • coding_agents
  • complete_apps
  • datasets
  • design
  • design_docs
  • docs
  • ecosystem
  • evals
  • evaluation_frameworks
  • examples
  • github
  • llm_wiki
  • mcp
  • memory
  • notebooks
  • observability
  • prompt_engineering
  • providers
  • resources
  • safety
  • scripts
  • skills
  • templates
  • tests
  • tutorials
  • utilities
  • web_apps
  • .env.example
  • .gitignore
  • CHANGELOG.md
  • CODE_OF_CONDUCT.md
  • CONTRIBUTING.md
  • LICENSE
  • llms-full.txt
  • llms.txt
  • MIGRATION_AND_PROVIDER_EXPANSION_PLAN.md
  • README.md
  • requirements.txt
  • ROADMAP.md
  • SECURITY.md
  • TRANSLATION.md

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More tutorials & guides

shareAI-lab/
learn-claude-code

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

77k
luongnv89/
claude-howto
luongnv89/claude-howtoTutorials & Guides

A visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.

41k
agentskills/
agentskills
agentskills/agentskillsTutorials & Guides

Specification and documentation for Agent Skills

25k