Sandbox
@vixues/LeAgent

Desktop agent stack for local workflows and tools

LeAgent is a desktop AI agent platform that plans, calls tools, self-corrects, and runs visual workflows. It also streams generative UI, loads Agent Skills, and connects to tools through MCP and built-in offline actions.

220 stars40 forksPythonUpdated 1mo ago
Who it's for

Builders who want to run an agent locally, wire it to workflows, and keep their work on their own machine.

What it delivers

You can turn prompts into tool use, workflow runs, and finished documents without stitching together separate agent pieces.

What it does

Planning agent runtime

Runs multi-turn sessions that plan, call tools, and recover when a turn needs to pause and resume.

Visual workflows

Lets the agent design and run ReactFlow DAGs, with typed nodes for registered tools and reusable templates.

Generative UI

Streams live UI blocks like KPI boards, slide decks, and galleries into chat and can export them.

Agent Skills support

Loads `SKILL.md` bundles, supports install from links or archives, and can use an HTTP skill registry.

Offline tool catalog

Ships 100+ built-in tools for documents, web research, data, code, databases, media, and automation.

MCP and automation hooks

Connects to MCP servers, webhooks, cron jobs, and outbound channels through declarative rules.

How to get it

  1. 1Prerequisites: git, uv, Node.js 20+ or 22+
    git clone https://github.com/vixues/LeAgent.git
    cd LeAgent
    ./start.sh                # backend :7860 + frontend :5173
  2. 2Run
    cd LeAgent/deploy
    cp .env.example .env      # set LEAGENT_SECRET_KEY + at least one provider key
    docker compose up -d --build
  3. 3Run
    curl -fsSL https://vixues.com.cn/install.sh | bash

README

LeAgent Logo

LeAgent

The open-source desktop AI agent that gets work done — an agent that plans and self-corrects, agentic visual workflows, generative UI, and 100+ offline tools in a single self-hostable stack.

CI Latest Release Python 3.11+ React 19 SQLite / PostgreSQL License PRs Welcome

中文文档 · 汉文 · Tutorial · Contributor guide · Releases

LeAgent Screenshot


LeAgent is an open-source desktop AI agent that doesn't just chat — it gets work done. Unlike cloud chatbots and code-only CLIs, LeAgent fuses three capabilities most agents keep apart: a streaming agent runtime that plans, calls tools, and self-corrects in one think-act loop; agentic visual workflows where the agent designs, runs, and refines ReactFlow DAGs (and every one of its tools is automatically a typed node); and a generative UI layer that streams live, interactive interfaces — KPI boards, slide decks, galleries — right into the chat. It ships 100+ built-in offline tools (documents, web, data, code, databases, media, game-art generation) plus a declarative rule engine, Agent Skills, and the Model Context Protocol — all running locally with zero external dependencies by default (SQLite, single process). Bring your own model keys, or run fully offline against a local Ollama / vLLM endpoint.

It is built for people who want a private, hackable, self-hosted alternative to closed agent products: your documents, sessions, and credentials never leave the machine unless you point a provider at a remote API.

Highlights

  • Agent runtime — multi-turn sessions with token streaming, tool execution, tiered model routing, layered prompt assembly, and a cognitive three-store memory (episodic / semantic / procedural). The QueryEngine session orchestrator drives both chat and background paths through one think-act loop, with durable checkpoints for pausing and resuming turns.
  • 100+ offline tools — documents, web research, data wrangling, code execution, databases, generative UI, charts, media, and coding projects (see the tool catalog below).
  • Visual workflows — a ReactFlow editor with YAML export, reusable templates, and automatic typed nodes for every tool. The engine stages ready batches, runs independent branches concurrently, and applies centralized retry/backoff and timeouts.
  • Agent SkillsAgent Skills v1.0 SKILL.md bundles with progressive disclosure and on-demand loading. Ship built-in skills, install from links or archives, or plug in an HTTP skill registry.
  • Research Paper Mode — open a PDF and the agent becomes a citation-grounded research analyst: structure & outline extraction, faithful section / whole-paper summaries, reference and LaTeX-formula extraction, and region translation — with reader tabs and matching agent tools, text extraction fully offline. (guide)
  • Generative UI — agents stream declarative UI trees (KPI boards, slide decks, galleries, steppers) that render inline in chat and export to PDF or PPTX.
  • Game-art asset pipeline — first-class, composable generation nodes (image / video / 3D mesh / VFX) with typed media sockets, a quality gate plus a bounded self-correction loop, and inline canvas previews. Runs end-to-end offline with no credentials. (docs)
  • Multi-provider LLM — DeepSeek, DashScope (Qwen), OpenAI, Anthropic, Azure OpenAI, Ollama, and vLLM, with cost-tiered routing and failover. DeepSeek is the most thoroughly validated and recommended for first use.
  • Sidebar desk pet — a customizable avatar with walk/jump animations and personality bubbles, synced to chat streaming and session state. Upload PNG / SVG / GIF or sprite sheets.
  • Integrations — MCP servers, inbound webhooks, scheduled cron jobs, outbound channels (IM/console), and a declarative YAML rule engine.
  • Zero-config default — SQLite out of the box and a single Docker container. Scale out with PostgreSQL and Milvus (vector-backed memory) when you need to.

What you can build

LeAgent is a complete agent platform, not a starter kit — the building blocks below are wired together and work out of the box, fully offline by default.

  • Office automation — point it at a folder of invoices, contracts, or reports: it OCRs and classifies the files, extracts structured data into Excel, and produces a polished Word or PPTX summary in a single turn.
  • Research & briefings — search across DuckDuckGo / SearXNG / Bing, scrape and download sources, then assemble a cited PDF report with charts and KPI dashboards.
  • Data analysis — load CSV/Excel or query the database, clean, merge, aggregate, run SQL and vector search, and narrate the findings with generated charts.
  • Coding companion — scaffold a project from a template, edit files across the tree, run a live dev server behind a preview proxy, and execute code in an isolated sandbox.
  • Game-art production — turn a text brief into images, video, 3D meshes, and VFX sprite sheets through a node graph that self-corrects against a quality gate and exports an engine-ready (Unity / Unreal / Godot) bundle.
  • Live, interactive answers — stream Generative UI (KPI boards, slide decks, galleries, steppers) straight into chat, then export it to PDF or PPTX.
  • Always-on automation — schedule cron jobs, trigger flows from inbound webhooks, fan out to IM channels (DingTalk / Feishu / WeChat Work), and gate behaviour with declarative YAML rules.

Architecture

LeAgent is an async Python (FastAPI) backend plus a React 19 single-page app, packaged as a modular monolith. The backend uses a strict, downward-only layered domain model — File → Code → Project — over a single persistence layer, and every agent turn (chat, SDK, background task, sub-agent, workflow node) converges on one think-act kernel.

LeAgent/
├── backend/                 # FastAPI backend (Python 3.11+, uv-managed)
│   └── leagent/
│       ├── agent/           # QueryEngine orchestrator, planner, subagents
│       ├── sdk/             # Versioned public Agent SDK (runtime, kernel, protocols)
│       ├── api/             # FastAPI routers (v1 + incubating v2)
│       ├── llm/             # LLM service, providers, transport, streaming, generation
│       ├── tools/           # 100+ tools across 13 categories
│       ├── workflow/        # Workflow engine, nodes, art-asset pipeline, templates
│       ├── context/         # Source-driven, relevance-gated prompt assembly
│       ├── prompts/         # Layered PromptBuilder, registry, templates
│       ├── memory/          # Episodic / semantic / procedural agent memory
│       ├── skills/          # Agent Skills v1.0 loader + registry
│       ├── rules/           # Declarative YAML rule engine
│       ├── mcp/             # Model Context Protocol
│       ├── file/ code/ project/   # Layered file → code → project domain model
│       ├── db/              # Persistence: engine, SQLModel models, repositories
│       └── services/        # DB, auth, chat, session, gen-ui, cron, ...
├── frontend/                # React 19 + TypeScript SPA (Vite, Zustand, React Query)
├── desktop/                 # Electron shell (bundled Python runtime)
├── deploy/                  # Dockerfile + SQLite-only Compose
├── config/                  # Demo workflows + workflow templates
├── docs/                    # Architecture, guides, deployment docs
└── start.sh / start.ps1     # Dev orchestrator (uv + npm)

Execution model

Regardless of where a request enters, it flows through one set of well-defined boundaries. Each ingress mints a single ExecutionRun (with a run_id and, for child scopes, a parent_run_id), goes through a thin facade, runs on the shared kernel, and persists to durable state with one owner per state class:

Ingress   HTTP/SSE · WebSocket · Cron · Background task · GenUI
   │
   ▼
Facade    ServiceManager.runtime_context · AgentRuntime · WorkflowService
   │
   ▼
Kernel    run_loop → QueryEngine → tool executor          (single think-act loop)
   │
   ▼
State     TieredSessionStore · CheckpointStore · WorkflowStateStore
   │
   ▼
Observe   EventManager (FLOW_*/TASK_*/AGENT_*) · OpenTelemetry
  • One kernel. Chat and every background path drive through leagent.sdk.kernel.run_loop. Turns pause to a durable checkpoint when they await user input, then resume exactly where they left off.
  • Three workflow shapes, one engine. Saved DAG flows, chat playbook step-cards (compiled to linear flows), and in-chat graph embeds all execute on the same WorkflowExecutor, which stages ready batches, runs independent branches concurrently, and applies centralized retry/backoff and timeouts.
  • Clear state ownership. The chat transcript lives in TieredSessionStore, paused turns in CheckpointStore (agent_checkpoints), and workflow runs in WorkflowStateStore — no shared mutable state across subsystems.

See AGENTS.md for the full subsystem map and docs/technical/execution-topology.md for the authoritative agent-loop / workflow-engine state contract.

Tool catalog

Every tool is auto-exposed as a typed workflow node, so anything the agent can call can also be wired into a visual flow.

CategoryWhat it covers
DocumentsRead/write Word, Excel, PPTX, PDF; OCR; classification; archives; text processing
WebSearch (DuckDuckGo / SearXNG / Bing), scraping, image & native media download
DataClean, merge, validate, transform, aggregate, SQL & vector search
CodeSandboxed in-process scripts and a subprocess code-execution agent
DatabaseSchema-aware querying over the managed database
GenerateWord / Excel / PPTX / PDF / report / checklist / template-fill generators
Canvas / GenUIStream and patch declarative UI trees; publish canvases
Charts & ImagesChart generation and image processing
MediaImage / video / 3D / audio generation backends
SkillsDiscover, install, and invoke Agent Skills
WorkflowSave, run, and inspect workflows from inside an agent turn
IntegrationMCP, webhooks, channels, and external service calls
UtilitiesCron, tasks, rule matching, folders, text splitting, pet bubbles, and more

Capabilities in depth

Agent runtime & memory

  • Multi-turn streaming sessions with tiered model routing, automatic context compaction (micro + auto), and abort-safe tool execution.
  • Hybrid reasoning — ReAct-style tool loops plus plan-and-execute — with sub-agent delegation and per-turn recovery.
  • Cognitive three-store memory: episodic (past turns), semantic (extracted facts), and procedural (tool success rates), with hybrid semantic + lexical recall that degrades gracefully when no vector store is configured.
  • Layered prompt assembly with relevance-gated policy/playbook sources, per-layer and global budgets, and provider-aware rendering (including Anthropic prompt-cache boundaries).

Workflow engine

  • One DAG executor backs saved flows, in-chat step cards, cron jobs, and agent-authored graphs.
  • Concurrent branch execution by ready batches, centralized retry/backoff, per-node timeouts, and durable pause/resume.
  • Every registered tool is lifted into a typed Tool.<name> node automatically — exposing a new capability visually needs zero glue code.

Tools & code execution

  • 100+ first-party tools across 13 categories, dispatched through one executor with a path sandbox and permission hooks.
  • Two-tier code execution: a fast in-process RestrictedPython sandbox for workflow scripts, and a subprocess sandbox (rlimits, timeouts, per-session workspace) for heavier work.
  • Files are first-class — tools return FileRefs through a unified file layer with HMAC-signed preview/download URLs.

Skills, MCP & integrations

  • Agent Skills v1.0 (SKILL.md) with progressive disclosure: ship built-ins, install from links/archives, or connect an HTTP skill registry.
  • Model Context Protocol client for external tool servers; inbound webhooks; outbound IM/console channels; and a declarative YAML rule engine for guardrails and automation.

Generative UI & media

  • Declarative UI trees stream and patch live over SSE, render inline in chat, and export to PDF or PPTX.
  • A first-class media plane (image / video / 3D / VFX / audio) behind a strategy-and-registry generation service with retry + failover, plus a deterministic offline floor that needs no credentials.

Multi-model routing

Cost-tiered routing (tier1 reasoning / tier2 fast) with automatic failover — bring cloud keys, or stay fully local.

ProviderNotes
DeepSeekRecommended default; auto-aliased to tier1 (v4-pro) + tier2 (v4-flash); reasoning content + prompt-cache metrics
DashScope (Qwen)Thinking + search modes
OpenAI / Anthropic / Azure OpenAICloud frontier models
Ollama / vLLMFully local / self-hosted OpenAI-compatible inference

Quick Start

Local dev (recommended for hackers)

Prerequisites: git, uv, Node.js 20+ or 22+

git clone https://github.com/vixues/LeAgent.git
cd LeAgent
./start.sh                # backend :7860 + frontend :5173

The dev orchestrator syncs the Python env with uv, installs frontend deps, and (unless skipped) installs the Playwright Chromium used by web tools.

Docker

cd LeAgent/deploy
cp .env.example .env      # set LEAGENT_SECRET_KEY + at least one provider key
docker compose up -d --build

API and interactive docs at http://localhost:8000/docs. The default image is a single SQLite-backed container; optional overlays add a local GPU vLLM service (docker-compose.vllm.yml).

Manual setup

# Backend
cd backend
uv sync --extra dev
uv run leagent init
uv run leagent app

# Frontend (separate terminal)
cd frontend
npm install && npm run dev

One-line install

curl -fsSL https://vixues.com.cn/install.sh | bash

Configuration

Set at least one provider key (env var or Settings → Environment secrets in the web UI, which writes ~/.leagent/.env). The most common knobs:

VariablePurpose
LEAGENT_SECRET_KEYApp secret for signed URLs and session crypto (openssl rand -hex 32)
DEEPSEEK_API_KEYDeepSeek provider — auto-aliased as tier1 (reasoning) / tier2 (fast)
OPENAI_API_KEY / ANTHROPIC_API_KEY / DASHSCOPE_API_KEYAdditional cloud providers
VLLM_ENDPOINT / LLM_OLLAMA_ENDPOINTLocal / self-hosted OpenAI-compatible inference
DATABASE_URLSwitch from SQLite to PostgreSQL
LEAGENT_DEBUGEnable debug logging

See deploy/.env.example for the full annotated list.

Desktop app (Beta — features still being refined)

Installers for each platform ship with every GitHub release — download and run. No separate Python, Node, or Docker install required; the build bundles its own Python runtime and backend.

PlatformDownloadNotes
Windows 10/11 (x64)LeAgent-Setup-*.exeNSIS installer; desktop + start-menu shortcut
macOS (Apple Silicon)LeAgent-*-arm64.dmgUnsigned — xattr -dr com.apple.quarantine /Applications/LeAgent.app after install
macOS (Intel)LeAgent-*.dmgSame Gatekeeper note as above
Linux (x64)LeAgent-*.AppImage / LeAgent-*.debAppImage: chmod +x then run. .deb: sudo dpkg -i

See all releases: https://github.com/vixues/LeAgent/releases

Tech stack

LayerTechnology
BackendPython 3.11+, FastAPI, Uvicorn/Gunicorn, SQLModel + Alembic, Pydantic v2, async I/O, OpenTelemetry
FrontendReact 19, TypeScript, Vite, Zustand, TanStack Query, ReactFlow, i18next (zh-CN / en-US / 汉文)
DesktopElectron (ESM main process), bundled Python backend
DataSQLite (default), PostgreSQL (optional), Milvus (optional vector memory)
Toolinguv (Python), npm (frontend), Playwright, black + ruff, ESLint

Operations

  • Ports. Local dev serves the backend on :7860 and the Vite frontend on :5173 (start.sh); the Docker image publishes the API on :8000.
  • Persistence & backup. State lives under LEAGENT_HOME: the SQLite database (WAL mode) plus the working-uploads, knowledge, and coding-project trees. A complete backup is the database and that directory.
  • Scaling. The default single-process / single-worker setup is correct for SQLite (single writer). To scale out, switch to PostgreSQL via DATABASE_URL and front the app with sticky sessions (the in-process execution registry and event bus are per-worker). Milvus is optional and only powers vector-backed memory recall.
  • Observability. Structured JSON logs (structlog), OpenTelemetry spans when an OTLP endpoint is configured, and Prometheus workflow/quality histograms. Interactive API docs at /docs; set LEAGENT_DEBUG=true for verbose tracing.

Documentation

The full documentation set lives in docs/ — start with the architecture overview.

Contributing

Issues and pull requests are welcome. Please:

  1. Open an issue for larger changes or ambiguous scope.
  2. Run tests for touched areas (cd backend && uv run pytest tests/ -v / cd frontend && npm run test).
  3. Follow AGENTS.md for coding conventions and i18n rules (every new UI string must exist in both zh-CN and en-US bundles).

See CONTRIBUTING.md for full guidelines and CODE_OF_CONDUCT.md for community standards.

License

Apache License 2.0 — see LICENSE.

Files in the repo

Repository payload23 top-level entries
  • .github
  • backend
  • config
  • deploy
  • desktop
  • docs
  • frontend
  • scripts
  • website
  • .gitignore
  • AGENTS.md
  • CHANGELOG.md
  • cliff.toml
  • CODE_OF_CONDUCT.md
  • CONTRIBUTING.md
  • LICENSE
  • mkdocs.yml
  • README_lzh.md
  • README_zh.md
  • README.md
  • SECURITY.md
  • start.ps1
  • start.sh

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More agents

Hmbown/
Codewhale

Open-source coding agent for your terminal, built in Rust and on a journey of continuous community improvement. Issues and PRs welcome.

41k

A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps,, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK

31k
TokenRhythm/
opensquilla

OpenSquilla — Token-Efficient AI Agent with same budget, higher intelligence density

7k

An open-source AI coding agent that lives in your terminal.

28k
Untrivial-ai/
agent-orchestrator

Run and supervise teams of coding agents from planning to merge. Any harness (Claude code, codex, +25 more). Desktop, web, mobile, and cloud agents.

11k