Sandbox
@aaronsb/knowledge-graph-system

Knowledge graph CLI and MCP server for grounded memory

Kappa Graph extracts concepts, relationships, and evidence from documents, then stores them in PostgreSQL with Apache AGE. It tracks grounding strength, preserves disagreements, and lets agents and humans query the graph through the CLI, MCP server, API, or FUSE filesystem.

124 stars20 forksPythonUpdated 2mo ago
Who it's for

Builders who want an agent-friendly knowledge graph with source tracing and confidence scores.

What it delivers

You can ask an agent to use grounded memory instead of re-reading or re-explaining the same sources.

What it does

Concept extraction from documents

Ingests PDFs, markdown, images, and text, then extracts concepts, relationships, and evidence.

Grounding scores

Assigns each concept a support measure so you can see what is well evidenced and what is weak or contested.

Contradiction preservation

Keeps disagreeing sources visible instead of forcing one answer.

MCP server for assistants

Lets AI assistants query the graph as persistent memory through MCP.

FUSE filesystem access

Mounts the graph so you can browse semantic data with standard Unix tools like `ls`, `grep`, and `find`.

Graph visualization

Includes a web explorer for searching, clustering, and tracing relationships in the graph.

How to get it

  1. 1The kg CLI, MCP server (for AI assistants), and optional FUSE filesystem
    curl -fsSL https://raw.githubusercontent.com/aaronsb/knowledge-graph-system/main/client-manager.sh | bash
  2. 2Run your own knowledge graph backend
    curl -fsSL https://raw.githubusercontent.com/aaronsb/knowledge-graph-system/main/install.sh | bash
  3. 3Or from source
    git clone https://github.com/aaronsb/knowledge-graph-system.git
    cd knowledge-graph-system
    ./operator.sh init    # Interactive setup
    ./operator.sh start   # Start containers

README

Kappa Graph — κ(G)

License GitHub stars Latest Release

A semantic knowledge graph that extracts concepts from documents, tracks how well-supported they are, and remembers where sources disagree.

κ(G) — vertex connectivity of a graph. The minimum number of connections you'd need to cut before the graph falls apart. A measure of how robust the structure is.

Also kg — the unit of mass. Because knowledge here has weight. Grounding scores measure how heavy an idea is: well-evidenced claims carry more than thin ones. Contested concepts weigh differently than unchallenged ones.

2D Force Graph Explorer

Quick Start

Install Client Tools

The kg CLI, MCP server (for AI assistants), and optional FUSE filesystem:

curl -fsSL https://raw.githubusercontent.com/aaronsb/knowledge-graph-system/main/client-manager.sh | bash

Or just the CLI: npm install -g @aaronsb/kg-cli

Deploy the Platform

Run your own knowledge graph backend:

curl -fsSL https://raw.githubusercontent.com/aaronsb/knowledge-graph-system/main/install.sh | bash

Or from source:

git clone https://github.com/aaronsb/knowledge-graph-system.git
cd knowledge-graph-system
./operator.sh init    # Interactive setup
./operator.sh start   # Start containers

See Quick Start Guide for details.

See It In Action

2D Force Graph Explorer Interactive graph exploration with smart search, concept clustering, and relationship visualization

CLI with Inline Image Evidence Command-line search returns concepts with source images rendered inline via chafa

Embedding Landscape with DBSCAN Clusters t-SNE embedding landscape with auto-detected clusters, named by topic via TF-IDF

What You Can Do

Ingest documents — PDFs, markdown, images, text. The system extracts concepts, relationships, and evidence automatically.

Search by meaning — "economic downturn" finds content about recessions, crashes, and crises even if those exact words aren't used.

Explore connections — Find paths between concepts. See how ideas relate across documents.

Check confidence — Every result includes grounding scores. Know what's well-supported vs. contested.

Trace sources — Every concept links back to the original text or image that generated it.

Query via AI — MCP server integration lets Claude and other assistants use the graph as persistent memory.

Navigate via filesystem — Mount the graph as a FUSE filesystem. Use ls, grep, find on semantic space.

Use Cases

Obsidian viewing the knowledge graph via FUSE Obsidian's graph view rendering knowledge graph relationships via the FUSE filesystem — no plugin needed

Research synthesis — Ingest papers, find connections across them, see where authors disagree. Grounding scores tell you which claims have broad support.

Technical documentation — Extract architecture concepts from diagrams, meeting notes, design docs. Query how components relate.

Agent memory — Give AI assistants persistent, grounded memory. They can check confidence before making claims.

Claude Desktop exploring the knowledge graph Claude Desktop using MCP to search, explore relationships, and validate claims against the knowledge graph

Spatial understanding — Ingest place photos. The graph learns physical relationships without coordinates.

Compliance/audit — Full provenance chain. Every concept traces to source evidence.

Architecture

Documents ──→ [FastAPI] ──→ LLM Extraction ──→ [PostgreSQL + AGE]
                  │                                    │
                  │                              [graph_accel]
                  │                            in-memory traversal
                  │                                    │
              [Garage S3]                        [AGE graph store]
               doc storage                     source of truth (ACID)
                  │                                    │
              [React + D3] ←──── REST API ────→ [FastAPI]
            web visualization                   query + ingest
                  │
           [CLI / MCP / FUSE]
           client interfaces
  • PostgreSQL + Apache AGE — Graph database with native openCypher queries. ACID transactions, schema integrity, vector search (pgvector).
  • graph_accel — In-memory graph traversal accelerator. A Rust PostgreSQL extension that maintains an adjacency structure in shared memory for instant BFS/shortest-path at any depth. AGE handles writes; graph_accel handles reads. Epoch-based invalidation ensures the read model is never stale. (ADR-201)
  • FastAPI — Extraction pipeline and REST API
  • React + D3 — Interactive graph visualization and exploration
  • TypeScript CLI — Command-line interface and MCP server for AI assistant integration
  • Ollama — Optional local inference (air-gapped operation)
  • Garage — S3-compatible object storage for document assets

Documentation

AudienceStart Here
Understanding the conceptsdocs/concepts/
Deploying and operatingdocs/operating/
Using the systemdocs/features/
Architecture decisionsdocs/architecture/

96 Architecture Decision Records document the design evolution.

Background

Most systems that store knowledge for retrieval — vector databases, RAG pipelines, knowledge graphs — optimize for finding relevant content. They can tell you what matches your query. They can't tell you how well-supported it is, whether sources disagree about it, or where the evidence actually came from.

kg adds an epistemic layer on top of the graph. Every concept carries a grounding score computed from supporting vs. contradicting evidence — not a label, but a measurement. A concept backed by 47 sources with 12 contradictions scores differently than one with a single unchallenged mention. When sources disagree, the system preserves both sides rather than picking a winner.

Semantic diversity provides a second signal. Well-established knowledge tends to connect across independent domains. Narrow claims that only reference each other score lower. In testing, Apollo 11 mission data showed 37.7% diversity across 33 concepts; moon landing conspiracy content showed 23.2% across 3.

The system also handles images. Feed it street view photos and the extracted relationships ("next to", "across from", "visible from") naturally encode spatial topology — no coordinates needed.

How It Compares

CapabilitykgGraphRAGZep/GraphitiCogneeVector DBs
Contradiction detectionNative (mathematical)LLM-dependentLimitedNoNo
Grounding scoresContinuous -1 to +1Source citations onlyNoNoSimilarity only
Semantic diversityYes (authenticity signal)NoNoNoNo
Epistemic statusPer-relationshipNoNoNoNo
Temporal trackingIngestion epochNoBi-temporalNoNo
Emergent ontologiesContinuous annealingOne-shot communitiesTemporal factsPartial (event-driven)N/A
Multimodal ingestProse-bridge (own vector space)Text-centricText-centricImages/audioEmbeddings only
FUSE filesystemYesNoNoNoNo
Air-gapped operationYes (Ollama)Cloud requiredCloud requiredLocal-capableSome local

Each neighbor leads on a different axis: Zep/Graphiti on bi-temporal fact tracking, GraphRAG on adoption (its Leiden communities are the closest analog to annealing, but computed once at index time), and Cognee — the nearest architectural neighbor — on self-improving memory (event-driven, no epistemic layer). kg's temporal axis is deliberately narrower: it records when evidence arrived (ingestion epoch), not when a fact was true — because it holds no truth values, only computed evidence, and preserves contradictions rather than expiring them. That's a stance, not a gap — see Computed Evidence over Asserted Truth. What we haven't found a shipped peer for is the combination: continuous ontology annealing, a semantic FUSE surface, and an epistemic layer that measures confidence rather than just retrieving content.

Why Try It

If you need:

  • Epistemic reliability — knowing how sure you should be, not just what the answer is
  • Contradiction awareness — preserving disagreement rather than hiding it
  • Full provenance — tracing every claim to source evidence
  • Local operation — running without cloud API dependencies
  • Unix integration — using standard tools on semantic data

kg was built for those requirements. Most alternatives optimize for retrieval accuracy or comprehensiveness. kg optimizes for knowing what you know and how well you know it.

License

Apache License 2.0 — Use, modify, distribute freely. Patent grant included.

Acknowledgments

Built with Apache AGE, Model Context Protocol, FastAPI, and local inference via Ollama.

Files in the repo

Repository payload38 top-level entries
  • .claude
  • .githooks
  • .github
  • api
  • appliance
  • cli
  • config
  • docker
  • docs
  • examples
  • fuse
  • graph-accel
  • operator
  • schema
  • scripts
  • specs
  • spike
  • tests
  • web
  • .env.example
  • .gitignore
  • CLAUDE.md
  • client-manager.sh
  • CODE_OF_CONDUCT.md
  • CONTRIBUTING.md
  • install.sh
  • LICENSE
  • Makefile
  • mkdocs.yml
  • operator.sh
  • package-lock.json
  • package.json
  • publish-wizard.sh
  • publish.sh
  • pytest.ini
  • README.md
  • requirements.txt
  • VERSION

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More tools

JuliusBrussee/
caveman

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

105k
1 add
MemPalace/
mempalace

The best-benchmarked open-source AI memory system. And it's free.

59k
stablyai/
orca

Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.

66k

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

132k

Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 550+ contributors

64k
headroomlabs-ai/
headroom

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

71k