Sandbox
@Kaelio/ktx

Local context layer for data agents

ktx ingests databases, semantic layers, and team docs, then turns them into context agents can query safely. It exposes a CLI and MCP server so Claude Code, Codex, Cursor, and similar tools can search approved metrics, join paths, and wiki knowledge before writing SQL.

1,582 stars101 forksTypeScriptUpdated 14d ago
ktx - Open-source Context Layer for Data Agents
Kaelio1.9k views • 3 months ago
Who it's for

Builders who want their agents to query warehouses with company context instead of guessing metric logic.

What it delivers

You can get read-only, company-aligned answers from your agent without re-explaining the warehouse on every prompt.

What it does

Ingests company knowledge

Pulls in wiki content, modeling code, BI metadata, and database details into one local context layer.

Builds semantic context

Creates searchable semantic-layer YAML and join graphs so agents can use approved metrics and relationships.

Serves agents through MCP

Runs a local MCP daemon that lets agent clients search wiki and semantic-layer content at execution time.

Keeps queries read-only

Plans read-only SQL against the warehouse and does not write back to the database.

Supports common data stacks

Works with warehouses like PostgreSQL, Snowflake, BigQuery, ClickHouse, MySQL, SQL Server, SQLite, DuckDB, Athena, and MongoDB.

Includes agent setup files

Provides repo-level instructions and skill assets for Claude Code, Codex, and Gemini-based workflows.

How to get it

  1. 1Run
    npm install -g @kaelio/ktx
    ktx setup
    ktx status
  2. 2Example ktx status after setup
    ktx project: /home/user/analytics
    Project ready: yes
    LLM ready: yes (claude-sonnet-4-6)
    Embeddings ready: yes (text-embedding-3-small)
    Databases configured: yes (warehouse)
    Context sources configured: yes (dbt_main)
    ktx context built: yes
    Agent integration ready: yes (codex:project)
  3. 3[!TIP] Already using an agent? Ask Claude Code, Codex, Cursor, or OpenCode from your…
    Run npx skills add Kaelio/ktx --skill ktx and use the ktx skill to install
    and configure ktx in this project.

README

ktx

The context layer for data agents

npm version Codecov Tests Documentation Join the ktx Slack community License Y Combinator P25

Quickstart · CLI Reference · Agent Setup · Slack

Built and maintained by Kaelio


ktx is a self-improving context layer that teaches agents how to query your warehouse accurately - from approved metric definitions, joinable columns, and business knowledge it builds and maintains for you.

[!NOTE] Run ktx with your own LLM API keys or a local agent sign-in — a Claude Pro/Max subscription through Claude Code, or your local Codex authentication. No extra usage billing from ktx.

Watch the ktx launch video (1:56)

Ingestion: ktx ingests databases, BI tools, modeling code, and docs through its context engine (source connectors, context builder, reconciliation, validation) into wiki Markdown and semantic-layer YAML

Serving: an agent queries ktx through MCP, which searches the wiki and semantic layer, returns approved metrics, and compiles them into read-only SQL run against the warehouse

Why ktx

General-purpose agents struggle on data tasks. They re-explore your warehouse on every question, invent their own metric logic, and return numbers that don't match approved definitions.

Traditional semantic layers don't fix this. They demand constant manual upkeep and don't absorb the rest of your company's knowledge.

ktx does both, automatically:

  • Learns from company knowledge. Ingests wiki content, organizes it, removes duplicates, and flags contradictions for human review.
  • Maps the data stack. Samples tables, captures metadata and usage patterns, detects joinable columns, and annotates sources so agents write better queries.
  • Builds a semantic layer. Combines raw tables and high-level metrics through a join graph that automatically resolves chasm and fan traps, so agents fetch metrics declaratively instead of rewriting canonical SQL each time.
  • Serves agents at execution. Exposes CLI and MCP tools with combined full-text and semantic search across wiki and semantic-layer entities.

How ktx compares

General-purpose agentTraditional semantic layerktx
Builds warehouse context automatically
Detects joinable columns + resolves fan/chasm trapsManual
Approved, reusable metric definitions
Absorbs wiki / Notion / team knowledge
Flags contradictions across sources
Ships CLI + MCP for agent executionPartial
Read-only by designn/an/a

Who is ktx for

Use ktx if you:

  • Want agents like Claude Code, Codex, Cursor, or OpenCode to query your warehouse with approved metric definitions
  • Have business knowledge scattered across dbt, Looker, Metabase, Notion, and team wikis
  • Need agents to reuse canonical SQL instead of inventing it on every prompt

Skip ktx if you:

  • You don't have a SQL warehouse - ktx sits on top of one
  • You only need one ad-hoc query - psql or a notebook will do

Works with PostgreSQL, Snowflake, BigQuery, ClickHouse, MySQL, SQL Server, SQLite, DuckDB, Amazon Athena, and MongoDB. Integrates with dbt, MetricFlow, LookML, Looker, Metabase, Sigma, Notion, and Google Drive.

Quick Start

npm install -g @kaelio/ktx
ktx setup
ktx status

ktx setup creates or resumes a local ktx project, configures providers and connections, builds context, and installs agent integration.

Example ktx status after setup:

ktx project: /home/user/analytics
Project ready: yes
LLM ready: yes (claude-sonnet-4-6)
Embeddings ready: yes (text-embedding-3-small)
Databases configured: yes (warehouse)
Context sources configured: yes (dbt_main)
ktx context built: yes
Agent integration ready: yes (codex:project)

[!TIP] Already using an agent? Ask Claude Code, Codex, Cursor, or OpenCode from your project directory:

Run npx skills add Kaelio/ktx --skill ktx and use the ktx skill to install
and configure ktx in this project.

[!IMPORTANT] If ktx status prints ktx mcp start --project-dir ..., run it before opening your agent client.

Upgrading

Re-run the global install with the @latest tag:

npm install -g @kaelio/ktx@latest

First commands

CommandPurpose
ktx setupCreate, resume, or update a ktx project
ktx statusCheck project readiness
ktx ingestBuild context for every configured connection
ktx sl "revenue"Search semantic sources
ktx wiki "refund policy"Search local wiki pages
ktx mcp startStart the MCP server for agent clients

See the CLI Reference for every command, flag, and option.

Project Layout

my-project/
├── ktx.yaml                         # Project configuration
├── semantic-layer/<connection-id>/  # YAML semantic sources
├── wiki/global/                     # Shared business context
├── wiki/user/<user-id>/             # User-scoped notes
├── raw-sources/<connection-id>/     # Ingest artifacts and reports
└── .ktx/                            # Local state and secrets, git-ignored

Commit ktx.yaml, semantic-layer/, and wiki/. Keep .ktx/ local.

Project resolution defaults to KTX_PROJECT_DIR, then the nearest ktx.yaml, then the current directory. Pass --project-dir <path> when scripting.

FAQ

  • Does ktx send my schema or query results to a hosted service? No. ktx runs locally. The only data leaving your machine is what you send to the LLM provider you configured.
  • Which LLM backends are supported? Anthropic API, Google Vertex AI, AI Gateway, the local Claude Code session through the Claude Agent SDK, and your local Codex authentication through the Codex SDK. See LLM configuration.
  • How is ktx different from a dbt or MetricFlow semantic layer? ktx ingests those layers and combines them with raw-table introspection and wiki content. Agents get one searchable surface instead of three disconnected ones - and ktx flags contradictions across sources.
  • Does ktx need a running server? There is no hosted service. The local MCP daemon runs on demand via ktx mcp start when an agent client needs it.
  • Is my warehouse safe? Yes. Connections are read-only - ktx never writes to your database.

Docs

Community

  • Slack — ask questions, share what you're building, and chat with maintainers.
  • GitHub Issues — report bugs and request features.
  • Contributing — set up the repo, run tests, and open a PR.

Development

git clone https://github.com/kaelio/ktx.git
cd ktx
pnpm install
uv sync --all-groups
pnpm run build
pnpm run check

ktx is a pnpm + uv workspace:

PathPurpose
packages/cliTypeScript CLI and published npm package source
packages/cli/src/contextCore context engine
packages/cli/src/llmLLM and embedding providers
packages/cli/src/connectorsDatabase scan connectors
python/ktx-slSemantic-layer query planning
python/ktx-daemonPortable compute service

Local development CLI:

pnpm run setup:dev
pnpm run link:dev
ktx-dev --help

Useful checks:

pnpm run type-check
pnpm run test
pnpm run dead-code
uv run pytest -q

Telemetry

ktx collects privacy-conscious usage telemetry to understand installs and improve setup, command reliability, and data-agent workflows. Catalog telemetry events do not record file paths, hostnames, SQL, schema names, table names, column names, error messages, raw environment values, or argv. Error reports use PostHog Error Tracking and can include stack frames and raw error messages, which may contain local file paths or the local username in those paths. ktx redacts secrets, credentials, database URLs, auth headers, argv, raw environment values, SQL text, row data, and user-typed prompt or MCP argument text from the explicit $exception payload. See Telemetry for the event catalog and opt-out options.

License

ktx is licensed under the Apache License, Version 2.0. See LICENSE.

Star History

ktx Star History Chart

Files in the repo

Repository payload32 top-level entries
  • .github
  • assets
  • docs
  • docs-site
  • examples
  • packages
  • python
  • scripts
  • skills
  • .gitignore
  • .pre-commit-config.yaml
  • .releaserc.cjs
  • AGENTS.md
  • biome.json
  • CLAUDE.md
  • codecov.yml
  • conductor.json
  • CONTRIBUTING.md
  • GEMINI.md
  • knip.json
  • LICENSE
  • package.json
  • pnpm-lock.yaml
  • pnpm-workspace.yaml
  • pyproject.toml
  • README.md
  • release-policy.json
  • SECURITY.md
  • skills.sh.json
  • tombi.toml
  • tsconfig.base.json
  • uv.lock

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

86k

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k