Sandbox
@mattgierhart/PRD-driven-context-engineering

Claude Code plugin for PRD-led context engineering

This repository packages a PRD-driven workflow for Claude Code around a markdown knowledge graph, stage skills, and event hooks. It keeps product decisions in `SoT/`, active work in `epics/`, and uses readiness scripts to gate progress and surface drift.

182 stars10 forksHTMLUpdated 17d ago
Who it's for

Builders who want Claude Code to keep product context, decisions, and implementation aligned across sessions.

What it delivers

You can move agent-led product work forward with remembered context instead of re-explaining the project every session.

What it does

Progressive PRD

`PRD.md` advances through gated stages from problem framing to launch, with each stage requiring specific decisions before moving on.

Markdown knowledge graph

`SoT/` stores durable product memory as linked markdown files with typed IDs instead of a database.

Claude Code skills and hooks

`.claude/skills` and `.claude/hooks` coordinate the workflow, load context, check gates, and manage session memory.

Readiness scoring

`scripts/readiness.py` computes whether the repo, epics, and PRD stage are ready to advance.

Agent squad

`.claude/agents` defines role agents that keep persistent memory and hand off work through files.

Self-install path

`install.sh` and `.claude/install-manifest.yaml` let the framework merge into an existing repo without overwriting product files.

How to get it

  1. 1You don't have to fork. The self-install path drops the framework into a fresh or…
    # Deterministic CLI (from a clone of this repo):
    bash install.sh --target /path/to/your/repo --profile product --dry-run   # preview
    bash install.sh --target /path/to/your/repo --profile product             # install

README

PRD-Led Context Engineering

Memory as Infrastructure — an operating system for building products with AI agents

GitHub stars License: MIT PRs Welcome Built for Claude Code

Your AI partner is brilliant in one session and amnesiac by the next. This repository is the fix: a fork-ready methodology that turns documentation into a knowledge graph humans and AI navigate together — so the 50th session is smarter than the 1st.

Quick Start · The Idea · The Lifecycle · The Skills · Live Demo Views

The Source-of-Truth Atlas — the knowledge graph rendered for humans

⭐ If this changes how you build with AI, star the repo — stars put this method in front of the next team drowning in context drift.


The Problem: AI forgets. Teams drift.

Every era of software solved memory its own way — and broke it its own way:

  • Waterfall had Static Memory. Write everything down upfront: certainty, at the cost of change.
  • Agile moved fast and created Fragmented Memory. Knowledge scattered across tickets, wikis, and chats; shared understanding evaporated.
  • AI-assisted building adds a third failure: Amnesiac Memory. The model that architected your system yesterday has never heard of it today.

The common mistake is treating these as tooling problems. They are memory problems.

PRD-Led Context Engineering builds Shared Memory: it treats AI as a team member, not a tool, and keeps documentation synchronized with code so humans and AI navigate the same truth.


The Idea: Memory as Infrastructure

This methodology comes from two converging experiences.

Leading human teams — alignment always followed the same pattern: rally around a single Source-of-Truth artifact and the team moves as one. Without it, even great talent drifts.

Partnering with AI — sometimes the model performs at a senior level, sometimes it hallucinates. The variable was never the model's intelligence. It was the Context Density provided: rich, structured context in; senior-level output out.

The convergence: documentation is not an afterthought. Documentation is the infrastructure of shared memory.

The Golden Rule: If it isn't part of the memory infrastructure, it isn't true.

So every durable decision gets a Unique ID (UJ-101, BR-004, API-045) in a Source-of-Truth file. That ID is a memory node with weight: when the AI references BR-004, it isn't guessing — it's retrieving a specific, validated decision you encoded. The linked network of IDs across files is the Knowledge Graph, and it lives in plain markdown, in your repo, under version control.

The 4 Pillars

  1. Just-in-Time Context — IDs let you load only what a task needs. No repo-dumps, no noise, fewer hallucinations.
  2. The Documentation Ecosystem — PRD ↔ EPICs ↔ SoT connected by links, skills, and hooks. When logic changes, both human and AI know.
  3. Context Validation — context is measured like code. Too sparse, the AI drifts; too dense, it chokes. The repo scores itself (see Readiness Scoring).
  4. Progressive Documentation — update in place, never copy. No PRD_v2.md, ever. One document, many versions, single current reality.

The Cognitive Shift

This changes how work is measured, not just how it's tooled:

Traditional AgilePRD-Led Context EngineeringThe Shift
SprintsContext WindowsWe don't time-box based on dates; we scope-box based on cognitive capacity.
User StoriesPromptsWe don't write descriptions; we engineer prompts that deterministically load context.
Tribal KnowledgeSource of TruthIf it isn't in the Knowledge Graph (SoT/), it doesn't exist.
StandupsDocumentation HooksNo status meetings. Event-driven hooks handle context loading, gate checks, and memory handoffs.
Project ManagementContext GovernanceWe don't task-manage people. The system gates execution until context is verified valid.

What's in the box

Everything below ships in this repo, works offline, and forks in one click:

FeatureWhat it gives you
🧠 The Knowledge Graph14 SoT files, 21 ID types, zero databases — durable memory in markdown
📈 The Progressive PRDA gated v0.1 → v1.0 lifecycle that stops AI from one-shotting your architecture
🛠 47 SkillsStage playbooks from problem framing to crossing the chasm — Dunford, Hormozi, Moore, Torres built in
📊 Readiness ScoringThe repo computes whether you're ready to advance — and what to fix first
🫀 The Development Graph@implements tags bridge code to specs; drift surfaces as a verdict, not a surprise
📰 The Human Review LayerEvery SoT file rendered as a styled, hyperlinked page its reviewer actually wants to read
🤖 The Agent SquadFour role agents with persistent memory, coordinated through files instead of meetings

🧠 Feature: A Knowledge Graph in plain markdown

The pitch: long-term product memory with no database, no SaaS, no lock-in — just files with discipline.

The architecture is 3 + 1 + SoT + Temp, designed to manage Context Density for both human cognitive load and AI context windows:

  1. Executive Functions — orient attention. Files load stable→volatile to maximize prompt-cache hits:
    • README.md — the Dashboard (where am I? what is active?)
    • PRD.md — the Strategy (why and what)
    • CLAUDE.md — the Physics (how the AI must behave)
  2. Focus Memoryepics/: the only variable state. An EPIC frames one problem as one context window.
  3. Long-Term MemorySoT/SoT.*.md: the immutable facts. Business Rules (BR-), User Journeys (UJ-), API Contracts (API-), and 18 more ID types. Nothing duplicated; everything referenced by ID.
  4. Short-Term Memorytemp/: the scratch pad. Files attach to the active EPIC and get harvested to SoT before the EPIC closes.

Just-in-Time Context: instead of dumping documentation into the context window, reference specific IDs (UJ-101, API-002). Fewer input tokens, deeper understanding.


📈 Feature: The Progressive PRD

The pitch: the "One-Shot" — asking AI to build the whole app at once — produces generic code and rapid drift. The Progressive PRD makes that impossible by design.

PRD.md is a gated workflow, not a document. The AI focuses on one stage at a time, and no stage advances until its Definition of Done is met:

VersionNameFocusDefinition of Done (DoD)
v0.1SparkProblem & OutcomesProblem defined, Outcomes measurable, Open Questions list.
v0.2Market DefinitionSegments & ICPSegments sized, "Not For" defined, Business Rules (BR-) created.
v0.3Commercial ModelValue & PricingCompetitors profiled, Pricing model, Monetization rules.
v0.4User JourneysPersonas & FlowsCore journeys mapped (UJ-), Dependencies (API-) noted.
v0.5Red Team ReviewRisks & FeasibilityRisks (Market/Tech) identified, Mitigations linked to tests.
v0.6ArchitectureTechnical StrategyStack selected, API contracts (API-) drafted, ARC- conformance rules, Cost guardrails.
v0.7Build ExecutionImplementation LoopCode tested (TEST-), SoT updated, code traced to specs (Development Graph), Epic loop execution.
v0.8Release & DeploymentOperational ReadinessRunbooks (RUN-), Monitoring (MON-, MON-DRIFT-), Rollback plan, Changelog system, MOPS handoff.
v0.9LaunchGo-to-MarketPositioning (Dunford), Offer (Hormozi), Channels (ORB), Launch metrics (KPI-), Feedback channels (CFD-), Tactical playbooks (AEO, alternatives, outreach, HN/Reddit).
v1.0GrowthMarket AdoptionAdoption stage (ADO-STAGE-), Beachhead (ADO-BEACHHEAD-), Whole product (ADO-WHOLE-), References (ADO-REF-), Continuous discovery, Case studies, Testimonials.

Why gates work: constrained focus prevents the AI from guessing the architecture before it understands the users; deep focus produces meaningful IDs; the result is not just a working product but a desirable one.

The paradox that makes it practical: gates provide focus; the ecosystem provides agility. Because documentation is modular and interlocked, you can revisit any stage just-in-time — customer feedback during Build doesn't restart the plan, it updates the BR- rules and lets hooks propagate the change.


🛠 Feature: 47 skills, one for every decision

The pitch: the lifecycle isn't advice — it's executable. Every stage ships with skills that know what to consume, what IDs to produce, and which gate they feed.

  • 41 stage skills (prd-v01-*prd-v10-*): problem framing, competitive landscape, pricing, persona definition, journey mapping, risk discovery, architecture design, epic scoping, test planning, release planning, GTM strategy, case studies…
  • 6 methodology operators (ghm-*): gate checks, SoT building, ID registration, insight harvesting, status sync.
  • Named frameworks, encoded: April Dunford positioning, Alex Hormozi offer construction, Owned/Rented/Borrowed channel allocation, Geoffrey Moore chasm crossing, Teresa Torres continuous discovery, Rob Fitzpatrick Mom Test interviews.
  • Three depth modesquick (founder gut-check, <15 min), standard (default), deep (investor-ready, with assumption logs) — so the method scales from solo founder to team.

Every skill emits Consumes / Produces sections in SoT IDs, which is what keeps the knowledge graph connected as you move through stages.


📊 Feature: A repo that scores its own readiness

The pitch: before advancing a stage or starting an EPIC, the repo already knows whether you're ready — and why not.

Readiness is a three-layer graph over the artifacts you already author:

  1. SoT files — scored on entry count, depth, cross-reference density, and orphan rate. A placeholder file scores 0.
  2. EPICs — inherit the readiness of every SoT file they reference. Dangling refs and unresolved assumptions surface as unmet criteria, each citing the file that caused it.
  3. PRD stage — answers "can we advance v0.X → v0.Y?" from gate criteria plus the layers below.

All three write to one file — status/readiness.json — with causal links intact: an EPIC's unmet criterion points at its caused_by SoT file; the top blockers are ranked by downstream impact. The highest-leverage fix is rarely the lowest-scoring file — it's the lowest-scoring file blocking the most EPICs. The system tells you which.

python scripts/readiness.py run        # compute all layers + print report
python scripts/readiness.py status     # print last-computed report
python scripts/readiness.py run --json # machine-readable output for hooks/CI

Exit codes 0/1/2 map to PASS / WARN / BLOCK (thresholds: warn=70, block=50, overridable per item). The ghm-gate-check skill delegates here for stage-advancement decisions.

🫀 Feature: Code that traces back to specs

Once building starts (v0.7), the code itself joins the knowledge graph. An AST pass extracts code nodes into status/devgraph.json; the @implements / @verifies tags you write under rule 04 become bridge edges linking each code unit to the spec it realizes. Readiness then measures reality, not just spec health:

  • implementation_coverage — which scoped specs actually have implementing code
  • architecture_conformance — do the ARC- rules still hold in the as-built system (drift = a violate verdict, not a surprise in review)

Untagged code shows up as an orphan node — a context leak you can see. The same devgraph.json powers the HeartBeat visualizer: a live pulse of built / unbuilt / drifted. See docs/DEVELOPMENT_GRAPH.md.

Deeper reading: .claude/rules/07-readiness-protocol.md · docs/READINESS_PROTOCOL.md


📰 Feature: The Human Review Layer

The pitch: markdown SoT files are optimized for agents and diffs. Humans reviewing a gate deserve a better reading surface — so every SoT file ships with a styled, hyperlinked HTML view in the format its natural reviewer already expects. Start at SoT/html/index.html (opens from file://, no build step, no JS).

The contract: markdown stays authoritative; the HTML is a render. Entry anchors equal unique IDs (SoT.BUSINESS_RULES.html#BR-001), and every cross-reference is a hyperlink — a reviewer walks the knowledge graph by clicking, the same way an agent walks it by ID.

The Atlas — index of all SoT viewsUser journey rendered as a journey map
The Atlas (index.html) — registry of every view, ID anatomy, graph patternsUser Journeys — trigger → steps → value moment, the way design reviews read flows
API contract rendered as an API referenceData model rendered as a schema browser
API Contracts — Swagger-style reference with method plates and status codesData Model — ER-style entity cards with keys and a relationship map
Customer feedback rendered as an insight cardAdoption stage rendered as the Moore curve
Customer Feedback — quote-first insight cards with decision stampsAdoption — Moore lifecycle curve with the chasm and a "you are here" marker

Each of the 13 pages serves a different reviewer: policy register for BR-, ADRs + topology diagram for TECH-/ARC-, Storybook-style specimens for DES-, Given/When/Then cards for TEST-, an ops console for DEP-/RUN-/MON-/SEC-, a retro playbook for LL-, a vendor context map for INT-. The full schema-per-ID-type and persona-per-view rationale lives in SoT/html/README.md.

Screenshots are generated — when the pages change, refresh them with python3 SoT/html/screenshot.py (Playwright + Chromium; see SoT/html/README.md for setup) and commit the regenerated PNGs with the change.

Next direction (concept): these pages render SoT outward for review. A proposed third artifact class — deliverables — would add an input mode where a human contributes judgment (rank, select, acknowledge) and the page emits paste-ready SoT markdown. See docs/DELIVERABLES_CONCEPT.md.


🤖 Feature: An agent squad with persistent memory

The pitch: four role agents that remember, coordinate through files, and get smarter every EPIC — no standups required.

  • horizon (Strategy, v0.1–v0.5) · studio (Design, v0.3–v0.6) · devlab (Build, v0.6–v0.8) · metro (Ops, v0.9–v1.0)
  • Memory that persists: each agent accumulates Feedback, Patterns, Decisions, and Handoff Notes in its MEMORY.md. A SubagentStop hook actively extracts memories from the conversation. During EPIC harvest, cross-EPIC insights are promoted to SoT/SoT.LESSONS_LEARNED.md as durable LL- entries.
  • Event-driven hooks instead of meetings: SessionStart injects read order, UserPromptSubmit checks context density, PreToolUse verifies an active EPIC before code writes, Stop reminds on SoT cascade updates. Behavior is standardized by HOOK_CONTRACT.md.
  • Multi-agent EPICs without the telephone game: a Synthesis Checkpoint forces the coordinator to produce self-contained worker prompts before implementation begins — workers never see degraded second-hand context.
  • File-based standups: the Squad Status section below shows agent activity and EPIC state at a glance, updated by ghm-status-sync.

Repository structure

/
├── README.md               # Dashboard, structure, and status
├── PRD.md                  # Product definition (Progressive PRD)
├── CLAUDE.md               # The agent's operating instructions
├── epics/                  # Active Context Windows (Tasks)
├── SoT/                     # Shared Memory Store (SoT.* files + html/ review layer)
├── temp/                    # Scratch Pad for explorations and audits
└── .claude/                 # Methodology runtime (skills, hooks, agents)
    ├── skills/              # 41 stage skills (prd-v*) + 6 operators (ghm-*)
    ├── hooks/               # Session/user/stop hooks + subagent memory hooks
    ├── agents/              # Role agents with persistent MEMORY.md
    ├── domain-profile.yaml  # ID registry + skill taxonomy
    └── settings.json        # Hook wiring and execution config

Agent Note: .claude/ can be replaced with .gemini/, .codex/, or any other agent structure, but the skills, hooks, and agent model here were built with Anthropic's documentation model in mind.

Fork Note: this README.md explains the methodology. When you fork for a product, copy README_template.md to README.md and customize it.


🚀 Quick Start

# 1. Fork this repo for your product, then:
cp README_template.md README.md     # your product dashboard replaces this page

# 2. Open the repo in Claude Code — hooks load the read order automatically.

# 3. Start the lifecycle at the beginning:
#    "Let's frame the problem"  →  triggers prd-v01-problem-framing
#    The skill produces CFD- evidence IDs and fills PRD.md v0.1.

# 4. Advance only through gates:
python scripts/readiness.py run      # are we ready for v0.2?

From there, the method drives itself: each stage's skills consume the previous stage's IDs, the readiness score tells you when to advance, and the knowledge graph grows with every decision. No subscriptions, no servers, no lock-in — fork and go.


How to use this repo today

The methodology is fork-native: everything runs from files in your repo, with no services to stand up.

  1. Fork it per product — this repo is the template; each product gets its own copy with its own PRD, SoT graph, and EPICs (Quick Start above).
  2. Or self-install into an existing repo — run the self-install path instead of forking the whole repo (see below).
  3. Work through the gates — let the stage skills drive: each one tells you what it consumes, what IDs it produces, and which gate it feeds. Run python scripts/readiness.py run before advancing.
  4. Keep the graph honest — decisions land in SoT/ with IDs before or during the change, never after. Review them as humans through SoT/html/.
  5. Harvest every EPICtemp/ notes and agent memories get promoted to durable LL- entries at EPIC close, so the next session starts smarter.

Adopt into an existing repo (self-install — prototype)

You don't have to fork. The self-install path drops the framework into a fresh or existing repo without clobbering product content — the subscription-native pattern borrowed from ZQadus/Xantham-system-blueprint: ship a blueprint a fresh Claude Code session executes, so all cost lands on your Pro/Max plan, not the metered API.

# Deterministic CLI (from a clone of this repo):
bash install.sh --target /path/to/your/repo --profile product --dry-run   # preview
bash install.sh --target /path/to/your/repo --profile product             # install

Or paste the one-line bootstrap from BLUEPRINT.md into a fresh Claude Code session and let the ghm-self-install wizard drive it. Both paths read .claude/install-manifest.yaml (framework vs. product file classes), are idempotent, and merge into an existing .claude/settings.json rather than overwriting it.

Still on the roadmap

The self-install path is the foundation — fuller distribution is next:

  • MCP server — the knowledge graph as a queryable service: look up any ID, traverse cross-references, and pull readiness scores from any MCP-capable agent, without loading files into context.
  • Packaged Claude Code plugin — the skills, hooks, and agent squad as a marketplace one-click, plus Xantham-style hardening (Docker audit sandbox, signed CHECKSUMS.sha256).

Watch the repo to catch these when they land. The fork and self-install paths both work end-to-end today.


Squad Status

Agent Activity

AgentRoleLast ActiveCurrent EPICStatus
horizonStrategy (v0.1-v0.5)idle
studioDesign (v0.3-v0.6)idle
devlabBuild (v0.6-v0.8)idle
metroOps (v0.9-v1.0)idle

EPIC Status

EPICStateLeadLast Updated
(no active EPICs)

Files in the repo

Repository payload21 top-level entries
  • .claude
  • .claude-plugin
  • .github
  • docs
  • epics
  • plugins
  • scripts
  • SoT
  • temp
  • tests
  • .gitignore
  • BLUEPRINT.md
  • CHANGELOG.md
  • CLAUDE.md
  • install.sh
  • LICENSE
  • MIGRATION.md
  • PRD-Methodology-Overview.pptx
  • PRD.md
  • README_template.md
  • README.md

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More plugins

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

138k
1 add

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.

82k
code-yeongyu/
oh-my-openagent

OmO: Just type "mass ulw" keyword with your prompt. Now you are the master of graph engineering.

69k

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

94k

Opinionated Oxlint rules for rejecting low-evidence TypeScript and JavaScript patterns

4.3k