An agentic skills framework & software development methodology that works.
Claude Code skill for structured research pipelines
Deepdive is a research skill for Claude Code that breaks an investigation into named phases: reframing, planning, search, triangulation, synthesis, and verification. It uses source files, claim ledgers, evidence filters, and citation checks to keep the work auditable and reusable.
Builders who want Claude Code to do documented research with source trails, checkpoints, and review gates.
You can turn a vague research question into a cited report you can revisit later without redoing the search.
What it does
Plan-review gate
Shows the reframed question, hypotheses, genre, and channels before search starts so you can approve or edit the plan.
Parallel sub-agent search
Splits research across named channels and sub-agents so multiple angles are explored at once.
Claims ledger and triangulation
Tracks each claim in `claims.csv` and requires multiple sources and roots before a claim is treated as settled.
Evidence and citation verification
Runs evidence filtering and multi-layer citation checks so quotes and claims are matched before synthesis and report writing.
Decision walkthrough
Ends with a guided fork-by-fork review that logs what you decided in `application.md`.
How to get it
- 1Run
git clone https://github.com/Socialpranker/deepdive.git \ ~/.claude/skills/deepdive
- 2"Validate this hypothesis"
# Clone git clone https://github.com/Socialpranker/deepdive.git cd deepdive # Package as .skill bundle zip -r ../deepdive.skill . -x ".*" -x "*.zip" # Upload via Claude.app → Settings → Skills → Add Skill
README
Deepdive
A structured meta-research skill for Claude Code
Stop ad-hoc Googling. Start documented investigation.
Docs · Install · How it works · Contribute
You: investigate the trade-offs between Postgres logical replication and CDC tooling
Claude: ✓ Reframed your question (3 hypotheses) + decision spec: what you'll do with the answer
✓ Picked genre: decision (comparison + validation)
✓ Wrote plan.md (17 sections)
✓ Checked your env: 4 APIs available, 2 fallback to HTML
✓ Launched 4 sub-agents across 12 channels
✓ Saved 23 sources to sources/ with quotes, provenance-checked
✓ Ran adversarial pass (3 counter-arguments + an "execute this" role)
✓ Report ready: research/postgres-replication-vs-cdc/2026-05-21_decision.md
✓ Walked through your decision forks — you picked CDC tooling, logged to application.md
What this is
A Claude Code skill that turns "research this topic" into a 13-phase pipeline with hypothesis testing, parallel sub-agent search, source triangulation, and adversarial review.
The output is a folder you can return to in a month. Every claim traces to a specific source file. The plan documents why you made every choice. No re-research needed.
New here? Start with the Quickstart — install → invoke → first result in ~5 min.
| Without this | With this |
|---|---|
|
One-shot prompt → wall of text Sources lost in chat history No way to detect bias No reuse next time Generic Google results Sources include... (vague) |
17-section Each source = file with verbatim quotes Mandatory adversarial pass + opposition queries Atomic theses in Every claim → |
Install in 30 seconds
For Claude Code (CLI)
git clone https://github.com/Socialpranker/deepdive.git \
~/.claude/skills/deepdive
That's it. Now type any of these in a Claude Code session:
- "Investigate X"
- "Изучи тему"
- "Validate this hypothesis"
For Claude Desktop (Skills enabled)
# Clone
git clone https://github.com/Socialpranker/deepdive.git
cd deepdive
# Package as .skill bundle
zip -r ../deepdive.skill . -x ".*" -x "*.zip"
# Upload via Claude.app → Settings → Skills → Add Skill
For other LLMs (Codex, Gemini, local)
The 13-phase methodology is portable. Load SKILL.md + references/*.md into the LLM's context manually. Skip the sub-agent parts and use separate chat sessions per subtopic.
How it works
The skill runs 13 phases in order:
| Phase | Name | What happens |
|---|
| 1 | Reframing | opus / high | | 2 | Genre & block selection | sonnet / medium | | 3 | Plan | opus / medium | | 3.5 | Capability Discovery | sonnet / low | | 3.7 | Plan-review gate | sonnet / low | | 4 | Search | sonnet / medium | | 5 | Claims-ledger + triangulation | haiku / low | | 5.5 | Evidence filter | sonnet / low | | 5.7 | Wiki reconcile | sonnet / low | | 6 | Synthesis + multi-angle red team | opus / high | | 6.5 | Verify | haiku / low | | 7 | Refresh targets | sonnet / medium | | 8 | Decision walkthrough | opus / high |
Each phase runs on a model matched to its task — Opus where reasoning multiplies (1/3/6), Haiku for the parallel fan-out (4). The skill announces the routing and an estimated cost up front, once.
Every phase is transparent: you see what's happening, you confirm key decisions, and you get a folder you can return to. Before any search fires, the plan-review gate (3.7) shows you the reframing, hypotheses, genre, and channels and lets you approve or edit them — strictness scales with mode (deep waits for an explicit go-ahead, medium is a soft check, shallow skips it). Editing the plan before execution is the single highest-leverage step in the whole pipeline — Gemini Deep Research calls plan review its "biggest lever over output quality," and a wrong plan executed perfectly still produces a wrong report.
Reframing (1) doesn't just restate the question — a router classifies its profile (factual / multi-step / relational / comparative / landscape) and that classification picks the decomposition method: factual questions get flat independent subquestions, multi-step ones ("X given Y") get least-to-most leveling, comparative ones get a shared axis matrix with mandatory opposition queries per candidate. Picking the wrong decomposition for a question's shape is a silent failure mode — the router makes the choice explicit instead of defaulting to "flat parallel" for everything.
Phase 4 (Search) isn't a single pass — it's a bounded loop with three cheap safeguards so it doesn't quietly waste budget or silently give up:
- Cheap goal-check — after each round, a Haiku pass tags every subquestion
met/partial/unmetwith a one-line reason. This is what the expensive Opus evaluation reads instead of re-deriving the gap from scratch, and it's what targets the next round's dispatch. - No-progress circuit breaker — two consecutive rounds that add nothing new to the source pool stop the loop immediately, regardless of remaining budget. The unresolved thread goes to Open Questions instead of burning tokens chasing a dead end.
- Least-to-most decomposition — for layered questions ("X given Y"), subquestions are leveled
L1 → L2instead of dispatched flat in parallel: L1 rounds run first, concrete facts they surface get carried forward, and L2 queries are launched already sharpened by that context. Independent subquestions still run flat.
Scoring (5) doesn't stop at the usual Credibility/Recency/Bias rating — it also flags input-level skepticism: a source that measures its own product, self-reports a benchmark, or is directly disputed by another collected source gets a strict caveat: marker (vendor / self-reported / disputed:sNN) before the claim reaches claims.csv, not after synthesis has already built on it. A claim whose key number carries that marker is capped at confidence: medium (or low for an unresolved dispute) — the same rule shape as primary-first sourcing. Vendor benchmarks are the numbers that most often get quietly repeated as fact; catching them on the way in, not in the red team pass at the end, is the point.
For medium/deep depth, the pipeline runs two more machine-checked passes most one-shot research skips entirely:
- Evidence filter (5.5) — a CRAG-style relevance classifier runs on every (claim, source) pair before synthesis and keeps only the quotes that actually support that specific claim. Dumping every found source into synthesis measurably hurts quality (Search-o1 dropped 33%→24% doing exactly that); this is the fix, not a nice-to-have.
- Faithfulness verification (6.5) — beyond checking that a cited link is alive, the skill checks that the source entails the claim it's attached to (RAGAS/ALCE-style claim⊨quote), and writes
SUPPORTED/PARTIAL/UNSUPPORTEDverdicts to.verify/faithfulness.json. Citation fabrication is common enough industry-wide — the Tow Center found a >60% error rate in AI-generated citations — that checking for it, not just for dead links, is a real differentiator.
None of this is enforced by discipline alone: scripts/validate_phases.py reads a finished run's mode: and checks that every phase mandatory for that mode actually left its file artifact (plan.md, claims.csv, evidence/, .verify/*.json, the dated report, ...). A skipped phase fails the check instead of silently passing — the model can't just claim "done." As of finish-up, this check is a blocker, not a suggestion: the skill won't report a research as done on a red gate, symmetrically to how a report isn't "done" without its verification header. sources.csv itself is now built the same deterministic way — scripts/build_sources_csv.py generates it from sources/NN.md frontmatter (with a --check mode for CI) instead of being assembled by hand each run.
A well-cited report that changes nothing is still a failure — the skill's answer to that is a decision spine running through the whole pipeline, not a bolt-on question at the end:
- Reframing (1) captures a decision spec, not just a topic: the action you'll take, who reads the report and what they do next, and at least one falsifiable if-then fork ("if the research shows X, I do A"). No fork means the research changes nothing — the skill downgrades to shallow "curiosity mode" with an explicit label instead of quietly running a full pipeline for a dead-end doc.
- Sourcing (4-5) tracks provenance, not just source count. A
root:field on every source flags what it's actually retelling — ten articles quoting one press release are one voice, not ten, so triangulation now requires ≥2 distinct roots in addition to ≥3 sources and ≥2 types. A snowball pass also chains citations backward (to the primary study everyone's retelling) and forward (who's citing it since), catching sources no keyword query would surface. - Synthesis (6) produces
memo.md— a one-page decision memo (recommendation, resolved forks, 3 key numbers with sources, the main risk, next actions) designed to be the thing that actually enters your decision process, with the full report as its evidence base. A conditional recommendation mapped to your forks is required; an unconditional "it depends" is not an acceptable ending. - Red team (6) gets a fourth role — the Executor — who plays your own decision-spec consumer and tries to actually act on the report using only its content, flagging every place a hedge phrase or a missing number blocks a real decision.
- Decision walkthrough (8) is mandatory at every depth, including shallow. The report isn't discussed, it's executed: the skill walks you through each fork one at a time ("research showed X [s03][s11] — does fork A fire? your call?") and logs the outcome — decided, blocked on missing data (one bounded gap-search, then honest
blocked), or deferred — toapplication.md. A global ledger tracks whether past research actually led anywhere, and Phase 1 reads it back on your next research to flag a pattern of dead-end runs.
Want to compare models head-to-head? The eval harness scores any run on 6 axes.
What's inside
106 Report Blocks10 categories: FRAME · EXPLAIN · COMPARE · MAP · VALIDATE · ANALYZE · CLOSE · PEOPLE · NUMBERS · CONTEXT Each block has its own template, anti-patterns, and composition rules. |
29 Search ChannelsNamed strategies with query patterns + paywall fallbacks:
|
460+ Stat Sources14 cross-industry + 19 industry categories. Each entry: URL · Type · Access · Quality · Limitations · Combine-with · Fallback. Categories: Plus a signal registry: 1072 endpoints verified live and filed by what question they answer — parliamentary APIs, advisory feeds, climate and macro endpoints. Loaded one file at a time, outside the base context budget. | ||||||||||||||
6 Report Genres
|
47+ API EndpointsFree no-auth APIs prioritized:
Patents and grants are first-class categories: Auth via env vars only — skill never asks for keys inline. |
Weekly Auto-SyncGitHub Actions cron validates all endpoints + discovers upstream additions:
| ||||||||||||||
Model RoutingPer-phase model selection — quality where it multiplies, cheap where it parallelizes:
~$2 instead of ~$8 on a deep run, and higher quality on critical phases. Override with |
Eval HarnessCompare research quality across models. Same question, different configs, scored on 6 axes:
Weighted sum with a citation floor — hallucinated sources can't win on depth. Verdict = quality per dollar. |
Citation Check
Ignores env proxies ( Verification runs four layers: liveness (does the source exist), faithfulness (does it actually entail the claim it's cited for), qualifier preservation (does the report still say what the ledger said), and construct provenance (do the frameworks, taxonomies and named "laws" the report uses exist outside it). Verdicts land in The fourth layer exists because the first three all join on Numbers get two independent passes: |
Phase-gate Validator
A skipped phase fails the check ( It also checks three things a missing-file test cannot: that Its own inputs are machine-built too: |
Example folder
Sample output for a typical decision-genre research:
research/<topic-slug>/
├── plan.md # 17-section plan
├── state.md # Round window, rewritten each round
├── sources.csv # Index with C/R/B scoring
├── claims.csv # Claim ledger + triangulation status
├── numbers.csv # Every figure the report stands on
├── outline.md # section → block → claim_id map
├── sources/ # One file per source
│ ├── 01_vendor-docs.md # Primary, total=14
│ ├── 02_benchmark-paper.md # Academic, total=12
│ ├── 03_industry-report.md # Industry, total=13
│ ├── 04_forum-thread.md # Forum, total=9 (opposition)
│ └── ... (19 more)
├── findings/
│ ├── F1_<atomic-thesis>.md # confidence: high
│ └── F2_<atomic-thesis>.md # confidence: medium
├── .verify/ # One producer per file, many consumers
│ ├── authority.json # Phase 5.5 — who may assert this
│ ├── citations.json # Layer 1 — liveness
│ ├── faithfulness.json # Layer 2 — entailment
│ ├── qualifiers.json # Layer 3 — scope drift
│ └── constructs.json # Layer 4 — named-construct provenance
├── memo.md # One-page decision memo (always)
├── application.md # Decision walkthrough verdict (always)
└── 2026-05-21_decision.md # Final report
Final report structure (assembled from the blocks chosen in plan.md):
## TL;DR
- Claim A holds under condition X [confidence: high]
- Claim B holds conditionally on threshold Y [confidence: medium]
- Claim C is disputed by opposition sources [confidence: low]
## Mental model
[How the underlying mechanism works...]
## Falsification criteria
What would disprove H1, H2, H3...
## Verdict conditional
Recommendation IF: <conditions met>
Different recommendation OTHERWISE: <conditions broken>
## Counter-arguments (steel-man)
CA1: "<the strongest opposing claim>" [source: s09]
→ Our answer: <conditions under which CA1 fails>
CA2: ...
Every claim is clickable to its source. A month later, you don't re-research — you read.
Contribute
The catalog is most valuable when it grows. Easy contributions:
| Time | Type | Example |
|---|---|---|
| 15 min | Add a stat source | Add SimilarWeb Pro to consumer_digital |
| 15 min | Improve a query pattern | Better arxiv channel queries for biology |
| 30 min | New search channel | Add patent-search with USPTO+EPO fallback |
| 1-2h | New industry category | Add industries/aerospace.md |
| 2-4h | New report block | Add decision-tree to compare.md |
| Half-day | LLM adapter | Add codex/ folder with adapted protocols |
FAQ
How is this different from ChatGPT Deep Research / Perplexity?
Those are products — closed UI, fixed flow, opaque source selection. This is open methodology — you control every step, the protocol is markdown you can fork, the source catalog is yours to extend.
They also don't separate sources into files, don't do explicit triangulation, don't run adversarial passes, and don't produce reusable atomic theses. Nor do they filter evidence for relevance before synthesis (feeding a model everything you found measurably hurts quality — Search-o1 dropped from 33% to 24% accuracy doing that) or verify that a cited source actually supports the claim it's attached to, rather than just existing (faithfulness, not just liveness). Citation fabrication is common enough industry-wide — the Tow Center found a >60% error rate in AI-generated citations — that checking for it is a real differentiator, not a nice-to-have.
Honesty about sources goes further than checking they exist: scoring flags a source that's measuring its own product, self-reporting a benchmark, or directly disputed by another collected source, and caps the confidence of any claim resting on that number — before it ever reaches the report. Vendor benchmarks getting quietly repeated as fact is a market-wide problem; catching it on input, not as an afterthought, is the same honesty principle as faithfulness applied one step earlier.
Does it work without Claude Code CLI?
Yes — on Claude Desktop with Skills enabled. Also works manually with any LLM by loading the markdown files into context (see "Use with other LLMs" below).
What's a research output look like?
See the example folder above. TL;DR: a folder with plan.md + sources/NN.md per source + findings/FN.md atomic theses + final <date>_<genre>.md report.
Every claim in the final report links to a specific sources/NN.md file.
Why so many files? Isn't this overkill?
For a 5-minute "what's the latest X" question — yes. That's wh
Files in the repo
- .claude-plugin
- .github
- assets
- docs
- eval
- references
- runner
- scripts
- tests
- .gitignore
- CODE_OF_CONDUCT.md
- CONTRIBUTING.md
- LICENSE
- phases.yaml
- pytest.ini
- QUICKSTART.md
- README.md
- SKILL.md
Discussion (0)
Ask about usage, or say what you built with itSign in to join the discussion.
No comments yet. Be the first to say what this is good for.
More skills

Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.
Topic in, narrated explainer video out. A Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar. Chinese or English; every frame drawn in code with Remotion.
Public repository for Agent Skills
Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)

Production-grade engineering skills for AI coding agents.