Sandbox
@withqwerty/football-docs

MCP server for football data provider docs

football-docs builds a searchable index of documentation for football data providers and tools, then serves it to coding agents over MCP. The server supports search, provider lookup, comparisons, and entity resolution, with provenance attached to results so you can trace each answer back to a source.

66 stars9 forksTypeScriptUpdated 7d ago
Who it's for

Builders who need their agent to look up football data docs instead of inventing API details.

What it delivers

You can ask your agent about football data providers and get sourced answers from real documentation.

What it does

Search provider docs

Searches across indexed docs with full-text queries and optional provider filters.

Resolve provider names

Maps aliases like Opta, Wyscout, or StatsBomb to canonical provider keys before searching.

Compare providers

Shows how different providers handle the same concept, such as coordinate systems or event types.

Resolve entities

Resolves players, teams, or coaches to cross-provider IDs through the Reep API.

Track provenance

Returns source URLs, versions, and crawl metadata with indexed results.

Request updates

Lets agents queue missing-doc and stale-doc requests through the MCP server.

How to get it

  1. 1Run
    claude mcp add football-docs -- npx -y football-docs

README

football-docs

Searchable football data provider and tooling documentation for AI coding agents. Like Context7 for football data.

Who it's for: Developers and analysts who use AI coding tools (Claude Code, Cursor, VS Code Copilot) to work with football data. Works with any tool that supports MCP.

What it does: Gives your AI agent a searchable index of documentation for 23 football data providers and tools — event types, qualifier IDs, coordinate systems, API endpoints, data models, identity surfaces, and cross-provider comparisons for the data providers (StatsBomb, Opta, Wyscout, Impect, SkillCorner, Sportradar, TheSportsDB, FMDB Pro, TransferRoom, and more), plus the open-source libraries people build with (kloppy, mplsoccer, socceraction, soccerdata, floodlight, fast-forward, unravelsports, and more). Your agent looks up the real docs instead of guessing from training data.

Why not just let the AI figure it out? LLMs get football data specifics wrong constantly — Opta qualifier IDs, StatsBomb coordinate ranges, API endpoint URLs, library method signatures. These are mutable facts that change across versions. football-docs gives the agent verified, sourced documentation with provenance tracking so you know where every answer came from.

Strategy

football-docs is intended to be a community-owned, source-transparent Context7 for football data. The public operating contract is in STRATEGY.md: what belongs here, what must stay out, how we handle public-safe provider facts, and how contributors should prove retrieval quality.

Provider identity facts

football-docs is the public source for provider identity-surface facts: access shape, ID schemes, matching fields, provider quirks, and provenance rules. Curated provider identity notes belong here when they can be stated as public facts about the provider. They should say whether a fact comes from public docs, public page evidence, licensed feed shape, or a reviewed public-safe observation, and must not include credentials, local paths, internal tooling details, or restricted payloads from any private project.

MCP (Model Context Protocol) is a standard for connecting AI coding tools to external data sources.

Quick start

Claude Code

claude mcp add football-docs -- npx -y football-docs

Cursor

Settings → MCP → Add server. Use this config:

{
  "mcpServers": {
    "football-docs": {
      "command": "npx",
      "args": ["-y", "football-docs"]
    }
  }
}

VS Code / Copilot

Add to .vscode/mcp.json:

{
  "servers": {
    "football-docs": {
      "command": "npx",
      "args": ["-y", "football-docs"]
    }
  }
}

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "football-docs": {
      "command": "npx",
      "args": ["-y", "football-docs"]
    }
  }
}

Tools

ToolDescription
search_docsFull-text search across all provider docs. Filter by provider. Results include provenance (source URL, version).
resolve_provider_idResolve provider names and aliases to canonical indexed provider keys before searching.
get_provider_docsRetrieve docs for a resolved provider, optionally filtered by topic or category.
list_providersList all indexed providers and their doc coverage.
compare_providersCompare how different providers handle the same concept.
request_updateRequest a new provider, flag outdated docs, or suggest a better doc source. Queues locally and points to the matching public GitHub issue template.
resolve_entityResolve players, teams, or coaches to cross-provider IDs via the Reep API.

Provider filters use the indexed provider keys shown by list_providers, but common aliases are accepted. Examples: fbref, understat, ClubElo, football-data.co.uk, and engsoccerdata search free-sources; Sofascore and ESPN search soccerdata; FMDB searches fmdb-pro; Transfer Room searches transferroom; Hudl Wyscout searches wyscout; Stats Perform / Opta F24 / WhoScored search opta; Metrica, Sportec / DFL, and TRACAB search databallpy; Second Spectrum searches kloppy; Hawk-Eye, SciSports, Signality, Respovision, GradientSports and OptaVision search fast-forward; unravel searches unravelsports; SportRadar API / Soccer Extended search sportradar; The Sports DB / TSDB search thesportsdb; StatsBomb Open Data searches statsbomb.

Example queries

  • "What is Opta qualifier 76?" (big chance)
  • "How does StatsBomb represent shot events?"
  • "Compare Opta and Wyscout coordinate systems"
  • "What player ID fields does Transfermarkt expose?"
  • "Does SportMonks have xG data?"
  • "What event types does kloppy map to GenericEvent?"
  • "How does SPADL represent a tackle?"

Indexed providers

ProviderChunksCategories
fast-forward250overview, getting-started, data-model, coordinate-system, orientations, layouts, transformations, distributed-compute, api-reference, 12 provider format pages
StatsBomb235event-types, data-model, coordinate-system, api-access, api-endpoints, charting-lineups, xg-model, iq-metrics, player/team stats, player-mapping, identity-surfaces
unravelsports202overview, installation, quickstart, concepts, graph converters, pressing intensity, formation detection, models, utils, american-football
Wyscout163event-types, data-model, coordinate-system, api-access, api-endpoints, charting-analysis-metrics, glossary, identity-surfaces
kloppy126data-model, usage, provider-mapping, tracking-rendering, event-derived-metrics
floodlight144core data objects, io parsers (Tracab, DFL, Kinexon, Opta, SkillCorner, StatsBomb, StatsPerform, Second Spectrum), transforms, metrics, models, visualisation, guides
SportMonks565full v3 endpoint reference (fixtures, livescores, leagues, seasons, states, types, statistics, brackets), syntax and includes, filtering, rate limits, error codes, changelog, plus curated event-types, data-model, api-access, charting-season-stories, identity-surfaces
databallpy63data-model, overview, usage
mplsoccer65overview, pitch-types, visualizations
Impect77overview, data-model, event-types, coordinate-system, concepts, kpi-definitions, identity-surfaces
SkillCorner49api-access, api-endpoints, data-model, physical-data, coordinate-system, concepts, identity-surfaces
Free sources62overview, fbref, understat, contextual-story-joins, xg-timelines
soccerdata40overview, data-sources, usage
TransferRoom43api-access, api-endpoints, charting-availability, data-model, identity-surfaces
Opta71event-types, qualifiers, coordinate-system, api-access, charting-game-state, charting-lineups, charting-passmaps, charting-set-pieces, charting-shot-placement, identity-surfaces
FMDB Pro35api-access, api-endpoints, data-model, identity-surfaces
Sportradar30api-access, api-endpoints, data-model, charting-and-stories, integration-notes
socceraction34SPADL format, VAEP, Expected Threat
BeSoccer14api-access, api-endpoints
Driblab27api-access, api-endpoints, data-model
TheSportsDB18api-access, api-endpoints, livescore, identity-surfaces
FotMob3identity-surfaces
Soccerdonna3identity-surfaces
Transfermarkt3identity-surfaces

2,322 searchable chunks across 24 providers and tools.

Impect documentation is built solely from the public ImpectAPI/open-data repository — a static Bundesliga 2023/24 snapshot, representative of Impect's structure and metric definitions rather than a complete or current mirror. Impect's commercial API is deliberately not documented here. Every enum value, KPI name and field name in docs/impect/ is validated against that repository in CI (pnpm impect:truth, src/__tests__/impect-open-data-validation.test.ts). Data source: Impect; use is subject to the repository's own Terms of Use.

Documentation validation

Docs for AI agents are only useful if they are correct, and prose about an API is exactly the kind of thing that drifts or gets invented. Where a machine-readable source of truth exists, this repo checks the docs against it in CI rather than trusting them.

ProvidersGround truthChecked by
kloppy, socceraction, soccerdata, mplsoccer, floodlight, databallpy, skillcorner, fast-forward, unravelsportsThe installed package itself — enum members, importable symbols, class constants, Literal parameter vocabulariessrc/__tests__/provider-truth.test.ts
Wyscout, SkillCorner, FMDB Pro, SportradarThe vendor's own publicly published OpenAPI spec — endpoint paths and methodssrc/__tests__/provider-truth.test.ts
BeSoccerThe vendor's published Postman collection — request vocabulary and parameterssrc/__tests__/provider-truth.test.ts
DriblabThe vendor's published API guide — endpoint paths, methods, parameter and field namessrc/__tests__/provider-truth.test.ts
ImpectThe public open-data repositorysrc/__tests__/impect-open-data-validation.test.ts

Truth files live in data/provider-truth/ and are generated, not hand-written:

pnpm provider:truth      # rebuild every package truth file (needs python3.11)
pnpm openapi:truth       # rebuild every spec-derived truth file

The specs those derive from are snapshots of publicly published, unauthenticated vendor documentation. Source URLs, fetch dates and refresh instructions are in specs/README.md. Wyscout's v3 and v4 specifications merge into one truth file, because its docs span both. The v2 legacy specification is no longer mirrored: the docs describe v3 and v4, and no documented fact derives from the legacy surface.

Each package gets its own pinned venv — co-installing them makes pip silently downgrade conflicting versions, which would produce truth that disagrees with the docs. Bump a pin in scripts/gen_all_truth.sh and the matching version in providers.json together, then re-run and fix whatever the tests flag.

A doc that names an enum member or importable symbol which does not exist in the real package fails the build. scripts/gen_openapi_truth.py derives the same kind of facts from a vendor OpenAPI spec, for providers documented that way.

Not every vocabulary is an enum. fast-forward's coordinate systems, orientations and layouts are lowercase strings on Literal-annotated parameters, so the truth files also record what each parameter accepts, and a doc writing coordinates="statsbomb" fails the same way an invented enum member would.

Contributing

Contributions are welcome from everyone. There are three ways to help:

  1. Open an issuerequest a new provider, flag outdated docs, or suggest a better doc source
  2. Use the request_update tool — AI agents can flag outdated or missing docs directly via the MCP server, which queues requests locally and points to the matching public GitHub issue template
  3. Open a PR — fix errors, add new providers, or improve existing docs

You don't need to be an expert. See CONTRIBUTING.md for the full guide.

For maintainers

Crawl pipeline

Provider doc sources are tracked in providers.json. The crawl pipeline discovers the best doc source (llms.txt > ReadTheDocs > GitHub README) and writes markdown with provenance frontmatter.

npm run discover                        # probe sources without crawling
npm run crawl                           # crawl all providers with sources
npm run crawl -- --provider kloppy      # crawl one provider
npm run ingest                          # rebuild search index from docs/
npm run ingest -- --provider kloppy     # re-ingest one provider (incremental)

Each crawled doc carries provenance metadata (source URL, source type, upstream version, crawl timestamp) that is surfaced in search results, so agents can distinguish between curated content and upstream documentation.

Cutting a release

Pushing the tag is the release. .github/workflows/release.yml runs on any v* tag and does the rest: it re-runs the full check suite, creates the GitHub Release, and publishes to npm.

# 1. Bump the version in package.json and server.json (three fields in total).
#    Land it on main through a pull request, as `chore: release vX.Y.Z`.
#
# 2. Tag the merge commit and push the tag.
git tag -a v0.11.0 -m "v0.11.0"
git push origin v0.11.0

Release notes come from the body of the chore: release vX.Y.Z commit, so write that message as the release notes you want readers to see. The workflow reads it from the second parent when the tag sits on a merge commit, strips the commit trailers, and falls back to GitHub's generated notes if the body is empty.

Three properties worth knowing, because each one exists to stop a specific failure:

  • The tag must match package.json. A mismatch fails the job before anything is created or published, so a version can never ship under another version's name.
  • The full suite runs again. A tag can be pushed to any commit, including one that never went through a pull request, so the release path cannot assume CI already passed on that tree.
  • Re-running is safe. An existing Release is left alone and an already-published version is skipped, so a failed job can simply be re-run.

Publishing to npm needs an NPM_TOKEN repository secret holding a granular access token with read and write on football-docs. An automation token satisfies the account's publish 2FA; an interactive npm login does not, which is why npm publish from a laptop stops for a one-time password. Without the secret the workflow still cuts the GitHub Release, then warns in the job summary that npm was skipped rather than failing silently.

Publishes from CI carry npm provenance, so the tarball on npm is attested to this repository and this workflow run.

Note that prepublishOnly runs pnpm build && pnpm ingest, which rebuilds data/docs.db. Publishing by hand therefore leaves that file dirty in the working tree; the content is unchanged, only SQLite's page layout differs, so git checkout data/docs.db clears it.

License

MIT

Files in the repo

Repository payload20 top-level entries
  • .github
  • bin
  • data
  • docs
  • scripts
  • specs
  • src
  • .editorconfig
  • .gitignore
  • AGENTS.md
  • biome.json
  • CONTRIBUTING.md
  • package.json
  • pnpm-lock.yaml
  • pnpm-workspace.yaml
  • providers.json
  • README.md
  • server.json
  • STRATEGY.md
  • tsconfig.json

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

86k

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k