Methodology: from raw records to this site

This site is not a republished spreadsheet. It combines four differently structured data systems, resolves what can safely be counted or mapped, classifies hundreds of thousands of records, prepares plain-English content, and then subjects the result to automated and human-directed review.

The aim of this page is to make those choices inspectable. It distinguishes source facts from our derived fields, explains where AI is used, records the important manual corrections, and is candid about the tests that passed, the tests that did not, and the uncertainty that remains.

Across navigation and search, grants is the umbrella term for the entries on this site. The interface then distinguishes funded awards from NIHR-supported projects. The latter are NIHR infrastructure-support relationships, not additional awards, and never contribute to funding totals.

Current published site

269,841 searchable records in the site release labelled 16 July 2026

Source retrieval dates differ; this is not a live feed
227,739 funded award records
42,102 NIHR infrastructure-support relationships
£97.6bn sum of positive published award values
29 topics grouped into 6 broad fields
The complete route How a source record becomes a public page
  1. 1 Collect Separate UKRI, NIHR, Europe PMC and CORDIS source snapshots
  2. 2 Harmonise Join source tables and map them to one common structure
  3. 3 Resolve Handle overlap, record type, recipient and geography
  4. 4 Represent meaning Turn each title and abstract into a semantic vector
  5. 5 Categorise Compare with curated examples and preserve reviewed decisions
  6. 6 Explain Prepare safe lay summaries offline
  7. 7 Challenge Run tests, review packs, audits and release-blocking gates
  8. 8 Publish Build static pages, search files, maps and downloads

Grant summaries are generated before publication. Opt-in semantic search is the one live AI exception, described below.

Source Original facts remain visible

Titles, abstracts, funder references and source links remain the authority.

Derived Interpretation is labelled

Topics, status, geography and summaries are explicitly derived fields.

Review Corrections are durable

Reviewed decisions are stored as overlays, not hidden edits to source records.

Caution Missing stays missing

We do not estimate an unpublished award value or invent a recipient or abstract.

01

Inputs

Sources and scope

Source acquisition happens in a separate grant-data workspace. Each source is collected and processed on its own terms before this public-site pipeline receives harmonised project and researcher files. That separation matters: it preserves provenance and avoids pretending that four very different databases are one uniform dataset.

UKRI Gateway to Research

Open source site ↗

Four separate feeds—projects, funding, people and organisations—are joined on UKRI project identifiers. Principal investigators, co-investigators, fellows, supervisors and students become linked researcher rows rather than being flattened into one field.

Public records
173,320
Current raw download
18 February 2026
Licence
OGL v3

NIHR Open Data

Open data portal ↗

We combine the funded portfolio, Research Schools projects and the infrastructure-supported projects dataset. The last of these records support relationships, so it is retained for discovery but separated from financial awards.

Funded awards
12,480
Support relationships
42,102
Current retrieval
18–19 June 2026

Europe PMC GRIST

Grant finder ↗

Europe PMC aggregates grants from medical and research charities. Its rows are often organised around researchers and affiliations rather than a definitive recipient, so project deduplication and recipient attribution require separate treatment.

Public charity records
27,395
Current retrieval
19 June 2026
Important caveat
Affiliation is not always recipient

European Commission CORDIS

CORDIS datasets ↗

Monthly Horizon 2020 and Horizon Europe project, organisation, topic and EuroSciVoc files are joined by CORDIS project ID. A project enters this site when at least one participant has an explicit UK or GB country code. It is counted once, while every UK participant remains visible.

UK-involved projects
15,046
Funding measure
Known UK participant contributions
Important caveat
Partial totals are lower bounds

What is included

The funded-award views cover records with a usable start year of 2006 or later. Records without a usable start year do not enter the current public award corpus. The coverage is substantial but not exhaustive of every UK funder: it is the union of the sources above after the decisions described on this page.

02

Harmonisation

Preparing the records

A shared public schema makes search and comparison possible, but the pipeline keeps enough source information to avoid unsafe joins. The preparation stage is a sequence of explicit decisions, not a blanket “clean data” command.

  1. 1

    Preserve source-aware identity

    Every public identifier receives a source prefix such as UKRI, NIHR or EPMC. Researcher rows join on both source and raw grant ID. This prevents the 1,428 raw identifiers that currently occur in more than one source from colliding.

  2. 2

    Map to common fields

    Titles, abstracts, funders, dates, amounts, currencies, organisations and people are mapped into common columns. Amount strings are parsed, researcher roles are unfolded, and obvious duplicate display rows are removed without collapsing different roles.

  3. 3

    Reduce known overlap—without claiming perfect deduplication

    Europe PMC produces one project row per grant ID while retaining its researcher-affiliation rows. Records from a maintained list of UKRI and NIHR funders are excluded from the Europe PMC charity feed. Spelling variants, transfers and less obvious cross-source overlaps can remain.

  4. 4

    Separate awards from support relationships

    NIHR infrastructure-supported projects remain searchable, but they do not inflate the funded award count, award-value total, funding trends or institution funding rankings. The rule uses the source dataset, not keywords in a title.

  5. 5

    Derive status and public totals

    “Upcoming”, “active”, “completed” and “unknown” are calculated from published start and end dates for every public export. The £97.6bn headline is the sum of positive published award-record values. It is not annual spend, cash expenditure or an inflation-adjusted figure.

Decision

Zero is not treated as a free grant

A missing or non-positive amount is displayed as “not disclosed”. In the current funded-award corpus, 64,671 records—28.4%—lack a positive published value.

Limitation

Status is a snapshot

Status is recalculated when a public release is exported, not live in the browser. It can lag the source dates between releases, and invalid source date combinations are not silently rewritten.

03

Attribution

Recipients, affiliations and geography

“Who received the award?” and “which organisations are connected to it?” are not always the same question. UKRI and NIHR normally identify a lead organisation. Europe PMC often provides several researchers’ affiliations without naming a grant-level recipient, so we deliberately avoid assigning the whole award to whichever institution happens to appear first.

Recipient decision flow
Source Explicit lead recipient?
YesUse the source recipient, unless a verified override applies
NoExamine the canonical Europe PMC organisations
Europe PMC Exactly one organisation?
YesUse it as the available recipient signal
NoShow associations only; do not allocate funding or map it

This produces 25,764 Europe PMC records with an attributable recipient and 1,631 association-only records in the current release. Association-only values remain in funder totals, but not in institution or regional allocations. Twenty-two independently checked exceptions are held as evidence-backed export overrides; source records remain untouched.

Mapping rules

  • The map represents where a recipient organisation is recorded as based, not the exact project site or department.
  • Europe PMC’s source country is not accepted as proof of recipient location.
  • Locations resolve through a promoted registry: manual overrides first; NIHR source coordinates; unique exact Companies House registered-office sectors for company-like recipients; UKRI postcodes; then exact ROR matches and other reviewed fallbacks.
  • A word such as “Cambridge” or “Oxford” in an organisation name is never treated as location evidence.
  • Companies House points show a registered-office postcode sector, which may be an accountant or legal office rather than a research site. ROR points normally represent an organisation’s recorded city.
  • Unresolved or competing affiliations stay searchable but are left off UK maps and region rankings.
04

Semantic classification

How grants are categorised

The public taxonomy contains 29 topics in 6 fields. Health topics are broadly informed by health-research classification practice; other topics draw on established research-field structures, then use plain labels intended for non-specialists. “Infrastructure” is a record facet rather than a topic, because a facility can support cancer, physics, computing or any other substantive field.

One classification, several layers of evidence
Source text Original title + full usable abstract where available Title alone otherwise; very long abstracts are split and recombined
Meaning 256-dimensional semantic vector OpenAI text-embedding-3-small
Comparison 2,900 curated exemplars 100 grants for each of 29 topics
50% mean similarity to the five nearest topic examples
30% similarity to the topic’s centre
20% similarity to the single closest example
Result primary topic, runner-up and ambiguity margin

Choosing the examples

The 100 exemplars per topic are not simply the largest awards or the records a first-pass model liked most. Candidate selection excludes title-only and infrastructure records, spreads examples across semantic sub-groups, and considers confidence, recency, previous review, funder and institution diversity, and duplicate titles. A separate set of 50 records per topic—1,450 in total—is held out from the exemplar pool for evaluation.

A versioned curation layer currently records 218 forced anchors, 125 forced held-out examples and 6,495 exclusions from anchor selection. These are structured human-directed, Codex-assisted review decisions, not an independently expert-labelled disciplinary gold standard.

Stability between releases

Once approved, a grant’s topic is stored against its source-aware ID. Routine refreshes reuse historic decisions and classify only new IDs. The registry also fingerprints the embedding model, dimensions, input fields, token limit and long-text pooling recipe. A normal export stops if the vectors and the category registry were built under different recipes.

The EC incremental category gate

New CORDIS IDs are embedded in a scratch release directory under the unchanged title-plus-full-objective contract and classified against the pinned exemplars. Historic vectors, assignments and exemplar IDs are fingerprinted and must remain unchanged. EuroSciVoc agreement is reported as a diagnostic only because it is a different taxonomy. Publication requires a prediction-blind, independently labelled EC held-out set with at least 10 projects in every public topic and at least 290 projects overall, at least 93% top-three accuracy, no more than 12% wrong-confident assignments, and a fingerprint-matching decision for every targeted conflict. No QA waiver is available for this release.

A genuine method change follows a separate migration route: produce a shadow classification, compare every proposed move with the locked labels, prioritise high-value, held-out, low-margin and cross-domain changes, then require an explicit, fingerprint-matching promotion decision. Durable manual moves remain authoritative after semantic scoring.

Current validation snapshot Useful for exploration, not authoritative disciplinary coding

The 10 July 2026 full-abstract migration changed 49,752 of 260,337 source-level assignments (19.1%). Leave-one-out testing scored 97.7% top-three accuracy, but the harder held-out set scored 76.8%. The confidently-wrong rate was 21.5%, above the intended 12% ceiling. No rows in the 50,666-row migration queue were recorded as individually reviewed.

Promotion therefore did not represent a clean pass of every QA gate. It used an explicit, dated waiver after side-by-side map inspection, recording the failed gates and accepted conflicts. This is why the site presents topics as navigational aids, preserves a runner-up topic, and warns that interdisciplinary and title-only records are the hardest cases.

05

Plain English

How lay summaries are prepared

The original funder title is retained; the site does not generate replacement “catchier” titles. For a subset of records with substantive source abstracts, DeepSeek V3 has prepared an accepted 150–200 word explanation. The prompt asks for a concrete opening, the problem or knowledge gap, what the work will do, and what could change if it succeeds—without forcing a daily-life application onto fundamental research.

93,766 records · 34.7% AI plain-English summary

Generated from a substantive source abstract, labelled as AI and passed through safety checks.

118,022 records Original abstract fallback

The source abstract is shown when no accepted AI summary is available.

58,053 records Title-only neutral fallback

No speculative research description is generated when the source provides too little evidence.

Counts refer to the 269,841 searchable records in the current served snapshot.

Editorial decisions in the prompt

  • Use plain, active language for an intelligent reader who is not a specialist.
  • Explain jargon rather than merely replacing it with another technical term.
  • Include a number only when the exact number appears in the source abstract.
  • Do not infer mechanisms, patient groups, locations, applications or delivery routes.
  • Avoid fictional characters, melodrama, anthropomorphism and extended metaphors.
  • State potential impact as a possibility where the source does not establish an outcome.

Model and prompt testing

DeepSeek V3 and Kimi K2 were first compared side by side. The prompt then went through versions 2–5 on a fixed 75-grant test set, including records with and without abstracts, before the final version was run again with DeepSeek. Review found the prose generally faithful and readable, but identified unsupported numbers and confident title-only inference as the highest-risk behaviours. The production decision was therefore to require a usable abstract and to add stricter source-only rules.

Making the work durable

Accepted summaries live in a sidecar dataset keyed by the source-aware grant ID. Enrichment merges that store before export, so a routine rebuild cannot erase paid generation work or replace it with an older fallback. Grant pages label the text and retain the source abstract, plus a link to the original funder record where the source publishes one.

06

Exceptions

Where manual correction was necessary

Automation handles volume; it does not make unusual records disappear. Manual and curated decisions are stored in versioned overlays and applied at the relevant curation or export stage, so source data is not silently rewritten. Category overlays retain the grant ID, action and—where moved—target topic. Recipient corrections additionally retain a reason and evidence URL. Provenance depth varies by correction workflow.

Problem found How it was found Durable response
Category boundary errors
Commercial technology, generic clinical trials and interdisciplinary records crossed topic boundaries.
Representative, newest, high-value and targeted boundary review packs; the Semantic Map used only to locate suspicious neighbourhoods. 1,478 explicit category moves and 62 “do not feature as representative” decisions currently override automated output.
Contaminated exemplars
Removing a bad anchor could allow another unreviewed record to refill the slot.
Direct inspection of the actual 100-record exemplar sets and the exemplar-builder refill route. Forced anchors, held-out cases, exclusions and topic-specific guards make review decisions persistent.
Thin and placeholder abstracts
AI filled evidence gaps with plausible but unsupported detail.
Placeholder scan, number audit and an attempted census of 5,289 summaries based on very short abstracts; 5,223 reviews completed. 2,029 IDs are gated in total, including 66 uncompleted audit cases held back as a safe default; shared hygiene prevents re-entry.
Europe PMC recipient ambiguity
The first affiliation was sometimes mistaken for the award holder.
Top-value recipient review against local affiliation evidence and authoritative funder or recipient pages where needed. 22 evidence-backed recipient overrides; multi-affiliation records remain association-only when unresolved.
NIHR support records counted as awards
Centre-support relationships inflated grant and institution totals.
Trace of the original NIHR infrastructure dataset and its record meaning. 42,102 relationships stay searchable under a separate record type and are excluded from award totals.
Unsafe or incomplete geography
A funder/currency country signal did not prove a UK recipient location.
Known-institution checks, overseas-marker tests and review of high-value unmatched names. Source-aware location registry with exact identifiers or names, provenance, conflict review and no city-token inference; unresolved organisations are not mapped.
07

Quality assurance

Testing and release gates

Testing is layered because no single accuracy number can validate ingestion, attribution, topic assignments, prose and a static website at once. Automated suites are run alongside review packs and release gates, while the known gaps and failed category gates are reported here rather than hidden behind a test total.

Automated

Record and export invariants

Tests cover source-aware joins, duplicate researcher display rows, award/support separation, recipient versus association-only logic, map exclusion, missingness calculations, organisation aggregation, clean directory rebuilds and safe grant URLs.

Automated

Category mechanics

Tests enforce exactly 100 unique exemplars and 50 non-overlapping held-out cases per topic, vector alignment, score construction, ambiguity flags, durable overrides, registry fingerprints and migration comparisons.

Review

Category quality

Leave-one-out and held-out tests, transition matrices, cross-domain change counts, high-value queues, boundary packs and side-by-side map views expose where classifications move. The present held-out result remains below the intended gate.

Review

Summary fidelity

Model comparison, five prompt generations, manual sample review, exact-number scanning, thin-text census, placeholder/refusal detection and idempotence tests focus on whether prose remains grounded in the source.

Independent checks

Public build and files

The Astro production build and search/Semantic Map component tests run separately from pipeline export. Export-time geography assertions are blocking; downloads are rebuilt from cleaned public chunks, although download-generation failure is reported rather than blocking the rest of an export.

The 100-largest-awards gate

Every normal export compares the proposed 100 largest displayed awards with the last committed snapshot. A new or materially changed amount, recipient, affiliation set or geography blocks release until a fingerprint-matching decision is recorded. The gate watches change; it does not mean all 100 baseline rows have been independently reverified.

The category migration gate

A change to the semantic recipe must produce a shadow report, a review queue and a decision tied to the exact report fingerprint. Promotion normally requires all absolute checks and conflicts to pass. The code also permits an explicit QA waiver with an approval flag and a non-empty reason. The current waiver goes further by recording the failed gates, accepted conflicts, reviewer, date and reason; no waiver was implied merely by proceeding with an export.

The EC consortium and category gate

The EC release additionally checks one public row per CORDIS project, exact equality between project-level known UK funding and the participant projection, explicit lower-bound labels, unchanged historic category rows and vector contract, balanced blind held-out accuracy, and completed targeted review. Unlike a general category migration, this EC release has no waiver route. Refreshed upstream data and shadow artefacts are retained when a gate fails, but public files are not installed.

08

Interpretation

Known limitations

01

Coverage is broad, not complete

The site covers four public systems, not every UK research funder or every award made outside those systems.

02

Source freshness is uneven

Retrieval dates differ and publishers can amend older records. The website is not a live or append-only register.

03

Published values are incomplete

Missing values are not estimated. Totals are not inflation-adjusted expenditure and can still be affected by unresolved cross-source overlap.

04

Recipient and place are not always knowable

A researcher affiliation is not necessarily an award holder. Map points may be source coordinates, postcode centroids, a ROR city or a Companies House registered-office area; none proves the location of every research activity.

05

One primary topic simplifies interdisciplinary work

A runner-up helps, but 29 topics cannot express every discipline, method or application. Current held-out validation shows substantial residual error.

06

AI prose can still be wrong

Safety gates reduce risk; they do not prove every sentence. The source abstract and original record remain authoritative.

07

Absence is not evidence of no research

A missing abstract, value, recipient, person or map point often reflects publishing practice rather than the importance or existence of the work.

09

Stewardship

Updates, reuse and contact

We aim to refresh source data roughly quarterly. A refresh can change old records as well as add new ones, so historic IDs and reviewed overlays are retained. Topic assignments and summaries are reused unless a new record, explicit correction or controlled migration requires change.

The current process does not yet emit a public release manifest containing a checksum and retrieval date for every source. Adding that manifest is a planned transparency improvement. Until then, the source-specific dates on this page are more informative than a single “as of” date.

Download the public data

Downloads are built from the cleaned grant chunks served by the site, so the downloadable view follows the same record-type, summary and attribution decisions.

Go to downloads

Report a problem

Please include the grant link, the field or sentence that appears wrong, and any source evidence that would help us reproduce the issue.

Email feedback

Attribution

Contains data from UK Research and Innovation (Gateway to Research), published under the Open Government Licence v3.0. Contains data from the National Institute for Health and Care Research (NIHR Open Data), Europe PMC GRIST and European Commission CORDIS. EC conversions use the European Central Bank annual GBP-per-EUR reference series. Derived topics, summaries, status, recipient decisions and geography are the responsibility of this project, not the source publishers.