Inputs
Sources and scope
Source acquisition happens in a separate grant-data workspace. Each source is collected and processed on its own terms before this public-site pipeline receives harmonised project and researcher files. That separation matters: it preserves provenance and avoids pretending that four very different databases are one uniform dataset.
UKRI Gateway to Research
Open source site ↗Four separate feeds—projects, funding, people and organisations—are joined on UKRI project identifiers. Principal investigators, co-investigators, fellows, supervisors and students become linked researcher rows rather than being flattened into one field.
- Public records
- 173,320
- Current raw download
- 18 February 2026
- Licence
- OGL v3
NIHR Open Data
Open data portal ↗We combine the funded portfolio, Research Schools projects and the infrastructure-supported projects dataset. The last of these records support relationships, so it is retained for discovery but separated from financial awards.
- Funded awards
- 12,480
- Support relationships
- 42,102
- Current retrieval
- 18–19 June 2026
Europe PMC GRIST
Grant finder ↗Europe PMC aggregates grants from medical and research charities. Its rows are often organised around researchers and affiliations rather than a definitive recipient, so project deduplication and recipient attribution require separate treatment.
- Public charity records
- 27,395
- Current retrieval
- 19 June 2026
- Important caveat
- Affiliation is not always recipient
European Commission CORDIS
CORDIS datasets ↗Monthly Horizon 2020 and Horizon Europe project, organisation, topic and EuroSciVoc files are joined by CORDIS project ID. A project enters this site when at least one participant has an explicit UK or GB country code. It is counted once, while every UK participant remains visible.
- UK-involved projects
- 15,046
- Funding measure
- Known UK participant contributions
- Important caveat
- Partial totals are lower bounds
What is included
The funded-award views cover records with a usable start year of 2006 or later. Records without a usable start year do not enter the current public award corpus. The coverage is substantial but not exhaustive of every UK funder: it is the union of the sources above after the decisions described on this page.
Harmonisation
Preparing the records
A shared public schema makes search and comparison possible, but the pipeline keeps enough source information to avoid unsafe joins. The preparation stage is a sequence of explicit decisions, not a blanket “clean data” command.
- 1
Preserve source-aware identity
Every public identifier receives a source prefix such as UKRI, NIHR or EPMC. Researcher rows join on both source and raw grant ID. This prevents the 1,428 raw identifiers that currently occur in more than one source from colliding.
- 2
Map to common fields
Titles, abstracts, funders, dates, amounts, currencies, organisations and people are mapped into common columns. Amount strings are parsed, researcher roles are unfolded, and obvious duplicate display rows are removed without collapsing different roles.
- 3
Reduce known overlap—without claiming perfect deduplication
Europe PMC produces one project row per grant ID while retaining its researcher-affiliation rows. Records from a maintained list of UKRI and NIHR funders are excluded from the Europe PMC charity feed. Spelling variants, transfers and less obvious cross-source overlaps can remain.
- 4
Separate awards from support relationships
NIHR infrastructure-supported projects remain searchable, but they do not inflate the funded award count, award-value total, funding trends or institution funding rankings. The rule uses the source dataset, not keywords in a title.
- 5
Derive status and public totals
“Upcoming”, “active”, “completed” and “unknown” are calculated from published start and end dates for every public export. The £97.6bn headline is the sum of positive published award-record values. It is not annual spend, cash expenditure or an inflation-adjusted figure.
Zero is not treated as a free grant
A missing or non-positive amount is displayed as “not disclosed”. In the current funded-award corpus, 64,671 records—28.4%—lack a positive published value.
Status is a snapshot
Status is recalculated when a public release is exported, not live in the browser. It can lag the source dates between releases, and invalid source date combinations are not silently rewritten.
Attribution
Recipients, affiliations and geography
“Who received the award?” and “which organisations are connected to it?” are not always the same question. UKRI and NIHR normally identify a lead organisation. Europe PMC often provides several researchers’ affiliations without naming a grant-level recipient, so we deliberately avoid assigning the whole award to whichever institution happens to appear first.
This produces 25,764 Europe PMC records with an attributable recipient and 1,631 association-only records in the current release. Association-only values remain in funder totals, but not in institution or regional allocations. Twenty-two independently checked exceptions are held as evidence-backed export overrides; source records remain untouched.
Mapping rules
- The map represents where a recipient organisation is recorded as based, not the exact project site or department.
- Europe PMC’s source country is not accepted as proof of recipient location.
- Locations resolve through a promoted registry: manual overrides first; NIHR source coordinates; unique exact Companies House registered-office sectors for company-like recipients; UKRI postcodes; then exact ROR matches and other reviewed fallbacks.
- A word such as “Cambridge” or “Oxford” in an organisation name is never treated as location evidence.
- Companies House points show a registered-office postcode sector, which may be an accountant or legal office rather than a research site. ROR points normally represent an organisation’s recorded city.
- Unresolved or competing affiliations stay searchable but are left off UK maps and region rankings.
Semantic classification
How grants are categorised
The public taxonomy contains 29 topics in 6 fields. Health topics are broadly informed by health-research classification practice; other topics draw on established research-field structures, then use plain labels intended for non-specialists. “Infrastructure” is a record facet rather than a topic, because a facility can support cancer, physics, computing or any other substantive field.
Choosing the examples
The 100 exemplars per topic are not simply the largest awards or the records a first-pass model liked most. Candidate selection excludes title-only and infrastructure records, spreads examples across semantic sub-groups, and considers confidence, recency, previous review, funder and institution diversity, and duplicate titles. A separate set of 50 records per topic—1,450 in total—is held out from the exemplar pool for evaluation.
A versioned curation layer currently records 218 forced anchors, 125 forced held-out examples and 6,495 exclusions from anchor selection. These are structured human-directed, Codex-assisted review decisions, not an independently expert-labelled disciplinary gold standard.
Stability between releases
Once approved, a grant’s topic is stored against its source-aware ID. Routine refreshes reuse historic decisions and classify only new IDs. The registry also fingerprints the embedding model, dimensions, input fields, token limit and long-text pooling recipe. A normal export stops if the vectors and the category registry were built under different recipes.
The EC incremental category gate
New CORDIS IDs are embedded in a scratch release directory under the unchanged title-plus-full-objective contract and classified against the pinned exemplars. Historic vectors, assignments and exemplar IDs are fingerprinted and must remain unchanged. EuroSciVoc agreement is reported as a diagnostic only because it is a different taxonomy. Publication requires a prediction-blind, independently labelled EC held-out set with at least 10 projects in every public topic and at least 290 projects overall, at least 93% top-three accuracy, no more than 12% wrong-confident assignments, and a fingerprint-matching decision for every targeted conflict. No QA waiver is available for this release.
A genuine method change follows a separate migration route: produce a shadow classification, compare every proposed move with the locked labels, prioritise high-value, held-out, low-margin and cross-domain changes, then require an explicit, fingerprint-matching promotion decision. Durable manual moves remain authoritative after semantic scoring.
The 10 July 2026 full-abstract migration changed 49,752 of 260,337 source-level assignments (19.1%). Leave-one-out testing scored 97.7% top-three accuracy, but the harder held-out set scored 76.8%. The confidently-wrong rate was 21.5%, above the intended 12% ceiling. No rows in the 50,666-row migration queue were recorded as individually reviewed.
Promotion therefore did not represent a clean pass of every QA gate. It used an explicit, dated waiver after side-by-side map inspection, recording the failed gates and accepted conflicts. This is why the site presents topics as navigational aids, preserves a runner-up topic, and warns that interdisciplinary and title-only records are the hardest cases.
Plain English
How lay summaries are prepared
The original funder title is retained; the site does not generate replacement “catchier” titles. For a subset of records with substantive source abstracts, DeepSeek V3 has prepared an accepted 150–200 word explanation. The prompt asks for a concrete opening, the problem or knowledge gap, what the work will do, and what could change if it succeeds—without forcing a daily-life application onto fundamental research.
Generated from a substantive source abstract, labelled as AI and passed through safety checks.
The source abstract is shown when no accepted AI summary is available.
No speculative research description is generated when the source provides too little evidence.
Counts refer to the 269,841 searchable records in the current served snapshot.
Editorial decisions in the prompt
- Use plain, active language for an intelligent reader who is not a specialist.
- Explain jargon rather than merely replacing it with another technical term.
- Include a number only when the exact number appears in the source abstract.
- Do not infer mechanisms, patient groups, locations, applications or delivery routes.
- Avoid fictional characters, melodrama, anthropomorphism and extended metaphors.
- State potential impact as a possibility where the source does not establish an outcome.
Model and prompt testing
DeepSeek V3 and Kimi K2 were first compared side by side. The prompt then went through versions 2–5 on a fixed 75-grant test set, including records with and without abstracts, before the final version was run again with DeepSeek. Review found the prose generally faithful and readable, but identified unsupported numbers and confident title-only inference as the highest-risk behaviours. The production decision was therefore to require a usable abstract and to add stricter source-only rules.
Making the work durable
Accepted summaries live in a sidecar dataset keyed by the source-aware grant ID. Enrichment merges that store before export, so a routine rebuild cannot erase paid generation work or replace it with an older fallback. Grant pages label the text and retain the source abstract, plus a link to the original funder record where the source publishes one.
Exceptions
Where manual correction was necessary
Automation handles volume; it does not make unusual records disappear. Manual and curated decisions are stored in versioned overlays and applied at the relevant curation or export stage, so source data is not silently rewritten. Category overlays retain the grant ID, action and—where moved—target topic. Recipient corrections additionally retain a reason and evidence URL. Provenance depth varies by correction workflow.
| Problem found | How it was found | Durable response |
|---|---|---|
| Category boundary errors Commercial technology, generic clinical trials and interdisciplinary records crossed topic boundaries. | Representative, newest, high-value and targeted boundary review packs; the Semantic Map used only to locate suspicious neighbourhoods. | 1,478 explicit category moves and 62 “do not feature as representative” decisions currently override automated output. |
| Contaminated exemplars Removing a bad anchor could allow another unreviewed record to refill the slot. | Direct inspection of the actual 100-record exemplar sets and the exemplar-builder refill route. | Forced anchors, held-out cases, exclusions and topic-specific guards make review decisions persistent. |
| Thin and placeholder abstracts AI filled evidence gaps with plausible but unsupported detail. | Placeholder scan, number audit and an attempted census of 5,289 summaries based on very short abstracts; 5,223 reviews completed. | 2,029 IDs are gated in total, including 66 uncompleted audit cases held back as a safe default; shared hygiene prevents re-entry. |
| Europe PMC recipient ambiguity The first affiliation was sometimes mistaken for the award holder. | Top-value recipient review against local affiliation evidence and authoritative funder or recipient pages where needed. | 22 evidence-backed recipient overrides; multi-affiliation records remain association-only when unresolved. |
| NIHR support records counted as awards Centre-support relationships inflated grant and institution totals. | Trace of the original NIHR infrastructure dataset and its record meaning. | 42,102 relationships stay searchable under a separate record type and are excluded from award totals. |
| Unsafe or incomplete geography A funder/currency country signal did not prove a UK recipient location. | Known-institution checks, overseas-marker tests and review of high-value unmatched names. | Source-aware location registry with exact identifiers or names, provenance, conflict review and no city-token inference; unresolved organisations are not mapped. |
Quality assurance
Testing and release gates
Testing is layered because no single accuracy number can validate ingestion, attribution, topic assignments, prose and a static website at once. Automated suites are run alongside review packs and release gates, while the known gaps and failed category gates are reported here rather than hidden behind a test total.
Record and export invariants
Tests cover source-aware joins, duplicate researcher display rows, award/support separation, recipient versus association-only logic, map exclusion, missingness calculations, organisation aggregation, clean directory rebuilds and safe grant URLs.
Category mechanics
Tests enforce exactly 100 unique exemplars and 50 non-overlapping held-out cases per topic, vector alignment, score construction, ambiguity flags, durable overrides, registry fingerprints and migration comparisons.
Category quality
Leave-one-out and held-out tests, transition matrices, cross-domain change counts, high-value queues, boundary packs and side-by-side map views expose where classifications move. The present held-out result remains below the intended gate.
Summary fidelity
Model comparison, five prompt generations, manual sample review, exact-number scanning, thin-text census, placeholder/refusal detection and idempotence tests focus on whether prose remains grounded in the source.
Public build and files
The Astro production build and search/Semantic Map component tests run separately from pipeline export. Export-time geography assertions are blocking; downloads are rebuilt from cleaned public chunks, although download-generation failure is reported rather than blocking the rest of an export.
The 100-largest-awards gate
Every normal export compares the proposed 100 largest displayed awards with the last committed snapshot. A new or materially changed amount, recipient, affiliation set or geography blocks release until a fingerprint-matching decision is recorded. The gate watches change; it does not mean all 100 baseline rows have been independently reverified.
The category migration gate
A change to the semantic recipe must produce a shadow report, a review queue and a decision tied to the exact report fingerprint. Promotion normally requires all absolute checks and conflicts to pass. The code also permits an explicit QA waiver with an approval flag and a non-empty reason. The current waiver goes further by recording the failed gates, accepted conflicts, reviewer, date and reason; no waiver was implied merely by proceeding with an export.
The EC consortium and category gate
The EC release additionally checks one public row per CORDIS project, exact equality between project-level known UK funding and the participant projection, explicit lower-bound labels, unchanged historic category rows and vector contract, balanced blind held-out accuracy, and completed targeted review. Unlike a general category migration, this EC release has no waiver route. Refreshed upstream data and shadow artefacts are retained when a gate fails, but public files are not installed.
Interpretation
Known limitations
Coverage is broad, not complete
The site covers four public systems, not every UK research funder or every award made outside those systems.
Source freshness is uneven
Retrieval dates differ and publishers can amend older records. The website is not a live or append-only register.
Published values are incomplete
Missing values are not estimated. Totals are not inflation-adjusted expenditure and can still be affected by unresolved cross-source overlap.
Recipient and place are not always knowable
A researcher affiliation is not necessarily an award holder. Map points may be source coordinates, postcode centroids, a ROR city or a Companies House registered-office area; none proves the location of every research activity.
One primary topic simplifies interdisciplinary work
A runner-up helps, but 29 topics cannot express every discipline, method or application. Current held-out validation shows substantial residual error.
AI prose can still be wrong
Safety gates reduce risk; they do not prove every sentence. The source abstract and original record remain authoritative.
Absence is not evidence of no research
A missing abstract, value, recipient, person or map point often reflects publishing practice rather than the importance or existence of the work.
Stewardship
Updates, reuse and contact
We aim to refresh source data roughly quarterly. A refresh can change old records as well as add new ones, so historic IDs and reviewed overlays are retained. Topic assignments and summaries are reused unless a new record, explicit correction or controlled migration requires change.
The current process does not yet emit a public release manifest containing a checksum and retrieval date for every source. Adding that manifest is a planned transparency improvement. Until then, the source-specific dates on this page are more informative than a single “as of” date.
Download the public data
Downloads are built from the cleaned grant chunks served by the site, so the downloadable view follows the same record-type, summary and attribution decisions.
Go to downloadsReport a problem
Please include the grant link, the field or sentence that appears wrong, and any source evidence that would help us reproduce the issue.
Email feedbackAttribution
Contains data from UK Research and Innovation (Gateway to Research), published under the Open Government Licence v3.0. Contains data from the National Institute for Health and Care Research (NIHR Open Data), Europe PMC GRIST and European Commission CORDIS. EC conversions use the European Central Bank annual GBP-per-EUR reference series. Derived topics, summaries, status, recipient decisions and geography are the responsibility of this project, not the source publishers.