Two things this document fixes in place:
data/evidence/biomarker_evidence.json,
loaded and validated by api/evidence_catalogue.py).Every clinician-facing statement in MetaboGuard carries one of these labels, in the API, the dashboard and this documentation:
| Class | Meaning | Example |
|---|---|---|
data observation |
Measured directly in the data file | “HbA1c is observed in 62% of adults.” |
model association |
Produced by our model on our sample; not validated, not causal | “This phenotype has a higher median HbA1c than the reference.” |
published evidence |
Catalogued source with URL, study design and evidence grade | “GALAD reached AUC 0.78 within 12 months in a prospective 7-centre study.” |
causal claim not established |
Default status for any mechanism statement | “Nothing here shows that adiposity caused this person’s disease.” |
Required per row: entry_id, cancer_site, marker_or_panel, marker_class, specimen,
intended_use, stage_or_lead_time, direction, study_design, sample_size,
performance, validation_status, evidence_grade, limitations,
primary_source_url, doi, related_verified_sources.
Optional: repo_reference, available_in_current_data, current_data_column, notes,
screening_recommendation_status, allowlisted_statements, denied_statements.
Two explicit placeholders, never inferred:
unknown — the value exists but has not been extracted from the source.n.a. — the field does not apply to this row.An empty string is a validation error, so a missing field can never be mistaken for a negative finding.
| Gate | Rule |
|---|---|
| Provenance | primary_source_url must be a structurally valid URL, or doi a valid DOI, or both explicit placeholders. |
| Clinician-facing | A row reaches a clinician view only with a real source and a graded evidence_grade. |
| Statements | allowlisted_statements require a real source; rows without one may hold no allowlisted statement. |
| Causal language | Causal phrasing is rejected unless study_design is a causal design (RCT, Mendelian randomisation). |
| Universal-denial | The claim that cancers have no specific biomarkers is rejected outright. |
python api/evidence_catalogue.py --strict exits non-zero on any hard issue.
20 rows: 17 clinician-ready, 3 research-only. Sites: pancreas, liver (HCC), multi-site.
| Row | Marker or panel | Design | Key figures | Grade |
|---|---|---|---|---|
ev-ca199-alone-pdac-bjsopen-2024 |
CA19-9 alone | Meta-analysis of prediagnostic studies | AUC 0.998 at diagnosis, 0.87 at 6 mo, 0.74 at 12 mo, 0.55 at 5 yr | phase 3 prediagnostic, not screening grade (source, DOI 10.1093/bjsopen/zrae046; USPSTF) |
ev-thbs2-ca199-pdac-scitranslmed-2017 |
THBS2 + CA19-9 | Case-control phase 2b, n=537 | c-statistic 0.97; 87% sensitivity at 98% specificity; no lead time | discovery only (source, DOI 10.1126/scitranslmed.aah5583) |
ev-five-marker-panel-pdac-bjsopen-2024 |
CA19-9 + CA125 + VWF + THBS2 + IL6ST | Panel evaluation within the meta-analysis | AUC 0.91 within 1 yr; 0.78 up to 4 yr | internal discovery only (source) |
ev-endpac-pdac-digdissci-2020 |
ENDPAC (age, glucose change, weight change) | Retrospective cohort, 13,947 NOD / 99 PDAC | AUC 0.75; PPV 2.0%; NPV 99.7% | emerging external validation, risk stratification (source, DOI 10.1007/s10620-020-06139-z) |
ev-recent-diabetes-weightloss-pdac-jamaoncol-2020 |
Recent-onset diabetes ≤4 yr + >8 lb weight loss | Prospective cohorts, 112,818 women + 46,207 men, 1,116 PDAC | Incidence ratio 10.57 (7.18–15.56); absolute 4-yr incidence 0.29% | moderate prospective association (source, DOI 10.1001/jamaoncol.2020.2948) |
ev-thrombocytosis-multisite-bjgp-2017 |
Thrombocytosis >400×10⁹/L | Retrospective cohort with controls, ~40,000 exposed / 10,000 controls | 1-yr cancer PPV 11.6% men, 6.2% women; 18.1% / 10.1% after a second raised count; mainly lung and colorectal | moderate for risk marking (source, DOI 10.3399/bjgp17X691109) |
ev-excess-body-fatness-multisite-nejm-2016 |
Excess body fatness | IARC working-group review | Sufficient evidence: colon, kidney, postmenopausal breast, corpus uteri, liver, pancreas, ovary | IARC sufficient evidence for risk association (source, DOI 10.1056/NEJMsr1606602) |
ev-galad-hcc-gastro-2024 |
GALAD (sex, age, AFP-L3, AFP, DCP) | Prospective phase 3, 7 centres, n=1,558 / 109 HCC | AUC 0.78 vs AFP 0.66 within 12 mo; 62% sensitivity at 82% specificity | phase 3 prospective, high-risk surveillance (source, DOI 10.1053/j.gastro.2024.09.008) |
ev-cancerseek-multisite-science-2018 |
CancerSEEK (proteins + cfDNA) | Case-control, clinically detected, 1,005 cases / 812 controls | Median sensitivity 70% at >99% specificity; stage I 43%; breast 33% | discovery only, spectrum-biased for screening (source, DOI 10.1126/science.aar3247) |
Earlier rows transcribed from RESEARCH_EVIDENCE.md (HbA1c, C-peptide, HOMA-IR, CA19-9
lead time, weight loss, adiposity, hs-CRP null, lipid evidence gap, NOD panels) are
retained unchanged.
Any MetaboGuard claim about detection performance must satisfy these standards before it leaves the research setting:
| Standard | Applies to | Source |
|---|---|---|
| PRoBE | biomarker study design and specimen provenance | PRoBE design paper, DOI 10.1093/jnci/djn326 |
| TRIPOD+AI | reporting of prediction-model development and validation | BMJ 2024, DOI 10.1136/bmj-2023-078378 |
| PROBAST+AI | risk-of-bias and applicability appraisal | probast.org |
| STARD | reporting of diagnostic accuracy studies | EQUATOR |
Consequence for this project. Current outputs are cross-sectional deviation scores, exploratory phenotypes and reliability audits. They do not meet PRoBE specimen requirements, have no prediagnostic lead time, and are reported as research only.
Representative examples:
Never say, write or display:
Global pancreatic cancer burden was 531,318 cases and 490,786 deaths in 2024 (source, DOI 10.3322/caac.70090), and a demographic constant-rate projection gives 998,663 cases and 936,038 deaths in 2050 (source, DOI 10.1001/jamanetworkopen.2024.43198) if incidence and mortality rates stay unchanged.
This is a demographic projection holding rates constant. It is not a causal forecast, not a prediction of what will happen, and not attributable to any risk factor.
data/evidence/biomarker_evidence.json with every required field.unknown / n.a. explicitly instead of guessing.python api/evidence_catalogue.py --strict and python -m unittest test_research_pass.allowlisted_statements, confirm the source URL or DOI resolves.