Algo-Model

Research decision record — 2026-08-04: professor feedback on scope and method

Status: accepted, implemented in this repository on 2026-08-04. Scope: MetaboGuard research direction, evidence handling, and analysis method.

These are recollected notes, written after the meeting from memory. They are not a transcript and contain no verbatim quotations. Where a statement is our own inference or engineering consequence rather than something a supervisor said, it is marked (our interpretation). Corrections from attendees are welcome and should be added as a dated amendment below rather than by rewriting this record.

Attendees referenced

Recollected feedback

1. Prof. Helmount — clinical value depends on early-stage detection

Recollected substance:

Explicitly not claimed, and not to be written anywhere in this repository (our interpretation of the correct scientific statement):

2. Prof. Nada — use clustering to discover phenotypes

Recollected substance:

Consequences (our interpretation):

Decisions taken

# Decision Rationale
D1 Terminology: no causal phrasing. Use risk-associated features, early-development signals, biological pathways. causal is reserved for the DoWhy estimation module, where it names a method, not a finding. Nothing in the current data supports causal claims.
D2 Early detection is framed as a panel + interaction problem; the repository must never state that cancers have no specific biomarkers. Scientific accuracy; see above.
D3 Clustering is exploratory phenotype discovery, label-free in fit and selection, with an explicit abstain result when stability criteria fail. A clustering that is not stable is not a finding.
D4 An evidence catalogue with mandatory provenance (URL/DOI, study design, validation status, evidence grade, limitations) gates any doctor-facing statement. Unknowns are recorded as unknown, never inferred. Prevents evidence drift and fabricated citations.
D5 A data reliability report with feature eligibility tiers (usable_now, qualified_use, unavailable, prohibited) runs before analysis and fails closed on hard violations. Reliability must be a gate, not a footnote.
D6 Future-risk and cancer-site outputs stay disabled and fail-closed; NHANES here is cross-sectional and TCGA is post-diagnosis context only. Unchanged from the pre-existing capability gates.
D7 Clinician-facing explanations must label every statement as data observation, model association, published evidence, or causal claim not established. Keeps the epistemic status visible at the point of reading.

What this does not change

Blocker that this feedback makes sharper, not smaller

Prof. Helmount’s early-stage requirement cannot be met with the current data at all: there is no follow-up time, no incident outcome and no stage information in the cross-sectional NHANES files, and TCGA is post-diagnosis. Cluster phenotypes are therefore a hypothesis generator for a future longitudinal study, not evidence of early detection.

Amendments

(none yet — append dated entries here rather than editing the notes above)