The endpoint belongs in the result name
The study changes endpoints and predictor sets explicitly. [2]
Our inference: an algorithm ranking without its endpoint and available features is underspecified.
Independent benchmark analysis / 2026 accepted journal manuscript
This published benchmarking study compares survival models across incident cardiovascular disease, breast cancer and Alzheimer’s disease, using several feature regimes. It supplies a concrete prevention-related evaluation rather than a general screening checklist. Our analysis asks how the endpoint, available information and validation design shape a model comparison. The reported outcome is risk discrimination: which people are ranked ahead of others. It does not by itself establish calibrated absolute probabilities or the benefit of an intervention. We retain the study’s name, conditions and uncertainty rather than branding its data as a new benchmark of our own.
01 / What is being tested?
Data origin. Longitudinal UK Biobank participant information and linked health outcomes; access-controlled research data. [1][2][3]
Cardiovascular disease, breast cancer and Alzheimer’s disease.
§2.2 [2]Age/sex; clinical risk; PRS/metabolomics; medical history; complete.
§2.3 [2]Cox, Lasso, Ridge, Elastic Net, random forest, LightGBM, XGBoost and deep learning.
Model implementations [2]Inner tuning is separate from outer evaluation; approximate cohort counts vary after endpoint exclusions.
§2.7 [2]Selected for <20% missingness across predictor matrices, largely requiring metabolomics coverage; before endpoint-specific baseline-disease exclusions.
§§2.4–2.5 [2]Exclude baseline occurrence of the corresponding endpoint.
[2]Fix one of the five feature regimes before comparison.
[2]Use nested cross-validation and fitted preprocessing.
[2]Report fold-level discrimination and uncertainty for that endpoint.
[3][2]02 / Measurement
Higher is better
Measures ordering of risk against observable event order among comparable pairs; the paper also reports Uno’s C.
C = concordant comparable pairs / comparable pairs, with the implementation’s tie handling
Compare identical endpoint, predictor regime and validation folds. C is not percent of patients correctly diagnosed. [2]
03 / Measured evidence
Paper-reported results / selected rows
2026 accepted manuscript and Supplementary Table 3; nested CV test performance for the Clinical Risk regime.
Study-specific implementations and tuning grids. These are discrimination values, not clinical accuracy percentages or universal algorithm rankings.
Source: Cardiovascular Disease worksheet; Clinical Risk rows [3][2]
04 / Our original analysis
The study changes endpoints and predictor sets explicitly. [2]
Our inference: an algorithm ranking without its endpoint and available features is underspecified.
The study focuses on discrimination rather than calibration. [2]
Our inference: an ordering result cannot alone support an absolute-risk treatment threshold.
The C-index compares observable pairs constructed from participant outcomes. [2]
Our inference: the number of pairs should not be used as though it were the number of independent patients when discussing uncertainty.
The selected clinical CVD estimates differ by 0.002 between deep learning and Cox-PH. [3]
Our arithmetic: 0.721 minus 0.719 equals 0.002; that difference is not a measured reduction in disease incidence.
05 / Scope of the evidence
UK Biobank recruitment is selective. The analysis further filters for predictor availability and missingness, largely removing participants without metabolomics. Transport to another population needs separate evidence. [2]
The comparison predicts disease events; it does not randomize or test a prevention policy. [2]
The paper is open, but participant records require UK Biobank researcher approval. [2]
The journal labels this an article in press; the cited supplement is the 2026 version, not the 2025 preprint. [1]
Evidence trail
Oexner et al. / King’s College London and UCL. Accepted peer-reviewed article; results differ in some uncertainty estimates from the 2025 preprint.
Oexner et al. / Journal of Big Data. Complete methods, endpoints, nested cross-validation and limitations.
Oexner et al. / Journal of Big Data. Published workbook containing condition-specific model results and confidence intervals.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.