Independent benchmark analysis / 2026 accepted journal manuscript

UK Biobank survival-model benchmark by Oexner et al.

Ranking future risk is only one part of a prevention decision.

This published benchmarking study compares survival models across incident cardiovascular disease, breast cancer and Alzheimer’s disease, using several feature regimes. It supplies a concrete prevention-related evaluation rather than a general screening checklist. Our analysis asks how the endpoint, available information and validation design shape a model comparison. The reported outcome is risk discrimination: which people are ranked ahead of others. It does not by itself establish calibrated absolute probabilities or the benefit of an intervention. We retain the study’s name, conditions and uncertainty rather than branding its data as a new benchmark of our own.

01 / What is being tested?

The task, before the score.

input
Baseline participant predictors under a specified feature regime
output
A survival-risk score for a defined incident-disease endpoint
unit
One participant with event/censoring time
setting
Nested five-fold cross-validation; no single fixed public test set

Data origin. Longitudinal UK Biobank participant information and linked health outcomes; access-controlled research data. [1][2][3]

Disease endpoints
3

Cardiovascular disease, breast cancer and Alzheimer’s disease.

§2.2 [2]
Feature regimes
5

Age/sex; clinical risk; PRS/metabolomics; medical history; complete.

§2.3 [2]
Primary model families
8

Cox, Lasso, Ridge, Elastic Net, random forest, LightGBM, XGBoost and deep learning.

Model implementations [2]
Validation
Nested 5-fold CV

Inner tuning is separate from outer evaluation; approximate cohort counts vary after endpoint exclusions.

§2.7 [2]
Analysis population
Approximately 240k

Selected for <20% missingness across predictor matrices, largely requiring metabolomics coverage; before endpoint-specific baseline-disease exclusions.

§§2.4–2.5 [2]
  1. 01

    Define incident outcomes

    Exclude baseline occurrence of the corresponding endpoint.

    [2]
  2. 02

    Specify available predictors

    Fix one of the five feature regimes before comparison.

    [2]
  3. 03

    Tune inside training folds

    Use nested cross-validation and fitted preprocessing.

    [2]
  4. 04

    Evaluate held-out ordering

    Report fold-level discrimination and uncertainty for that endpoint.

    [3][2]

02 / Measurement

Harrell’s concordance index

Higher is better

Measures ordering of risk against observable event order among comparable pairs; the paper also reports Uno’s C.

Scoring definition

C = concordant comparable pairs / comparable pairs, with the implementation’s tie handling

Compare identical endpoint, predictor regime and validation folds. C is not percent of patients correctly diagnosed. [2]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Cardiovascular disease with clinical-risk features

2026 accepted manuscript and Supplementary Table 3; nested CV test performance for the Clinical Risk regime.

Harrell’s C · C-index
00.51
Reported
Minimally penalised Cox-PH95% CI 0.717–0.721.
0.719
Deep learning95% CI 0.719–0.722.
0.721
XGBoost95% CI 0.560–0.569.
0.564

Study-specific implementations and tuning grids. These are discrimination values, not clinical accuracy percentages or universal algorithm rankings.

Source: Cardiovascular Disease worksheet; Clinical Risk rows [3][2]

04 / Our original analysis

What follows from the design?

01

The endpoint belongs in the result name

Published evidence

The study changes endpoints and predictor sets explicitly. [2]

Our interpretation

Our inference: an algorithm ranking without its endpoint and available features is underspecified.

02

Discrimination leaves an unanswered prevention question

Published evidence

The study focuses on discrimination rather than calibration. [2]

Our interpretation

Our inference: an ordering result cannot alone support an absolute-risk treatment threshold.

03

Validation units are participants, not pair counts

Published evidence

The C-index compares observable pairs constructed from participant outcomes. [2]

Our interpretation

Our inference: the number of pairs should not be used as though it were the number of independent patients when discussing uncertainty.

04

A small score gap is context dependent

Published evidence

The selected clinical CVD estimates differ by 0.002 between deep learning and Cox-PH. [3]

Our interpretation

Our arithmetic: 0.721 minus 0.719 equals 0.002; that difference is not a measured reduction in disease incidence.

05 / Scope of the evidence

Where this benchmark stops.

Selected source cohort and assay coverage

UK Biobank recruitment is selective. The analysis further filters for predictor availability and missingness, largely removing participants without metabolomics. Transport to another population needs separate evidence. [2]

No intervention endpoint

The comparison predicts disease events; it does not randomize or test a prevention policy. [2]

Controlled data access

The paper is open, but participant records require UK Biobank researcher approval. [2]

Accepted manuscript version

The journal labels this an article in press; the cited supplement is the 2026 version, not the 2025 preprint. [1]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Paper and aggregate supplements public; participant data restricted
License
Article CC BY 4.0; participant data under UK Biobank access terms
Conditions
Approved-researcher access is required for participant data. Article licensing does not grant dataset access.
[1][2]

Evidence trail

Read the originals.

  1. Comprehensive benchmarking of machine learning methods for risk prediction modelling from large-scale survival data: a UK Biobank study ↗

    Oexner et al. / King’s College London and UCL. Accepted peer-reviewed article; results differ in some uncertainty estimates from the 2025 preprint.

  2. UK Biobank survival benchmark: accepted manuscript ↗

    Oexner et al. / Journal of Big Data. Complete methods, endpoints, nested cross-validation and limitations.

  3. UK Biobank survival benchmark: Supplementary Table 3 ↗

    Oexner et al. / Journal of Big Data. Published workbook containing condition-specific model results and confidence intervals.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗