Benchmark analysis / 5 min read

Why the endpoint and feature set belong beside a survival-model ranking

Use the UK Biobank study to make model comparisons conditional and reproducible.

The short answer

A ranking of algorithms without its data conditions is incomplete. In the UK Biobank survival benchmark, the endpoint and predictor regime are deliberate experimental axes. Our coverage explorer makes those axes visible so readers can inspect the question behind a score. The analysis is based on the 2026 accepted manuscript and its matching results workbook, not an independent rerun or a new dataset release.

Treat each condition as a separate comparison

A clinical-risk model for cardiovascular disease and a complete-feature model for Alzheimer’s disease differ in both target and inputs. Sorting their C-indices together would not answer which algorithm is best for either task. A valid comparison first holds the scientific condition fixed, then examines the competing implementations within it.

We recommend a result key comprising endpoint, feature regime, eligibility rule, split and model configuration. This key is more informative than a short experiment nickname. It also provides a place to record missing cells. If a method was not run under a condition, show that absence rather than borrowing a score from a neighboring condition. The explorer’s rows are task descriptions, not invented performance observations.

Separate algorithm family from implementation evidence

The paper documents implementations, package versions, fixed settings and tuned hyperparameters. A reported value therefore belongs to a configured method within this study. It is not a timeless estimate of every implementation bearing the same algorithm-family name. Search spaces, stopping rules and preprocessing can matter to an empirical comparison.

Our original interpretation is to preserve the strongest supported wording: under this condition and protocol, these implementations produced these measurements. A broader statement that one algorithm is universally superior needs broader evidence. Keeping that qualifier does not weaken a useful finding. It makes the result portable as a hypothesis to test in the next cohort rather than a rule to assume.

Read fold uncertainty without inventing independent patients

The published workbook includes test discrimination and confidence intervals for each condition. The nested design reuses a common cohort across folds and comparisons. A reader should not treat the collection of task combinations as independent patient populations or assume that a large number of pairwise comparisons creates an equally large number of independent observations.

When reproducing the study, retain fold-level results and the exact uncertainty procedure. If comparing two models, preserve their matched fold or participant structure rather than reconstructing a significance claim from rounded means alone. Our selected results preserve the source intervals. We have not performed an additional hypothesis test, and the site does not mark a winner using an invented threshold.

Make the next experiment answer a new question

Once the within-condition comparison is understood, ask what uncertainty remains. It may concern transport to a different population, availability of the predictors, calibration or performance in a subgroup. Repeating the same benchmark with a new plotting style does not answer those questions. A useful follow-up specifies which assumption is being challenged.

Our suggested decision record links the selected benchmark condition to the intended use and lists the evidence still needed. Keep the original result and the follow-up result separate, even if both use a C-index. This allows a reader to trace what changed and why. It turns a model ranking into a sequence of inspectable research decisions without claiming that the benchmark alone validates a prevention service.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. UK Biobank survival benchmark: accepted manuscript ↗Oexner et al. / Journal of Big Data. Complete methods, endpoints, nested cross-validation and limitations.
  2. UK Biobank survival benchmark: Supplementary Table 3 ↗Oexner et al. / Journal of Big Data. Published workbook containing condition-specific model results and confidence intervals.
  3. Comprehensive benchmarking of machine learning methods for risk prediction modelling from large-scale survival data: a UK Biobank study ↗Oexner et al. / King’s College London and UCL. Accepted peer-reviewed article; results differ in some uncertainty estimates from the 2025 preprint.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →