The short answer
Harrell’s C measures agreement between risk ordering and observable event ordering. It is not the fraction of patients correctly diagnosed and not a calibrated probability of future disease. The UK Biobank study explicitly emphasizes discrimination. Our original explanation shows why that is useful evidence while leaving additional questions for anyone who wants to act at a numerical risk threshold.
Understand what the denominator contains
The C-index compares pairs whose outcome ordering is observable under the survival-data rules. Censoring matters because a participant who has not yet had an observed event cannot always be ordered relative to another participant. The implementation determines comparability and handles ties. Calling the result ordinary patient-level accuracy would erase those rules.
Our suggested reporting language is risk-order concordance for the stated endpoint and follow-up setting. Keep the model’s risk direction explicit: higher scores must consistently mean higher predicted risk. For a replication, use the documented implementation rather than an improvised pair counter, and record how censoring and tied predictions are treated. A mathematically familiar metric still requires a concrete operational definition.
Use an abstract example to separate ordering from probability
Imagine two illustrative prediction systems that rank the same abstract individuals in exactly the same order. One assigns probabilities of 1%, 2% and 3%; the other assigns 20%, 40% and 60%. These numbers are invented solely to explain the concept. The ranking can remain identical while the proposed absolute probabilities differ greatly.
This is why a concordance result alone cannot establish whether a numerical intervention threshold is well calibrated. A monotonic change in scores can preserve ordering while changing which people cross an absolute threshold. The example is not a medical recommendation, a risk calculator or a reported result from the study. It isolates the mathematical distinction that a prevention decision must account for separately.
Interpret the selected study measurements at their own scale
Our results panel uses the 2026 supplement’s cardiovascular-disease condition with clinical-risk predictors. It retains the C-index scale and confidence intervals. The difference between the displayed deep-learning and Cox values is an arithmetic score difference within that experiment. It is not a measured reduction in disease events or an estimate of additional people helped.
We also avoid converting the displayed confidence intervals into a new significance claim. The source’s validation and statistical procedures belong to the study. A practical decision might additionally consider computation, predictor availability and uncertainty about transport. Those considerations can be analyzed, but they should not be hidden inside a single discrimination number or described as outcomes the benchmark directly measured.
Ask for the evidence needed by the intended decision
If the intended use involves an absolute-risk threshold, the next questions concern calibration in the relevant population and observation horizon. If resources are allocated by ranking, the question may instead concern the consequences of selecting a particular part of the ranked list. In either case, acting on predictions introduces a decision question beyond the ordering metric itself.
Our proposed evaluation brief therefore has separate lines for discrimination, probability calibration, decision policy and observed outcomes after action. A missing line remains missing evidence; it is not filled by a strong result elsewhere. This framework makes the prevention relevance of the survival benchmark clearer while preserving its limits. It also explains why we present the work as an analysis of a published study rather than a universal prevention-performance score.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- UK Biobank survival benchmark: accepted manuscript ↗Oexner et al. / Journal of Big Data. Complete methods, endpoints, nested cross-validation and limitations.
- UK Biobank survival benchmark: Supplementary Table 3 ↗Oexner et al. / Journal of Big Data. Published workbook containing condition-specific model results and confidence intervals.
- Comprehensive benchmarking of machine learning methods for risk prediction modelling from large-scale survival data: a UK Biobank study ↗Oexner et al. / King’s College London and UCL. Accepted peer-reviewed article; results differ in some uncertainty estimates from the 2025 preprint.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.