A new paper published on arXiv, "The C-index illusion: discrimination without calibration in published survival models," challenges the common practice of evaluating survival analysis models solely on their discrimination, as measured by the C-index. The research demonstrates that this metric can be misleading because it ignores model calibration and time-dependent accuracy. The study reproduced three published survival-ML models from diverse domains, finding that a model with nearly identical discrimination to a published one failed a formal calibration test. The paper also highlights how misinterpreting censoring as non-informative can lead to significant biases in risk estimations for financial models. AI
IMPACT Highlights potential flaws in common ML model evaluation metrics, urging caution in interpreting discrimination scores.
RANK_REASON Academic paper published on arXiv discussing methodology for evaluating machine learning models. [lever_c_demoted from research: ic=1 ai=1.0]
- churn model
- C-index
- Competing Risk of Death and ESRD in Incident CKD Patients
- hard disk drive failure
- Holm-corrected family-wise error rate
- International Conference on Machine Learning
- loan prepayment
- peer-to-peer credit default
- Survival models for heterogeneous populations derived from stable distributions
- User Disengagement-Oriented Target Enforcement for Multi-Tenant Database Systems
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →