Researchers have developed PhenoBench, an executable benchmark designed to evaluate AI models using data from the Human Phenotype Project, which includes over 13,000 participants. This benchmark comprises 90 clinically grounded tasks across 15 domains and 26 input modalities, allowing for the assessment of how different measurements inform health-related questions. Initial evaluations showed that while tabular foundation models generally outperformed standard task-specific models, their aggregate improvement was modest. Language models also demonstrated informative predictions on certain tasks without cohort-specific fitting, though they exhibited capability gaps and rarely surpassed models trained on the same data. AI
IMPACT Establishes a new evaluation framework for AI models in healthcare, enabling more standardized comparisons of their performance on complex human health data.
RANK_REASON The cluster describes a new benchmark and evaluation framework for AI models using a deeply phenotyped human cohort, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Human Phenotype Project
- Influence Flower
- Litmaps
- PhenoBench
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →