A new study published on arXiv reveals that multimodal large language models (MLLMs) systematically overrate facial attractiveness compared to human judgments. Researchers compared ratings from 2,513 human participants with those from four commercial AI models: Claude, Gemini, GPT, and Grok. While the AI models showed strong correlations with human rank-ordering of attractiveness, they consistently rated faces more favorably and within a narrower range than humans. Only face age was a consistent predictor of attractiveness across both humans and MLLMs, with other factors like ethnicity and gender showing inconsistent patterns among the models. Grok, in particular, exhibited the lowest agreement with human ratings. AI
IMPACT Suggests current MLLMs may not be reliable for subjective tasks like beauty assessment, highlighting differences in AI and human perception.
RANK_REASON Research paper published on arXiv detailing findings about AI model capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Claude
- DagsHub
- Gemini
- generative pre-trained transformer
- Gotit.pub
- Grok
- Hugging Face
- Mohit Mendiratta
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →