A new study published on arXiv investigates the sensitivity of Large Language Models (LLMs) used for resume screening to presentation variations. Researchers found that even when the underlying candidate qualifications remain the same, changes in wording, structure, or stylistic polish can lead to different screening decisions. For instance, Llama-3.1-8B, despite achieving strong validity scores, reversed nearly 30% of its decisions when presented with differently formatted resumes. Similarly, Mistral-7B-v0.3 showed a higher flip rate of over 41% with a lower validity score. The study highlights the need for resume screening evaluations to consider not only accuracy but also the stability of decisions when faced with presentation changes. AI
IMPACT Highlights potential biases and unreliability in LLM-based screening tools, suggesting a need for more robust evaluation methods.
RANK_REASON Research paper published on arXiv detailing LLM performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →