A new research paper published on arXiv explores the limitations of stochastic sampling in large language models (LLMs). The study, titled "Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs," suggests that the variation in outputs from a single LLM run multiple times does not effectively reveal what the model doesn't know. Researchers compared a single model run 100 times with an ensemble of 24 LLMs run once, using benchmarks like MMLU, HellaSwag, and GSM8K. Their findings indicate that while self-consistency provides per-question uncertainty, it lacks cross-question structural insights, which are only surfaced by a diverse ensemble of models. AI
IMPACT This research suggests that relying on single model output variation for uncertainty estimation may be insufficient, highlighting the value of model ensembles for a more robust understanding of LLM knowledge gaps.
RANK_REASON Research paper published on arXiv detailing findings about LLM sampling. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- GSM8K
- HellaSwag
- Hugging Face
- Marchenko--Pastur
- Massive Multitask Language Understanding
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →