A new dataset called WildSEEK has been developed to evaluate language models' performance on real-world information-seeking queries. The dataset includes over 3,000 manually annotated queries, with a focus on risk-sensitive domains like health and finance, and distinguishes between factoid and analytical queries. Analysis of over 1.8 million user queries using classifiers trained on WildSEEK revealed that more than a third are high-risk and analytical, with LLM responses frequently failing in areas such as sycophantic behavior, overreliance, US-centric bias, and poor handling of vulnerable populations. AI
IMPACT Provides a framework for assessing LLM reliability, safety, and fairness in information access, crucial for responsible deployment.
RANK_REASON The cluster contains a research paper introducing a new dataset and evaluation framework for language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
- WildSEEK
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →