A new study published on arXiv has evaluated the accuracy of large language models (LLMs) in retrieving bibliographic information for environmental science literature. Researchers compared the performance of several LLMs, including Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini, using both abstract-only and full-text prompts for 50 randomly selected articles from leading environmental science journals. The study found that abstract-based prompts generally yielded higher accuracy than full-text prompts, with accuracy also varying by LLM platform, journal, and the position of the reference in the output list. Overall, LLM-assisted literature retrieval in this field was found to be moderately accurate but inconsistent. AI
IMPACT Highlights limitations in current LLM capabilities for accurate scientific literature retrieval, suggesting caution for researchers relying solely on these tools.
RANK_REASON Academic paper evaluating LLM performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
- ChatGPT
- Claude
- DeepSeek
- Energy and Environmental Science
- Environmental Science & Technology
- Gemini
- Google Scholar
- Grok
- Lancet Planetary Health
- Nature Climate Change
- Nature Sustainability
- Perplexity
- Scopus
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →