A new two-part study published on arXiv explores the capabilities of large language models (LLMs) in assisting scientific research. The first paper details how mid-2025 models like ChatGPT, Claude, and DeepSeek performed in generating project plans and evaluating proposals, finding that while human reviewers rated AI-generated proposals similarly to human-written ones, AI reviewers showed a preference for AI-generated content. The second paper focuses on literature review assistance, revealing that mid-2025 LLMs selected only a small percentage of overlapping references with human experts and frequently produced references with metadata errors, though a 2026 model, ChatGPT Pro 5.5, showed improved reliability. AI
IMPACT LLMs show potential to assist in scientific research tasks like project planning and literature review, but require careful verification due to biases and potential for errors.
RANK_REASON The cluster consists of two academic papers published on arXiv detailing research into AI capabilities.
- arXiv
- ChatGPT
- ChatGPT 4o
- ChatGPT Pro 5.5
- Claude
- Claude Opus-4.8
- DeepSeek
- Gemini
- OpenAI Deep Research
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →