A new study published on arXiv investigates how different reviewer guideline designs impact the effectiveness of Large Language Model (LLM)-based automated peer review. The research found that official conference guidelines, refined through established practices, yield results most aligned with human judgments. Conversely, guidelines generated to imitate human reviewers were less effective, and strict rubric-style scoring degraded performance. The study emphasizes the value of subjective and holistic scoring over rigid rubric enforcement for automated peer review. AI
IMPACT This research suggests that refining LLM guidelines based on established scientific practices can improve automated peer review accuracy.
RANK_REASON The cluster contains an academic paper detailing research findings on LLM-based automated peer review. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- conference practice
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Human judgments of positive and negative causal chains
- Influence Flower
- LLM
- peer review
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →