Researchers have developed a new method for training large language models to reduce hallucinations in long-form text generation. This approach utilizes a rubric-based reward system that specifies required and optional information for an answer, rather than relying on simpler global richness proxies. Experiments show that a balanced combination of grounding, rubric coverage, and relevance rewards yields the best results, improving in-distribution support and out-of-distribution transfer. AI
IMPACT Introduces a novel training methodology for LLMs that could lead to more reliable and informative long-form text generation.
RANK_REASON This is a research paper published on arXiv detailing a new method for training LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →