Researchers have developed a new pipeline to transform scientific papers into multi-turn generation trajectories for continued pre-training of language models. This method reconstructs the writing process of a paper, including requests, plans, and section-level deliberations, while keeping the original text verbatim. The resulting corpus, derived from arXiv papers, is approximately twice the size of the source text. This approach also enables the creation of instruction datasets and a new academic writing benchmark called PAW-Bench. Experiments show that continued pre-training on this corpus, followed by supervised fine-tuning, significantly improves writing capabilities without compromising general reasoning or long-document comprehension. AI
IMPACT Enhances LLM training data by leveraging structured scientific papers, potentially improving academic writing and long-document understanding.
RANK_REASON The cluster describes a new method for processing scientific papers to create training data for language models, detailed in an arXiv preprint. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →