Researchers have introduced DataOrchestra, a novel framework designed to optimize the pretraining data processing for large-language models (LLMs). Unlike existing methods that apply uniform processing strategies, DataOrchestra dynamically orchestrates a per-example pipeline, deciding whether to drop, keep, or clean each data chunk. For cleaning, it selects specific operations and generates instructions for tool models, leading to stable performance gains across multiple benchmarks when models are pretrained from scratch. This approach also proves effective for specialized tasks like math continued pretraining, outperforming stronger baselines while reducing computational costs by skipping unnecessary processing steps. AI
IMPACT This framework could lead to more efficient and effective LLM training by tailoring data processing to individual examples.
RANK_REASON The cluster contains an academic paper detailing a new framework for data processing in LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →