Two new arXiv papers explore methods for selecting and characterizing synthetic data used in large language model training. The first paper, "Training-Aware Target Coverage," introduces a method to identify synthetic data that adds beneficial information without introducing errors, demonstrating its effectiveness on the Qwen2.5-Math-1.5B-Instruct model for mathematical reasoning tasks. The second paper, "Synthetic Data Characterization via Training Dynamics," analyzes synthetic data by studying sample-level learnability across different LLM families and scales, comparing it to human-written data and evaluating data selection strategies. AI
IMPACT These papers offer new methodologies for improving LLM training efficiency and performance through better synthetic data utilization.
RANK_REASON Two academic papers published on arXiv detailing new methods for synthetic data selection and characterization for LLMs.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- ScienceCast
- GSM8K
- Qwen2.5-Math-1.5B-Instruct
- Training-Aware Target Coverage
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →