Two new research papers explore methods to improve the alignment of large language models (LLMs) with human preferences. The first paper, BALIGN, introduces a strategy for selecting preference data to mitigate the "alignment tax," which causes LLMs to forget pre-trained capabilities. BALIGN uses a composite risk score based on factors like reference model log-probability margin and token length differences to filter out problematic data. The second paper, CRPO, addresses the issue of English-centric preference data by proposing a cross-lingual framework that leverages English preferences to improve alignment in other languages. CRPO uses a hierarchical structure and relative ranking of responses to enhance adaptation and performance across multiple languages. AI
IMPACT These methods aim to improve LLM performance and usability by addressing critical alignment challenges, potentially leading to more capable and reliable AI systems.
RANK_REASON Two academic papers published on arXiv detailing novel methods for LLM alignment.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →