Researchers have developed StalePO, a new optimization method designed to improve machine translation systems when trained on outdated preference data. Traditional methods can degrade performance or fail to correct specific errors when using post-edits from older models. StalePO addresses this by ensuring likelihood movement is downward on both responses, anchoring the policy to its base response, and applying KL constraints at the token level. This approach has shown significant gains, improving quality checks by up to 14.9 percentage points on English-to-Hindi translation and 4.6 percentage points on English-to-Turkish translation. AI
IMPACT Improves machine translation quality by enabling effective use of legacy training data.
RANK_REASON The cluster contains a research paper detailing a new method for machine translation. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- DagsHub
- Direct Preference Optimization
- English-to-Hindi
- English-to-Turkish
- Gotit.pub
- Hugging Face
- LLM-as-a-Judge
- ScienceCast
- South Marquesan
- StalePO
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →