Researchers have developed a new multimodal Reinforcement Learning from Human Feedback (RLHF) framework to translate historical Han-Nom manuscripts into modern Vietnamese. This approach leverages both the visual information from manuscript images and aligned Han-Nom text to improve translation quality, addressing challenges like degraded pages and limited parallel data. The framework integrates multiple language models and vision encoders, and experiments showed that Direct Preference Optimization (DPO) outperformed Proximal Policy Optimization (PPO) and Khi-squared Optimization (KTO) in various metrics, including BLEU-4 and BERTScore, demonstrating the effectiveness of preference optimization for low-resource historical translation. AI
IMPACT This research advances multimodal AI capabilities for low-resource historical translation tasks.
RANK_REASON The item describes a research paper detailing a new method for historical manuscript translation. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- bert-base-chinese
- BERTScore
- BLEU-4
- Character Error Rate
- CLIP ViT-L/14@336
- Direct Preference Optimization
- Hugging Face
- KTO
- Proximal Policy Optimization
- ROUGE L Score
- supervised fine-tuning
- T5-Small
- vinai/phobert-base
- word error rate
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →