Researchers have developed a novel two-stage training framework for multimodal disaster severity assessment that integrates Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). This approach utilizes a unified data construction pipeline to derive two datasets: ReasoningSet for SFT and PreferenceSet for DPO-based alignment. Experiments show significant improvements in both classification accuracy and explanation quality, with subsequent DPO alignment further enhancing interpretability. The framework's robustness was demonstrated through cross-model validation on InternVL-3-8B and LLaVA-1.5-7B, leading to better detection of underrepresented damage cases and stronger alignment between model reasoning and human judgment. AI
IMPACT Enhances the reliability and interpretability of AI systems for critical applications like disaster management.
RANK_REASON The cluster contains an academic paper detailing a new methodology and experimental results. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- InternVL-3-8B
- LLaVA-1.5-7B
- PreferenceSet
- ReasoningSet
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →