Researchers have developed FigEx2, a novel framework designed to extract and caption information from scientific compound figures. This system addresses the issue of figures lacking captions, which are often discarded by existing pipelines. FigEx2 utilizes visual conditioning to jointly generate panel-specific bounding boxes and descriptive text, improving localization and scientific accuracy through an Entity-Attention KL regularizer and a panel-level Entity-F1 reward. The framework demonstrates strong performance on its curated BioSci-Fig-Cap dataset and outperforms existing models on the MedICaT dataset for captioning, also showing zero-shot transfer capabilities to different scientific domains. AI
IMPACT Enhances the accessibility and utility of scientific literature by enabling automated extraction and captioning of complex figures.
RANK_REASON The cluster contains a research paper detailing a new framework and methodology for processing scientific figures. [lever_c_demoted from research: ic=1 ai=1.0]
- BioSci-Fig-Cap
- Entity-Attention Kullback-Leibler (KL) regularizer
- Entity-F1
- FigEx2
- Group Relative Policy Optimization (GRPO)
- Jifeng Song
- PubMed Central
- Qwen3 VL 8B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →