A new research paper explores the challenge of optical chemical structure recognition (OCSR) in real-world documents, highlighting a significant gap between performance on synthetic data and actual patent and journal figures. The study fine-tuned various vision-language models (VLMs), including Qwen2.5-VL-7B, InternVL3-8B, and GLM-4.1V-9B, using mixtures of synthetic and real-world data. Results indicate that incorporating labeled real training images is crucial for improving accuracy, with the best configuration achieving a 0.84 exact match on one benchmark. The research also found that the effectiveness of vision-tower adaptation strategies, like LoRA, varies depending on the base VLM used. AI
IMPACT Highlights the importance of diverse, real-world training data for improving AI model performance on specialized tasks like chemical structure recognition.
RANK_REASON Research paper detailing model performance and training data impact on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- CLEF-IP
- GLM-4.1V-9B
- Hugging Face
- InternVL3 8B
- Qwen2.5-VL-7B
- United Overseas Bank
- United States Patent and Trademark Office
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →