Researchers have developed Synth-JDoc, a novel dataset designed to improve the optical character recognition (OCR) capabilities of Large Vision Language Models (LVLMs), particularly for Japanese text. This dataset synthesizes document images directly from text using HTML and CSS, incorporating diverse layouts with both vertical and horizontal writing styles. To enhance realism and robustness, the synthesized documents include embedded images generated by text-to-image models and are subjected to noise and degradation filters. Experiments show that models fine-tuned on Synth-JDoc outperform those trained on previous synthetic datasets, significantly improving LVLM performance on reading vertically written Japanese text. AI
IMPACT Enhances LVLM capabilities for processing complex Japanese documents, potentially improving applications requiring accurate OCR.
RANK_REASON The item describes a new dataset and research paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →