Researchers have developed Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a new Japanese document understanding model. This model is built upon Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 and focuses on structured document parsing through capability injection and forgetting control techniques. Experiments explored different training methods, including parsing-centric supervised fine-tuning (SFT), mixed SFT combining parsing and VQA data, and parsing-centric reinforcement learning (RL) using a DAPO-based approach. The mixed SFT method effectively balanced structured parsing performance with VQA capability preservation, while RL further enhanced parsing accuracy. AI
IMPACT This model advances structured document parsing capabilities for Japanese, potentially improving information extraction and analysis in specific contexts.
RANK_REASON The cluster describes a new research paper detailing the development of a specific language model for document understanding. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
- Stockmark-Nemotron-3-Nano-Omni-JapanDocReader
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →