English(EN)Baidu Releases Unlimited OCR, a 3B Model That Keeps the KV Cache Flat for Long-Document Parsing
百度发布Unlimited OCR,具有恒定的KV缓存用于长文档
作者PulseAugur 编辑部·[15 个来源]·
百度发布了Unlimited OCR,这是一个30亿参数的混合专家模型,专为高效的长文档解析而设计。该模型利用参考滑动窗口注意力(R-SWA)来保持恒定的KV缓存,克服了传统OCR模型在处理长输出时面临的内存和速度限制。这项创新使得Unlimited OCR能够在一个前向传播中处理数十页文档,并在OmniDocBench v1.5等基准测试中取得了最先进的性能。
AI
<p>Baidu open-sourced Unlimited OCR, a 3B-parameter MoE model that parses dozens of document pages in a single forward pass. Its Reference Sliding Window Attention (R-SWA) holds the KV cache constant, so memory and latency stay flat as output grows. It scores 93.23 on OmniDocBenc…
Baidu has open-sourced Unlimited OCR, a 3B-parameter AI model that keeps KV cache flat when parsing long documents. The model uses Reference Sliding Window Attention to maintain constant memory regardless of output length, scoring 93.23% on OmniDocBench. https://www. marktechpost…
Lobsters — AI tag
TIER_1English(EN)·github.com via metahost·
<p> </p> <p><strong>What:</strong> The <strong>Unlimited OCR release</strong> (Baidu, arXiv 2606.23050) is a <strong>3-billion-parameter open OCR model</strong> whose decoder replaces standard attention with <strong>Reference Sliding Window Attention (R-SWA)</strong> — the trick …
🧠 # Baidu ha presentato Unlimited-OCR, un modello # OCR end-to-end pensato per affrontare uno dei limiti principali degli OCR basati su # LLM : la gestione di documenti lunghi. 👉 I dettagli: https://www. linkedin.com/posts/alessiopoma ro_baidu-ocr-llm-ugcPost-7475861096952819712-…
Baidu Releases Unlimited OCR, a 3B Model That Keeps the KV Cache Flat for Long-Document Parsing Baidu open-sourced Unlimited OCR, a 3B-parameter MoE model that parses dozens of document pages in a ... #AI #Paper #Summary #AI #Shorts #Applications #Artificial #Intelligence #Editor…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1ue51uk/unlimitedocr_is_now_on_modelscope_a_33b/"> <img alt="Unlimited-OCR is now on ModelScope! A 3.3B multilingual OCR model for one-shot parsing across single images, multi-page documents, and PDFs. License…
🚀 # GitHub and # Baidu introduce "Unlimited OCR: One-Shot Long-Horizon Parsing," proving that even # AI can get lost in its own overcomplicated jargon maze. 🙄 With promises of "direct agents" and "automate any workflow," it's like they've discovered the fax machine of the digital…
Baidu has unveiled Unlimited-OCR, a new model that solves a fundamental bottleneck in long-document transcription. By introducing Reference Sliding Window Attention, it compresses memory from linear to constant growth, achieving 93.92 percent on the OmniDocBench benchmark. The 3B…
🔥 Unlimited OCR: One-Shot Long-Horizon Parsing Researchers have developed a new OCR (Optical Character Recognition) system that can parse long-horizon text with unprecedented accuracy. This technology has significant implications for document scanning and data extraction, and cou…