PulseAugur
实时 09:32:09
English(EN) Baidu Releases Unlimited OCR, a 3B Model That Keeps the KV Cache Flat for Long-Document Parsing

百度发布Unlimited OCR,具有恒定的KV缓存用于长文档

百度发布了Unlimited OCR,这是一个30亿参数的混合专家模型,专为高效的长文档解析而设计。该模型利用参考滑动窗口注意力(R-SWA)来保持恒定的KV缓存,克服了传统OCR模型在处理长输出时面临的内存和速度限制。这项创新使得Unlimited OCR能够在一个前向传播中处理数十页文档,并在OmniDocBench v1.5等基准测试中取得了最先进的性能。 AI

影响 为长文档OCR设定了新标准,可能加速企业采用AI进行文档处理。

排序理由 主要AI实验室(百度)发布的新模型,采用了新颖的技术方法。

在 MarkTechPost 阅读 →

AI 生成摘要 · Google Gemini · 来自 15 个来源。 我们如何撰写摘要 →

百度发布Unlimited OCR,具有恒定的KV缓存用于长文档

报道来源 [15]

  1. Hugging Face Trending Models TIER_1 (ET) · baidu ·

    baidu/Unlimited-OCR

    image-text-to-text · 47 downloads · 55 likes

  2. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    百度发布Unlimited OCR,一个3B模型,可保持KV缓存扁平化以解析长文档

    <p>Baidu open-sourced Unlimited OCR, a 3B-parameter MoE model that parses dozens of document pages in a single forward pass. Its Reference Sliding Window Attention (R-SWA) holds the KV cache constant, so memory and latency stay flat as output grows. It scores 93.23 on OmniDocBenc…

  3. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    百度发布Unlimited-OCR:恒定KV缓存为长文档带来SOTA性能

    Baidu Unveils Unlimited-OCR: Constant KV Cache Delivers SOTA Performance on Long Documents

  4. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    百度已开源Unlimited OCR,一个30亿参数的AI模型,在解析长文档时能保持KV缓存扁平化。该模型使用参考滑动窗口注意力机制

    Baidu has open-sourced Unlimited OCR, a 3B-parameter AI model that keeps KV cache flat when parsing long documents. The model uses Reference Sliding Window Attention to maintain constant memory regardless of output length, scoring 93.23% on OmniDocBench. https://www. marktechpost…

  5. Lobsters — AI tag TIER_1 English(EN) · github.com via metahost ·

    Unlimited-OCR:单次长时域OCR

    <p><a href="https://lobste.rs/s/5ej4m6/unlimited_ocr_one_shot_long_horizon_ocr">Comments</a></p>

  6. dev.to — LLM tag TIER_1 English(EN) · pueding ·

    百度无界OCR将KV缓存保持在40+页:参考滑动窗口注意力

    <p> </p> <p><strong>What:</strong> The <strong>Unlimited OCR release</strong> (Baidu, arXiv 2606.23050) is a <strong>3-billion-parameter open OCR model</strong> whose decoder replaces standard attention with <strong>Reference Sliding Window Attention (R-SWA)</strong> — the trick …

  7. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    🧠 # 百度发布Unlimited-OCR,一款端到端#OCR模型,旨在解决#LLM驱动OCR的主要局限性之一:d的管理

    🧠 # Baidu ha presentato Unlimited-OCR, un modello # OCR end-to-end pensato per affrontare uno dei limiti principali degli OCR basati su # LLM : la gestione di documenti lunghi. 👉 I dettagli: https://www. linkedin.com/posts/alessiopoma ro_baidu-ocr-llm-ugcPost-7475861096952819712-…

  8. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Unlimited-OCR: One-shot Long-horizon OCR https:// lobste.rs/s/5ej4m6 # ai https:// github.com/baidu/Unlimited-OCR

    Unlimited-OCR: One-shot Long-horizon OCR https:// lobste.rs/s/5ej4m6 # ai https:// github.com/baidu/Unlimited-OCR

  9. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    百度发布Unlimited OCR,一个保持KV缓存扁平化的3B模型,用于长文档解析百度开源了Unlimited OCR,一个30亿参数的MoE模型

    Baidu Releases Unlimited OCR, a 3B Model That Keeps the KV Cache Flat for Long-Document Parsing Baidu open-sourced Unlimited OCR, a 3B-parameter MoE model that parses dozens of document pages in a ... #AI #Paper #Summary #AI #Shorts #Applications #Artificial #Intelligence #Editor…

  10. r/LocalLLaMA TIER_1 English(EN) · /u/Sporeboss ·

    Unlimited-OCR现已登陆ModelScope!一个3.3B的多语言OCR模型,支持单张图片、多页文档和PDF的一次性解析。许可证:MIT

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1ue51uk/unlimitedocr_is_now_on_modelscope_a_33b/"> <img alt="Unlimited-OCR is now on ModelScope! A 3.3B multilingual OCR model for one-shot parsing across single images, multi-page documents, and PDFs. License…

  11. Mastodon — fosstodon.org TIER_1 中文(ZH) · [email protected] ·

    🌘 GitHub - baidu/Unlimited-OCR: 无限OCR时代:拥抱单次长视角分析的革命 ➤ 构建高性能、长文本的工业级OCR分析解决方案 ✤ https://github.com/baidu/Unlimited-OCR 百度开源“Unlimited-OCR”项目,旨在进一步拓展文档分析技术的边界

    🌘 GitHub - baidu/Unlimited-OCR:無限 OCR 時代:迎接單次長視野解析的革命 ➤ 打造高效能、長文本的工業級 OCR 解析方案 ✤ https:// github.com/baidu/Unlimited-OCR 百度開源了「Unlimited-OCR」專案,旨在進一步推進文檔解析技術的邊界。該工具專注於「單次長視野解析」(One-shot Long-horizon Parsing),能夠高效處理單頁與多頁文件的 OCR 需求。該模型不僅支援 Huggingface Transformers 的標準推理,還針對高效能需求提供了…

  12. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🚀 # GitHub 和 # 百度推出“无限制OCR:单次长时域解析”,证明即使是# AI 也会迷失在自己过度复杂的行话迷宫中。🙄 W

    🚀 # GitHub and # Baidu introduce "Unlimited OCR: One-Shot Long-Horizon Parsing," proving that even # AI can get lost in its own overcomplicated jargon maze. 🙄 With promises of "direct agents" and "automate any workflow," it's like they've discovered the fax machine of the digital…

  13. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    百度发布Unlimited-OCR,新模型解决长文档转录中的根本瓶颈。通过引入参考滑动窗口注意力机制

    Baidu has unveiled Unlimited-OCR, a new model that solves a fundamental bottleneck in long-document transcription. By introducing Reference Sliding Window Attention, it compresses memory from linear to constant growth, achieving 93.92 percent on the OmniDocBench benchmark. The 3B…

  14. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Unlimited-OCR: One-shot Long-horizon OCR https://github.com/baidu/Unlimited-OCR # AI # OCR # ComputerVision

    Unlimited-OCR: One-shot Long-horizon OCR https://github.com/baidu/Unlimited-OCR # AI # OCR # ComputerVision

  15. Mastodon — mastodon.social TIER_1 English(EN) · AI_Tech_News_UK ·

    🔥 无限OCR:单次长视域解析 研究人员开发了一种新的OCR(光学字符识别)系统,可以解析长视域文本

    🔥 Unlimited OCR: One-Shot Long-Horizon Parsing Researchers have developed a new OCR (Optical Character Recognition) system that can parse long-horizon text with unprecedented accuracy. This technology has significant implications for document scanning and data extraction, and cou…