PulseAugur
中
实时 06:55:57
English(EN) RenderRank: Learning to Rerank Text with Compressed Visual Tokens

新方法提高了视觉文档重新排序的效率和准确性

研究人员开发了两种新颖的视觉文档重新排序方法:RenderRank 和 RidgeRank。RenderRank 利用从文档图像中提取的压缩视觉令牌来学习与查询相关的评分,显著减少了输入令牌数量,同时在多个数据集上优于基于文本的重新排序器。RidgeRank 通过融合检索器分数和重新排序器分数并采用浅层线性读出器来提高效率,以一小部分计算成本实现了接近交叉编码器的准确性。这两种方法都旨在提高多模态语言模型在文档检索任务中的重新排序速度和准确性。 AI

影响 这些方法可以显著加快多模态人工智能系统中文档检索和分析的速度。

排序理由 arXiv 上发表了两篇研究论文,详细介绍了视觉文档重新排序的新方法。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新方法提高了视觉文档重新排序的效率和准确性

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv 上发表了两篇研究论文,详细介绍了视觉文档重新排序的新方法。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
2 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Heuiseok Lim ·

    RenderRank:学习使用压缩视觉令牌重新排序文本

    Rendering document text as images allows vision-language models to encode documents as visual tokens, which can reduce input sequence length compared with text input. This reduction in input length is particularly useful for reranking, where each query involves scoring multiple c…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Dongfang Zhao ·

    RidgeRank:通过分数融合和浅层线性读出实现高效视觉文档重排序

    Multimodal language models rerank visual document retrieval results accurately, but scoring every candidate page at full cost makes them slow. Some methods that compress these rerankers need relevance labels to regain accuracy, and they rank by the reranker score alone. RidgeRank…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    RenderRank:学习使用压缩视觉令牌重新排序文本

    Rendering document text as images allows vision-language models to encode documents as visual tokens, which can reduce input sequence length compared with text input. This reduction in input length is particularly useful for reranking, where each query involves scoring multiple c…