PulseAugur
中
实时 22:51:30
English(EN) Effective Dense Retrieval using Only In-Context Examples

新研究探讨 LLM 驱动的检索增强和隐私

研究人员正在探索使用大型语言模型 (LLM) 增强信息检索的新颖方法。一种名为 RARS 的方法通过在文档结构的不同级别上显式分配相关性,专注于分层检索的多分辨率相关性。另一种名为 MERGE 的方法采用小型 LLM 的集成来丰富查询,然后由大型模型进行最终综合,旨在提高在标准基准测试上的性能。此外,一种名为 RICE 的免训练技术表明,LLM 可以通过仅使用上下文示例来提示生成有效的密集检索表示。PILLAR 通过将稀疏和密集检索阶段与私有信息检索技术相结合,提供了一种隐私保护的检索增强生成方法。 AI

影响 这些研究论文探讨了使用 LLM 改进信息检索系统的技术,有望带来更准确、更高效的搜索和推荐能力。

排序理由 多篇 arXiv 论文介绍了信息检索领域的新方法和研究。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 12 个来源。 我们如何撰写摘要 →

新研究探讨 LLM 驱动的检索增强和隐私

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇 arXiv 论文介绍了信息检索领域的新方法和研究。
Source corroboration
12 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
12 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [12]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Fuzhen Zhuang ·

    学习多分辨率相关性以实现分层生成检索

    Generative retrieval with semantic identifiers (SIDs) makes successive decisions over a document hierarchy. Relevant documents for the same query may share coarse prefixes and diverge at finer depths, with branching patterns varying across queries. These paths reveal how relevanc…

  2. arXiv cs.CL TIER_1 English(EN) · Tzu-I Ho, Yung-Yu Shih, Shang-Yu Su, Dongzhe Wang, Yun-Nung Chen ·

    MERGE:通过生成式增强实现多 LLM 集成检索

    arXiv:2609.37574v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limit…

  3. arXiv cs.CL TIER_1 English(EN) · Nour Jedidi, Abdul Basit Ali, Hang Li, Jimmy Lin ·

    仅使用上下文示例实现有效的密集检索

    arXiv:2609.38099v1 Announce Type: cross Abstract: Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for…

  4. arXiv cs.AI TIER_1 English(EN) · Truong Son Nguyen (Arizona State University), Daniel Blackley (George Mason University), Ni Trieu (Arizona State University), Evgenios M. Kornaropoulos (George Mason University) ·

    PILLAR:用于增强检索的私有倒排索引词汇查找

    arXiv:2609.36326v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k docume…

  5. arXiv cs.AI TIER_1 English(EN) · Ryan C. Barron, Cade W. Trotter, Maksim E. Eren, Kim {\O}. Rasmussen, Liz D. Miller, Benjamin J. Migliori ·

    生成的查询扩展仍有助于强大的稀疏检索:一项使用 SPLADE-v3 的对照研究

    arXiv:2609.37911v1 Announce Type: cross Abstract: Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronge…

  6. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Shengbo Guo ·

    利用生成模型探索论坛帖子检索

    Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Fo…

  7. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Rish Tandon ·

    使用生成模型探索论坛帖子检索

    Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Fo…

  8. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jimmy Lin ·

    仅使用上下文内示例实现有效的密集检索

    Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examp…

  9. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Benjamin J. Migliori ·

    生成查询扩展仍有助于强大的稀疏检索:一项使用 SPLADE-v3 的对照研究

    Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term list…

  10. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yun-Nung Chen ·

    MERGE:通过生成式增强进行检索的多LLM集成

    Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, …

  11. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yun-Nung Chen ·

    MERGE:通过生成式增强实现多 LLM 集成检索

    Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, …

  12. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · SeongKu Kang ·

    概念补充稠密语义:学习紧凑稀疏空间用于文本图像检索

    Cross-modal retrieval has been advanced by vision-language pre-trained models that encode images and texts into a shared dense embedding space. While dense representations effectively capture overall semantic similarity, they often obscure fine-grained visual-textual information …