New research explores LLM-driven retrieval enhancements and privacy
ByPulseAugur Editorial·[12 sources]·
Researchers are exploring novel methods to enhance information retrieval using large language models (LLMs). One approach, RARS, focuses on multiresolution relevance for hierarchical retrieval by explicitly allocating relevance across different levels of document structure. Another method, MERGE, employs an ensemble of smaller LLMs to enrich queries before a final synthesis by a larger model, aiming to improve performance on standard benchmarks. Additionally, a training-free technique called RICE demonstrates that LLMs can be prompted to produce effective dense retrieval representations using only in-context examples. PILLAR offers a privacy-preserving approach to retrieval-augmented generation by combining sparse and dense retrieval stages with private information retrieval techniques.
AI
IMPACT
These research papers explore novel techniques for improving information retrieval systems using LLMs, potentially leading to more accurate and efficient search and recommendation capabilities.
RANK_REASON
Multiple arXiv papers introducing new methods and studies in information retrieval.
Generative retrieval with semantic identifiers (SIDs) makes successive decisions over a document hierarchy. Relevant documents for the same query may share coarse prefixes and diverge at finer depths, with branching patterns varying across queries. These paths reveal how relevanc…
arXiv:2609.37574v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limit…
arXiv cs.CL
TIER_1English(EN)·Nour Jedidi, Abdul Basit Ali, Hang Li, Jimmy Lin·
arXiv:2609.38099v1 Announce Type: cross Abstract: Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for…
arXiv cs.AI
TIER_1English(EN)·Truong Son Nguyen (Arizona State University), Daniel Blackley (George Mason University), Ni Trieu (Arizona State University), Evgenios M. Kornaropoulos (George Mason University)·
arXiv:2609.36326v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k docume…
arXiv cs.AI
TIER_1English(EN)·Ryan C. Barron, Cade W. Trotter, Maksim E. Eren, Kim {\O}. Rasmussen, Liz D. Miller, Benjamin J. Migliori·
arXiv:2609.37911v1 Announce Type: cross Abstract: Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronge…
Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Fo…
Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Fo…
Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examp…
arXiv cs.IR (Information Retrieval)
TIER_1English(EN)·Benjamin J. Migliori·
Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term list…
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, …
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, …
Cross-modal retrieval has been advanced by vision-language pre-trained models that encode images and texts into a shared dense embedding space. While dense representations effectively capture overall semantic similarity, they often obscure fine-grained visual-textual information …