New research probes LLM long-context retrieval and RAG strategies · 9 sources tracked
ByPulseAugur Editorial·[12 sources]·
Recent research explores how large language models (LLMs) handle long contexts, with studies investigating the mechanisms behind improved performance. One paper examines Chain-of-Thought (CoT) reasoning, finding that it enables targeted retrieval and more compact representations compared to broad retrieval. Another study analyzes how different positional encoding choices, like RoPE and sliding-window attention, shift models from positional to semantic retrieval, impacting performance on tasks like question answering. Additionally, research compares various retrieval-augmented generation (RAG) strategies, highlighting the importance of re-ranking and late interaction for scientific question answering, and introduces new frameworks for learning retrieval actions and self-evaluative exploration to enhance knowledge retrieval.
AI
IMPACT
Advances in retrieval and context handling are crucial for improving LLM performance on complex, long-horizon tasks and domain-specific knowledge.
RANK_REASON
Multiple arXiv papers detailing novel research into LLM retrieval mechanisms and RAG strategies.
arXiv:2609.38958v1 Announce Type: new Abstract: Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanism…
arXiv:2609.38530v1 Announce Type: new Abstract: Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention (SWA NoPE). However, how these choices…
arXiv cs.AI
TIER_1English(EN)·Bhagyesh Rathi, Eshan Chawla, William B. Andreopoulos·
arXiv:2609.38473v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is large and the trade-offs between variants are not well under…
arXiv cs.CL
TIER_1English(EN)·Jingyuan Ma, Lynx Aster, He Zhang, Siyao Song, Weijie Yuan, Zhe Zhang, Kai Jia, Zhifang Sui·
arXiv:2609.37082v1 Announce Type: new Abstract: Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages…
arXiv cs.IR (Information Retrieval)
TIER_1English(EN)·William B. Andreopoulos·
Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is large and the trade-offs between variants are not well understood, especially on domain-specific corpora at re…
arXiv:2609.28653v1 Announce Type: cross Abstract: Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can …
LLM-based retrievers and rerankers have advanced passage ranking, yet both paradigms interact with the corpus in a single pass and commit to the resulting candidate set, leaving relevant documents permanently unrecoverable once missed. We introduce Seek, Self-Evaluative Explorati…
Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can improve small language models (SLMs) as next-actio…
<p>Anthropic’s own benchmark tells an uncomfortable story about dumping everything into the prompt: even with a generous context window, standard retrieval still missed the right chunk 5.7% of the time, and fixing that took a second retrieval technique, not a bigger window [1]. T…
<h4>What a controlled retrieval experiments revealed about ranking, evidence, latency, and failure modes.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*LIOvevb1bkfEFw5HWs6gGg.png" /></figure><p>Most RAG demos are deceptively simple.</p><p>Ingest a docume…
dev.to — LLM tag
TIER_1English(EN)·Ruchita Nimkar·
<h2> What happens when a question looks simple, but answering it correctly requires more than retrieving a few documents? </h2> <p>For my <strong>TigerGraph Hackathon</strong> project, I explored this question by implementing and comparing <strong>RAG, GraphRAG, and Agentic Graph…