LongBench-v2
PulseAugur coverage of LongBench-v2 — every cluster mentioning LongBench-v2 across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Referential Dangling: A New Failure Mode in LLM Prompt Compression
A new paper identifies a significant failure mode in hard prompt compression techniques, termed "referential dangling." This occurs when methods designed to reduce context length by selecting high-scoring text segments …
-
New LLM inference techniques target efficiency and edge deployment · 7 sources tracked
Multiple research papers introduce novel techniques to enhance Large Language Model (LLM) inference efficiency. Cascade optimizes serving by managing latency budgets for heterogeneous requests, improving goodput and red…
-
New research tackles LLM KV cache compression for efficient long-context inference · 10 sources tracked
Multiple research papers submitted to arXiv in August 2026 propose novel methods for compressing the key-value (KV) cache in large language models (LLMs) to mitigate memory and bandwidth bottlenecks during long-context …
-
New methods target LLM KV cache compression for efficiency
Researchers are developing advanced techniques to compress the Key-Value (KV) cache in Large Language Models (LLMs), a major contributor to memory costs during inference. New methods like JoLT and FlashJoLT utilize tens…
-
New research explores adaptive LLM evaluation and self-improvement techniques · 10 sources tracked
Researchers are developing new methods to evaluate and improve large language models (LLMs). One approach, ATLAS, uses item response theory to significantly reduce the number of items needed for accurate LLM evaluation,…
-
New framework LongCrafter enhances LLM long-context understanding
Researchers have introduced LongCrafter, a novel framework designed to generate diverse and high-quality data for fine-tuning large language models (LLMs) to improve their long-context understanding. This framework addr…
-
New TTT-NTP method boosts LLM performance on long-context tasks
Researchers have introduced a new method called Test-Time Training with Next-Token Prediction (TTT-NTP) that enhances the performance of pre-trained long-context language models. This technique adapts existing LLM check…
-
New framework guides LLMs to choose between RAG and long-context processing
Researchers have developed a new framework called Pre-Route to help large language models decide whether to use retrieval-augmented generation (RAG) or long-context (LC) processing for document understanding. This proac…
-
Telegraph English compresses prompts with structured symbols, outperforming LLMLingua-2
Researchers have developed a new prompt compression protocol called Telegraph English (TE), which rewrites natural language into a structured dialect using logical symbols. Unlike methods that delete tokens, TE decompos…