LongBench-v2
PulseAugur coverage of LongBench-v2 — every cluster mentioning LongBench-v2 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
KV Cache Compression Research Identifies Temporal Aggregation as Key Factor
Researchers have investigated the impact of temporal aggregation and ranking preservation on decoding-time KV cache compression in large language models. They found that exponential moving average (EMA) aggregation can …
-
New TTT-NTP method boosts LLM performance using next-token prediction
Researchers have introduced a new method called Test-Time Training with Next-Token Prediction (TTT-NTP) that enhances the performance of pre-trained long-context language models. This technique leverages the inherent ne…
-
PolicyLong advances LLM context extension with on-policy data evolution
Researchers have introduced PolicyLong, a novel method for extending the context windows of large language models by dynamically constructing training data. Unlike previous offline methods that use a fixed model to gene…
-
New KV cache compression techniques aim to boost LLM long-context performance
Researchers are developing new methods to compress the key-value (KV) cache in large language models, a major bottleneck for long-context inference. Minima-KV uses a mixed-format approach, storing recent pages in FP8 an…
-
New GC-OPD method aligns LLM teacher preferences with task success
A new method called GC-OPD has been developed to address a discrepancy in on-policy distillation for large language models. This method reconciles the teacher model's token-level likelihood preferences with the actual t…
-
Referential Dangling: A New Failure Mode in LLM Prompt Compression
A new paper identifies a significant failure mode in hard prompt compression techniques, termed "referential dangling." This occurs when methods designed to reduce context length by selecting high-scoring text segments …
-
New research explores LLM efficiency and reasoning improvements
Several research papers explore methods to enhance the efficiency and reliability of large language models (LLMs). Hugging Face's LFM2.5-DSpark demonstrates up to 3.2x faster inference speeds by using speculative decodi…
-
New research tackles LLM KV cache optimization for efficiency · 10 sources tracked
Recent research papers introduce novel techniques to optimize KV cache management in large language models, addressing memory bottlenecks and improving inference efficiency. Methods like vToken, GCache, LinearKV, KVDiag…
-
New methods target LLM KV cache compression for efficiency
Researchers are developing advanced techniques to compress the Key-Value (KV) cache in Large Language Models (LLMs), a major contributor to memory costs during inference. New methods like JoLT and FlashJoLT utilize tens…
-
New research explores adaptive LLM evaluation and self-improvement techniques · 10 sources tracked
Researchers are developing new methods to evaluate and improve large language models (LLMs). One approach, ATLAS, uses item response theory to significantly reduce the number of items needed for accurate LLM evaluation,…
-
New framework LongCrafter enhances LLM long-context understanding
Researchers have introduced LongCrafter, a novel framework designed to generate diverse and high-quality data for fine-tuning large language models (LLMs) to improve their long-context understanding. This framework addr…
-
New TTT-NTP method boosts LLM performance on long-context tasks
Researchers have introduced a new method called Test-Time Training with Next-Token Prediction (TTT-NTP) that enhances the performance of pre-trained long-context language models. This technique adapts existing LLM check…
-
New framework guides LLMs to choose between RAG and long-context processing
Researchers have developed a new framework called Pre-Route to help large language models decide whether to use retrieval-augmented generation (RAG) or long-context (LC) processing for document understanding. This proac…
-
Telegraph English compresses prompts with structured symbols, outperforming LLMLingua-2
Researchers have developed a new prompt compression protocol called Telegraph English (TE), which rewrites natural language into a structured dialect using logical symbols. Unlike methods that delete tokens, TE decompos…