LLMLingua-2
PulseAugur coverage of LLMLingua-2 — every cluster mentioning LLMLingua-2 across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New tools and research tackle GPU optimization for AI workloads
Several research papers and a new open-source tool address challenges in optimizing AI workloads on GPUs. COMPASS-ABS aims to reduce fragmentation in shared GPU clusters for deep learning training, improving resource ut…
-
Prompt compression struggles with non-English languages, study finds
A new study published on arXiv titled "Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors" investigates the effectiveness of prompt compression techniques across different languages. …
-
New research quantifies Cyrillic tokenization overhead in AI systems
A new research paper titled "Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems" quantifies the significant tokenization overhead for underrepresented Cyrillic-script languages like Ukrainian comp…
-
Edge RAG systems can save energy with adaptive compression, study finds
A new research paper explores adaptive compression techniques for retrieval-augmented generation (RAG) systems operating on edge devices. The study, conducted on an NVIDIA Jetson AGX Thor, demonstrates that dynamically …
-
AI models struggle to predict human attention in text, fusion offers improvement
A new research paper explores the challenge of predicting human attention in text, establishing benchmarks for "floor" (naive truncation) and "ceiling" (split-half oracle) scores. The study found that current frontier l…
-
Prompt compression fails non-English languages, new paper finds
A new paper titled "Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors" reveals that prompt compression techniques, designed to reduce LLM inference costs by removing low-information …
-
Free self-hosted gateway bypasses Claude limits, shares subscriptions
A developer has created a free, self-hosted gateway designed to circumvent usage limits and improve session continuity for users of Anthropic's Claude models. The tool, which has gained significant traction on GitHub, o…
-
New middleware cuts AI coding agent prompt tokens by up to 47%
Researchers have developed a new middleware that optimizes prompts for AI coding agents by preprocessing them on the edge. This system uses a local Llama 3.2 model to translate non-English text to English and rewrite pr…
-
Telegraph English compresses prompts with structured symbols, outperforming LLMLingua-2
Researchers have developed a new prompt compression protocol called Telegraph English (TE), which rewrites natural language into a structured dialect using logical symbols. Unlike methods that delete tokens, TE decompos…