n-gram
PulseAugur coverage of n-gram — every cluster mentioning n-gram across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
APEX system optimizes LLM inference with adaptive speculative decoding
Researchers have developed APEX, a novel system designed to enhance the efficiency of large language model inference. APEX employs a learned controller that dynamically adjusts speculative decoding strategies based on t…
-
Developer explores adaptive mixing for CPU-based code completion
A developer has created a code completion tool called pycomplete that runs on a laptop CPU. The tool predicts the next token by blending an n-gram model, a cache, and a small transformer, with a fixed mixing constant of…
-
Researchers explore small AI models using n-gram techniques
A user on the r/LocalLLaMA subreddit is inquiring about the existence of experimental small language models (9 billion parameters or less) that utilize n-gram or engram techniques. The user notes a lack of such models o…
-
AI analyzes speech and language to detect loneliness in older adults
Researchers have developed a multimodal framework to detect loneliness in older adults by analyzing speech and language patterns. The study, which involved 310 older adults, combined linguistic features like psycholingu…
-
New research reveals code watermarks vulnerable to obfuscation; OpenStamp offers robust alternative
Researchers have identified significant vulnerabilities in current methods for watermarking AI-generated code, demonstrating that common code obfuscation techniques can effectively neutralize N-gram-based watermarks. A …
-
AI model architecture innovations beyond transformers discussed
A discussion on Reddit's r/LocalLLaMA forum explores potential architectural innovations in AI models beyond well-known advancements like n-grams and quantization. Users are seeking insights into fundamental changes to …
-
New AI benchmark contamination metric proposed to counter paraphrasing
A new method for evaluating AI model contamination has been proposed, focusing on the n-gram overlap between a benchmark and its training corpus. The current method, which measures exact string matches, can be misleadin…
-
N-gram vs. Experts: Understanding LLM Architecture Trade-offs
A Reddit post explains the difference between n-gram and Mixture of Experts (MoE) architectures in large language models. MoEs are described as performing reasoning tasks by selecting specific feed-forward blocks, while…
-
N-gram tables could enable massive AI models on single servers
A discussion on Reddit explores the potential impact of n-gram tables on the AI landscape. The user speculates that n-gram tables could enable the operation of models with over a trillion parameters on single servers wi…
-
Code completer's ghost text feature fails eval due to self-prediction gap
A developer has created a code completion tool called pycomplete that utilizes a transformer model blended with n-gram models and a cache. While the tool's next-token prediction accuracy is 54.6%, its performance on gen…
-
N-gram models better predict reading time than transformers, study finds
A new paper proposes that traditional n-gram language models may be better predictors of naturalistic reading time than complex transformer models. The research suggests that while transformers excel at next-word predic…
-
DeepSeek's Engram Module Made Tokenizer-Agnostic
Researchers have developed a tokenizer-agnostic engram module for large language models, building upon DeepSeek's original design. The new approach replaces XOR-based hashing with polynomial hashing, creating a joint em…
-
Qwen3.6-27B benchmark reveals DFlash leads speculative decoding speedups
A recent benchmark compared speculative decoding methods across vLLM and SGLang frameworks using the Qwen3.6-27B model on a single RTX PRO 6000 Max-Q GPU. The DFlash method emerged as the most effective, offering speedu…
-
Speculative decoding research boosts LLM inference speed on consumer hardware
Researchers are exploring speculative decoding techniques to accelerate large language model (LLM) inference. Two papers, one from arXiv and another from dev.to, detail methods for improving efficiency on consumer hardw…
-
New speculative decoding methods boost LLM inference speed and efficiency · 6 sources tracked
Researchers have introduced DominoTree, a novel method for speculative decoding that significantly accelerates LLM inference by using a conditional tree-structured approach. This method achieves up to 6.6x speedup on Qw…
-
New Hamm-Grams Algorithm Enhances Malware Detection with Robust Features
Researchers have developed a new algorithm called Hamm-Grams, designed to improve malware detection and classification by creating more robust features than traditional n-grams. These hamm-grams are a type of regular ex…