Full Attention
PulseAugur coverage of Full Attention — every cluster mentioning Full Attention across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New Autonomy-of-Heads method boosts LLM efficiency without data
Researchers have developed a novel data-free method called Autonomy-of-Heads (AoH) to improve the efficiency of long-context Large Language Models. AoH identifies retrieval and streaming heads by analyzing the spectral …
-
Bole system accelerates hybrid-attention LLM inference with tree speculation
Researchers have developed Bole, a new system designed to accelerate inference for hybrid-attention large language models. These models combine full attention with recurrent linear attention to manage long contexts more…
-
Xiaomi details MiMo-V2.5 AI model efficiency optimizations
Xiaomi has detailed the engineering optimizations behind its MiMo-V2.5 series of AI models, focusing on achieving efficiency for long-context reasoning and multimodal tasks. The models employ Hybrid Sliding Window Atten…
-
New attention mechanisms boost LLM efficiency and reduce hallucination · 10 sources tracked
Researchers are developing novel attention mechanisms to improve the efficiency and capabilities of large language models (LLMs) and multimodal large language models (MLLMs). These advancements focus on optimizing spars…
-
HydraHead architecture fuses attention types for improved long-context LLMs
Researchers have introduced HydraHead, a novel architecture that hybridizes Full Attention and Linear Attention at the head level within transformer models. This approach leverages interpretability to identify critical …
-
New research explores hybrid and sparse attention mechanisms for LLMs
Researchers are exploring novel methods to optimize attention mechanisms in large language models, particularly for handling long contexts. The HydraHead architecture, for instance, hybridizes Full Attention (FA) and Li…