lambada
PulseAugur coverage of lambada — every cluster mentioning lambada across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Transformer models use 'off-axis' computation for concepts, study finds
A new research paper explores the internal workings of Transformer models, revealing that their intermediate states are not random noise but rather a functional component for computing concepts. The study found that a 1…
-
New linearized attention model achieves higher accuracy and lower perplexity
Researchers have developed a linearized version of 2-simplicial attention, which rewrites the trilinear score into an inner product. This new form allows for linear cost in sequence length while maintaining global reach…
-
New SCSE method improves Looped Transformers for text tasks
Researchers have introduced Source-Centered State Evolution (SCSE), a novel method designed to enhance Looped Transformers. SCSE addresses the challenge of maintaining consistent hidden states across varying recurrent d…
-
HOLA enhances linear attention models with a complementary memory system
Researchers have developed a novel approach called HOLA (Hippocampal Linear Attention) to enhance the memory capabilities of linear attention and state-space language models. This method introduces a complementary 'hipp…
-
New HOLA architecture enhances linear attention language models with dual memory system
Researchers have developed HOLA (Hippocampal Linear Attention), a novel architecture that enhances linear attention language models by incorporating a complementary memory system. This system addresses the issue of info…
-
New NC-FFN architecture enhances transformer interpretability and efficiency
Researchers have developed a novel parameter-neutral replacement for transformer feed-forward networks, termed NC-FFN, which utilizes explicit fuzzy set operations. This new architecture demonstrates strong parameter ef…