Penn Treebank
PulseAugur coverage of Penn Treebank — every cluster mentioning Penn Treebank across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New research offers advanced low-rank compression for LLMs · 3 sources tracked
Three new research papers introduce advanced techniques for compressing large language models (LLMs) using low-rank decomposition. The first paper, 'Per-Matrix Optimality Is Not Enough,' proposes a three-level optimizat…
-
New Riemannian Language Models achieve 2x perplexity improvement
Researchers have introduced Riemannian Language Models (RiLM), a novel approach to parameter-efficient language modeling that eliminates the need for an output matrix. This method leverages geodesic decoding, where cont…
-
New research decouples MoE routing and aggregation for better performance
Researchers are exploring new approaches to optimize sparse Mixture-of-Experts (MoE) models, moving beyond traditional methods. One study introduces MOSAIC, a framework that integrates architecture and systems co-design…
-
Adaptive MoE Gating Applied Post-Hoc to Qwen3.6-35B Shows Limited Gains
Researchers have developed a post-hoc adaptive Mixture of Experts (MoE) gating method for the Qwen3.6-35B model, aiming to improve efficiency without retraining. Their approach, implemented as an inference-time patch fo…
-
Kan Extension Transformers unify attention, diffusion, and self-conditioning
Researchers have introduced Kan Extension Transformers (KETs), a new framework that unifies various Transformer implementations under a categorical lens. KETs view Transformer layers as weighted structured extension ope…
-
Energy-Gated Attention enhances Transformer models by prioritizing salient tokens
Researchers have introduced Energy-Gated Attention (EGA), a novel mechanism designed to improve transformer models by focusing on spectrally salient tokens. This approach mimics principles from fluid dynamics, prioritiz…
-
Researchers explore weight decay, in-context learning, and acceleration for Transformer models
Researchers have developed several new methods to improve the efficiency and theoretical understanding of Transformer models. One paper provides a functional-analytic characterization of weight decay, demonstrating its …