Penn Treebank
PulseAugur coverage of Penn Treebank — every cluster mentioning Penn Treebank across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New research decouples MoE routing and aggregation for better performance
Researchers are exploring new approaches to optimize sparse Mixture-of-Experts (MoE) models, moving beyond traditional methods. One study introduces MOSAIC, a framework that integrates architecture and systems co-design…
-
Adaptive MoE Gating Applied Post-Hoc to Qwen3.6-35B Shows Limited Gains
Researchers have developed a post-hoc adaptive Mixture of Experts (MoE) gating method for the Qwen3.6-35B model, aiming to improve efficiency without retraining. Their approach, implemented as an inference-time patch fo…
-
Kan Extension Transformers unify attention, diffusion, and self-conditioning
Researchers have introduced Kan Extension Transformers (KETs), a new framework that unifies various Transformer implementations under a categorical lens. KETs view Transformer layers as weighted structured extension ope…
-
Energy-Gated Attention enhances Transformer models by prioritizing salient tokens
Researchers have introduced Energy-Gated Attention (EGA), a novel mechanism designed to improve transformer models by focusing on spectrally salient tokens. This approach mimics principles from fluid dynamics, prioritiz…
-
Researchers explore weight decay, in-context learning, and acceleration for Transformer models
Researchers have developed several new methods to improve the efficiency and theoretical understanding of Transformer models. One paper provides a functional-analytic characterization of weight decay, demonstrating its …