PulseAugur
EN
LIVE 10:49:29
ENTITY WikiText-103

WikiText-103

PulseAugur coverage of WikiText-103 — every cluster mentioning WikiText-103 across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
12
27 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
12
27 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

11 day(s) with sentiment data

RECENT · PAGE 1/2 · 27 TOTAL
  1. TOOL · CL_196106 ·

    New CurveFP datatypes promise lower cost and better performance for language models

    Researchers have introduced CurveFP, a novel family of low-precision datatypes designed to reduce the cost of language models. CurveFP optimizes scalar fidelity and the arithmetic induced by products through a closed-pr…

  2. RESEARCH · CL_193811 ·

    New research decouples MoE routing and aggregation for better performance

    Researchers are exploring new approaches to optimize sparse Mixture-of-Experts (MoE) models, moving beyond traditional methods. One study introduces MOSAIC, a framework that integrates architecture and systems co-design…

  3. TOOL · CL_185364 ·

    Attention-free architecture shows promise in language generation and scaling

    A new research paper, "Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention," explores an attention-free architecture for language modeling. The study demonstrates that this architecture can mat…

  4. TOOL · CL_180526 ·

    TextNCA: Neural Cellular Automata for Language Modeling Explored

    Researchers have developed TextNCA, a language model based on Neural Cellular Automata that utilizes hierarchical local attention. While not outperforming a similarly sized Transformer model on the WikiText-103 benchmar…

  5. TOOL · CL_174100 ·

    New SCSE method improves Looped Transformers for text tasks

    Researchers have introduced Source-Centered State Evolution (SCSE), a novel method designed to enhance Looped Transformers. SCSE addresses the challenge of maintaining consistent hidden states across varying recurrent d…

  6. RESEARCH · CL_171882 ·

    New research explores efficient few-step generation for text and images

    Two new research papers introduce novel approaches to generative modeling, focusing on improving the efficiency and quality of few-step generation for text and images. The first paper, "Latent-Kernel Discrete Flow Maps …

  7. RESEARCH · CL_169569 ·

    New research tackles evaluation and architecture for masked diffusion language models

    Two new research papers introduce novel evaluation protocols and architectures for masked diffusion language models (MDLMs). The first paper, "CaRE," proposes a compute-aware framework to standardize evaluations, reveal…

  8. RESEARCH · CL_160678 ·

    Naju model introduces decoupled retention and writing for long-sequence memory

    Researchers have introduced Naju, a novel native discrete state-space model designed for enhanced long-sequence memory. Unlike previous models that struggle to balance retention and overwriting of information, Naju deco…

  9. TOOL · CL_156552 ·

    SHUFFLESPARSE learned permutations boost structured sparse network accuracy

    Researchers have developed SHUFFLESPARSE, a novel permutation primitive designed to enhance structured weight sparsity in neural networks. This method aims to close the accuracy gap between structured and unstructured s…

  10. TOOL · CL_156418 ·

    New framework tackles MoE LLM trilemma with dynamic clustering and compression

    Researchers have developed a new framework to address the trilemma faced by Mixture-of-Experts (MoE) Large Language Models (LLMs), which involves load imbalance, parameter redundancy, and communication overhead. Their m…

  11. TOOL · CL_154002 ·

    New Split-FG method trains deep networks with 35% less memory

    Researchers have developed a novel method called Split Forward Gradient (Split-FG) to train deep neural networks more efficiently by reducing memory usage. This technique splits a network into a trunk and an output head…

  12. RESEARCH · CL_147469 ·

    New geometric framework for function-preserving continual learning unveiled

    Researchers have introduced a novel function-preserving operator for continual learning called gate-zero growth. This method adds new residual blocks via a zero-initialized gate, which, under specific conditions, leads …

  13. TOOL · CL_111778 ·

    New tPC-RTRL method learns long-range dependencies in recurrent systems

    Researchers have developed a novel method called Temporal Predictive Coding combined with Real-Time Recurrent Learning (tPC-RTRL) to enhance the learning capabilities of recurrent neural networks. This approach addresse…

  14. TOOL · CL_98119 ·

    Gaussian Mixture Attention offers linear-time sequence mixing

    Researchers have introduced Gaussian Mixture Attention (GMA), a novel sequence mixing technique designed to overcome the quadratic scaling bottleneck of standard Transformer attention. GMA replaces explicit token-to-tok…

  15. TOOL · CL_93842 ·

    New IGLU activation function offers improved gradient flow

    Researchers have introduced IGLU, a novel parametric activation function for deep neural networks designed to improve gradient flow and optimization stability. Derived from a mixture of GELU gates under a half-normal di…

  16. TOOL · CL_93350 ·

    New Hybrid Architecture Boosts Long-Context Language Model Efficiency

    Researchers have introduced a Parallel Hybrid Architecture (PHA) that combines Gated State Spaces (GSS), Grouped Query Attention (GQA), and Feed-Forward Networks (FFNs) to improve long-context language modeling. This ar…

  17. RESEARCH · CL_90780 ·

    New RAG and Long-Context Models Leverage Knowledge Graphs

    Two new research papers introduce advanced methods for improving retrieval-augmented generation (RAG) and long-context language modeling. The first paper, "A Unified Framework for Context-Aware and Relation-Aware Graph …

  18. TOOL · CL_87135 ·

    LongSpike: New SNN Framework Enhances Long Sequence Learning

    Researchers have introduced LongSpike, a new Spiking Neural Network (SNN) framework that utilizes fractional-order State-Space Modeling (f-SSM) to enhance the learning of long sequences. This approach overcomes the limi…

  19. RESEARCH · CL_79133 ·

    Chiaroscuro Attention optimizes transformer compute with dynamic token routing

    Researchers have developed CHIAR-Former, a novel 4-layer transformer model that optimizes compute usage by dynamically routing tokens. Instead of applying self-attention uniformly, CHIAR-Former analyzes token spectral e…

  20. RESEARCH · CL_53609 ·

    Kan Extension Transformers unify attention, diffusion, and self-conditioning

    Researchers have introduced Kan Extension Transformers (KETs), a new framework that unifies various Transformer implementations under a categorical lens. KETs view Transformer layers as weighted structured extension ope…