OLMoE-1B-7B
PulseAugur coverage of OLMoE-1B-7B — every cluster mentioning OLMoE-1B-7B across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New research identifies and solves subspace contention in MoE+LoRA fine-tuning
A new research paper titled "Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning" explores the limitations of combining Mixture-of-Experts (MoE) routing with Low-Rank Adaptation (…
-
Instella-MoE: New open-source MoE language model released
A new technical report introduces Instella-MoE, an open-source Mixture-of-Experts (MoE) language model with 16 billion total parameters. Trained on AMD Instinct GPUs, the model incorporates innovations like Gated Multi-…
-
MoE Models Show Fragile Moral Encoding Despite Redundant Representations
A new arXiv paper titled "Output Dilution: Redundant but Fragile Representations in MoE Models" investigates the encoding of moral content in Mixture-of-Experts (MoE) models. Researchers found that while MoE models like…
-
New research quantifies quantization damage in Mixture-of-Experts models
A new research paper explores the impact of quantization on Mixture-of-Experts (MoE) models, specifically focusing on how numerical disturbances can cause route flips. The study proposes a method to quantify this route-…
-
New research decouples MoE routing and aggregation for better performance
Researchers are exploring new approaches to optimize sparse Mixture-of-Experts (MoE) models, moving beyond traditional methods. One study introduces MOSAIC, a framework that integrates architecture and systems co-design…
-
AMD releases open Instella-MoE-16B LLM with 2.8B active parameters
AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model. This model features 16 billion total parameters but only activates 2.8 billion per token, utilizing architectural innovations like…
-
MoE models show mixed inference performance on consumer and edge hardware
A recent study investigated whether Mixture-of-Experts (MoE) language models offer practical inference advantages on consumer and edge hardware. The research found that while MoE models theoretically reduce per-token co…
-
New method validates LLM circuits using ablation tests
Researchers have developed a new method for discovering circuits within large language models by clustering attention head co-activation statistics. This approach, termed "closure-validated circuit discovery," uses caus…
-
Regret Pre-training boosts language model knowledge grounding
Researchers have developed a new self-supervised learning framework called Regret Pre-training to improve causal language models. This method leverages future information typically unavailable during standard causal tra…
-
MobileMoE models set new efficiency standard for on-device LLMs
Researchers have introduced MobileMoE, a new family of on-device Mixture-of-Experts (MoE) language models designed for mobile deployment. These models, with sub-billion active parameters, establish a new performance fro…
-
MoE models misroute tokens on complex reasoning tasks, study finds
Researchers have identified a significant issue in Mixture-of-Experts (MoE) language models where the routing mechanism, which directs tokens to specific experts, often selects suboptimal paths. While the standard route…