mixture of experts
PulseAugur coverage of mixture of experts — every cluster mentioning mixture of experts across labs, papers, and developer communities, ranked by signal.
- instance of alphaXiv 90%
- instance of Gotit.pub 90%
- instance of CatalyzeX 90%
- instance of DagsHub 90%
- instance of large-language models 90%
- instance of Innu-aimun 90%
- instance of Apache Software License 2.0 90%
- instance of Mixture of Experts (MoE) 90%
- instance of Kimi Delta Attention 90%
- instance of GPT OSS 20B 90%
- instance of Mixtral 8x7B 90%
- used by FreeToken 90%
- 2026-05-11 research_milestone A new paper proposes an enhanced Mixture-of-Experts framework for faster time series forecasting model training. source
19 day(s) with sentiment data
-
New MoE framework enhances time series forecasting with integrated expert losses
Researchers have developed a new Mixture-of-Experts (MoE) framework for time series forecasting that improves training efficiency and predictive performance. This framework integrates expert-specific losses directly int…
-
Colibri engine enables 744B parameter LLMs on desktop via novel weight streaming
A new inference engine called Colibri allows users to run extremely large Mixture-of-Experts (MoE) models, such as those with 744 billion parameters, on standard desktop hardware. Instead of compressing the model to fit…
-
New Gait Recognition Method Uses Mixture of Experts to Handle Occlusions
Researchers have introduced GaitMoE, a novel approach to gait recognition that addresses challenges posed by occlusions in real-world scenarios. This method frames gait recognition as an action detection problem, utiliz…
-
New Colla-Q framework balances MoE expert performance via activation entropy
Researchers have introduced Colla-Q, a novel quantization framework designed to mitigate performance degradation in Mixture-of-Experts (MoE) models. This method utilizes activation entropy to balance the bit allocation …
-
New MoE-JEPA model sets state-of-the-art in synthetic image detection
Researchers have developed MoE-JEPA, a novel dual-stream architecture for detecting synthetic and manipulated images. This model enhances a V-JEPA 2 backbone with a Residual Mixture-of-Experts mechanism and a noise stre…
-
New engine serves 35B MoE models from SSDs on consumer hardware
Researchers have developed a new inference engine called Edge0 that enables large Mixture-of-Experts (MoE) models to run on consumer hardware by efficiently utilizing Solid State Drives (SSDs). The system employs a "pre…
-
SOTER model advances generative foundation models for wearable physiological data
Researchers have developed SOTER, a novel generative foundation model specifically designed for wearable human physiological time-series data. This model addresses the unique challenges of such data, including irregular…
-
Deep learning model enhances small-molecule structure identification with mixed-condition training
Researchers have developed a multimodal deep learning approach to improve the identification of small-molecule structures using spectroscopic data. By incorporating domain knowledge from chemistry and spectroscopy into …
-
New MoME technique enhances LLM efficiency with context-aware memory
Researchers have introduced Mixture of Memory Embeddings (MoME), a novel context-aware memory mechanism designed to enhance the efficiency of large language models. Unlike previous methods that assign a single memory en…
-
Open-source Iris search agent challenges closed-source rivals with advanced context management
AllSpark Research has launched Iris, an open-source search agent that challenges closed-source competitors. Iris utilizes a Mixture-of-Experts architecture and features a 256K context window, with models available under…
-
Mixture of Experts: From 1991 concept to DeepSeek-V3 efficiency
Mixture of Experts (MoE) architecture, first proposed in 1991 by Jacobs et al., offers a solution to the scale vs. cost dilemma in large language models. Unlike dense models where all parameters are activated for every …
-
ExpertHTR framework unifies handwritten text recognition with multi-task learning
Researchers have introduced ExpertHTR, a novel framework designed to unify handwritten text recognition (HTR) across diverse datasets. This system employs multi-task learning and a sparse Mixture-of-Experts architecture…
-
New FWP routing strategy optimizes quantized MoE models
Researchers have developed a novel routing strategy for quantized Mixture-of-Experts (MoE) models, aiming to optimize throughput while managing quality degradation. The new method, called Fragility-Weighted Perplexity (…
-
New methods tackle memory peaks for long-context MoE training
Researchers have developed four novel techniques to address memory limitations in training Mixture-of-Experts (MoE) models with long contexts. These methods, PipelinedLLEP, Ring-DTP, Selective Checkpoint Offload (SCO), …
-
Graph-guided MoE framework enhances multi-modal tumor survival prediction
Researchers have developed a novel graph-guided Mixture of Experts (MoE) framework to improve multi-modal tumor survival prediction. This approach addresses limitations in existing methods by effectively integrating div…
-
VIDRAFT's AX-RAY detects AI causal leakage, powers 700B cybersecurity model
VIDRAFT has developed AX-RAY, an AI safety diagnostic system designed to detect "causal leakage," a flaw where models use shortcuts instead of genuine reasoning. This system is now powering a 700 billion parameter Mixtu…
-
MoE models get smarter pruning, retrieval, and inference efficiency
Researchers are exploring advanced techniques for Mixture-of-Experts (MoE) language models to improve their efficiency and performance. One paper introduces HOPE (Higher-Order Pruning of Experts), a novel pruning object…
-
Cohere releases 218B MoE translation model, North Small Translate
Cohere has quietly released North Small Translate, a 218-billion-parameter Mixture-of-Experts (MoE) model specifically designed for machine translation. This sparse model, with 25 billion active parameters per token, su…
-
DeepSeek V4.1 Flash debuts efficient mixture-of-experts LLM
DeepSeek has unveiled its new DeepSeek V4.1 Flash model, which utilizes a mixture of experts (MoE) architecture to achieve high performance while requiring fewer computational resources. This approach allows the model t…
-
M3-Former uses LLMs and MoE for advanced vessel trajectory prediction
Researchers have introduced M3-Former, a novel multimodal trajectory prediction framework that leverages large language models (LLMs) to improve long-term forecasting of vessel movements. The framework integrates static…