mixture of experts
PulseAugur coverage of mixture of experts — every cluster mentioning mixture of experts across labs, papers, and developer communities, ranked by signal.
- instance of DagsHub 90%
- instance of ScienceCast 90%
- instance of alphaXiv 90%
- instance of large-language models 90%
- instance of Innu-aimun 90%
- instance of Apache Software License 2.0 90%
- instance of GPT OSS 20B 90%
- instance of Mixtral 8x7B 90%
- uses Kimi Delta Attention 90%
- instance of Llama 4 90%
- instance of Kimi Delta Attention 90%
- instance of OLMoE-1B-7B 90%
- 2026-05-11 research_milestone A new paper proposes an enhanced Mixture-of-Experts framework for faster time series forecasting model training. source
27 day(s) with sentiment data
-
Mixture of Experts model boosts image compression efficiency
Researchers have developed a novel Mixture of Experts (MoE) based entropy model, termed MoEE, for learned image compression. This approach allows the model to selectively activate only the necessary parameters for a giv…
-
MammoMix uses Mixture-of-Experts for robust mammogram breast detection
Researchers have developed MammoMix, a new framework utilizing the Mixture-of-Experts (MoE) paradigm to improve the detection of breast lesions in mammograms. This approach trains individual expert models on specific da…
-
UniF-MoE framework unifies adaptive MoE computation for improved efficiency
Researchers have introduced UniF-MoE, a novel framework for Mixture-of-Experts (MoE) computation that unifies various adaptive strategies. This approach decomposes experts into blocks, allowing for shared computation fi…
-
New research decouples MoE routing and aggregation for better performance
Researchers are exploring new approaches to optimize sparse Mixture-of-Experts (MoE) models, moving beyond traditional methods. One study introduces MOSAIC, a framework that integrates architecture and systems co-design…
-
DistMoE enables rehearsal-free distributed tuning for multimodal LLMs
Researchers have introduced DistMoE, a novel mixture-of-experts approach designed for distributed visual instruction tuning of Multimodal Large Language Models (MLLMs). This method augments standard feedforward networks…
-
PressureMesh system estimates 3D human poses using multi-device pressure data
Researchers have developed PressureMesh, a novel system for estimating 3D human poses using pressure data from multiple devices. The system utilizes an end-to-end network called MDP-Net, which incorporates a Mixture of …
-
AdapterMoE architecture improves crop disease recognition efficiency
Researchers have developed AdapterMoE, a novel two-stage hard-routing Mixture-of-Experts architecture designed for multi-crop disease recognition. This system aims to improve efficiency and flexibility by using a Router…
-
New AI model uses sparse routing for retinal pathology analysis
Researchers have developed a new deep learning architecture for analyzing retinal fundus images that utilizes sparse conditional computation. This model pairs a Guided Context Gating (GCG) spatial attention front-end wi…
-
New Wiener Filtering Technique Reduces Hallucinations in Vision-Language Models
Researchers have developed a novel technique called Wiener Representation Filtering to reduce hallucinations in vision-language models (VLMs). This training-free method operates post-hoc by editing the representation sp…
-
MoE expert caching evaluation methods found to be misleading
Researchers have identified critical flaws in trace-driven evaluation methods for Mixture-of-Experts (MoE) models, which can lead to misleading conclusions about expert caching policies. The study highlights how replay …
-
MoE-Prism framework enhances LLM serving with elastic expert routing
Researchers have developed MoE-Prism, a framework designed to enhance the efficiency of Mixture-of-Experts (MoE) models in serving large language models (LLMs). This system allows for request-level compute elasticity by…
-
Lightweight fine-tuning prunes MoE models, reducing size and latency
Researchers have developed a method to prune experts in Mixture-of-Experts (MoE) models using lightweight fine-tuning techniques. By applying parameter-efficient adapters like LoRA, they can identify and remove less cri…
-
New LorExperts and BTExperts methods compress MoE models effectively
Researchers have developed two new methods, LorExperts and BTExperts, for compressing Mixture-of-Experts (MoE) language models. These techniques aim to reduce the computational cost of deploying MoE models by compressin…
-
New CoCo method enhances interpretability of AI reward models
Researchers have introduced a new method called Contribution-Contrast (CoCo) to improve the interpretability of Mixture-of-Experts (MoE) reward models. Unlike previous methods that focused on routing weights, CoCo analy…
-
New TEXAS method enhances Mixture-of-Experts LLM adaptation
Researchers have developed a new method called TEXAS (Task-Expert-Aware Supervision) to improve the adaptation of Mixture-of-Experts (MoE) large language models. This technique identifies task-relevant experts by compar…
-
New EntropyMoE architecture optimizes tokenizer-free LLMs with sparse expert routing
Researchers have introduced EntropyMoE, a novel Mixture-of-Experts (MoE) architecture designed for tokenizer-free large language models that process data in dynamic byte patches. Unlike existing models that apply unifor…
-
inclusionAI releases lightweight Ling-3.0-tiny MoE model for local deployment
inclusionAI has released Ling-3.0-tiny, a new hybrid reasoning Mixture-of-Experts (MoE) model with 7.9 billion total parameters and 1.3 billion activated parameters per token. This model is designed for efficient local …
-
New tools enable LLM fine-tuning on low-spec hardware
New tools and techniques are emerging to enable fine-tuning and running large language models (LLMs) on consumer-grade hardware. Soup CLI, an open-source Python tool, utilizes layer streaming to fine-tune an 8B LLM on a…
-
New SkillMemo framework enhances robotic manipulation generalization
Researchers have developed SkillMemo, a novel framework designed to improve the compositional generalization of embodied visuomotor models in robotics. This framework addresses the limitations of current models, which a…
-
AI application offers personalized, tax-aware investment advice
Researchers have developed a novel application for retail portfolio management that utilizes a three-phase reinforcement learning system. This system aims to provide personalized, tax-aware investment recommendations by…