MoE models
PulseAugur coverage of MoE models — every cluster mentioning MoE models across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Noesis architecture enhances Graph-RAG with adaptive parallelism and cross-KB routing · 2 sources tracked
Researchers have introduced Noesis, a novel Graph-RAG architecture designed to overcome limitations in grounding large language models with domain-specific knowledge. Noesis employs four key algorithms: Bidirectional Gr…
-
Lumabri enables P2P MoE model execution; Google Meet adds in-person meeting transcription
Lumabri is a new open-source project that enables users to run Mixture of Experts (MoE) models on a peer-to-peer network using the Colibri protocol. The project, developed by JustVugg, aims to decentralize AI model exec…
-
llama.cpp, PyTorch, and new MoE model see significant updates
The llama.cpp project has released updates enhancing WebGPU acceleration and simplifying FlashAttention implementation for more efficient local LLM inference. Concurrently, PyTorch's MPSInductor now supports unsigned in…
-
New research details physics of multimodal pretraining, efficiency recipes
Researchers have explored the fundamental mechanisms and design space of multimodal pretraining, focusing on how different modalities interact during unified training. Their experiments reveal insights into knowledge fl…
-
xHC method expands transformer streams beyond N=4 for improved LLM pre-training
Researchers have introduced xHC (Expanded Hyper-Connections), a novel method for scaling transformer models beyond the typical limit of N=4 streams. This new approach addresses bottlenecks in previous Hyper-Connections …
-
Budget GPU advice sought for local LLM inference
A user on the r/LocalLLaMA subreddit is seeking advice on purchasing hardware for running large language models on a limited budget. They are considering either a Radeon VII with 32GB VRAM or two P100 GPUs offering a co…