Mixture-of-Transformers
PulseAugur coverage of Mixture-of-Transformers — every cluster mentioning Mixture-of-Transformers across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New framework uses generation to boost MLLM visual understanding
Researchers have introduced GAS, a novel training framework that leverages generation as auxiliary supervision to enhance visual understanding in multimodal large language models (MLLMs). This approach adapts Next Embed…
-
Flex-π model integrates 3D geometry and object semantics with RGB data
Researchers have developed Flex-$\pi$, a 6-billion parameter world-action model that integrates 3D geometry and object semantics alongside RGB data. This model leverages a pre-trained video-generation VAE to encode 3D p…
-
New GAS framework enhances visual understanding in MLLMs with zero inference overhead
Researchers have developed a new framework called GAS that uses generation as an auxiliary supervision method to enhance visual understanding in multimodal large language models (MLLMs). GAS employs Next Embedding Predi…
-
StatePlay model generates mechanically consistent game worlds by predicting internal states
Researchers have introduced StatePlay, a novel game world model designed to generate visually realistic and mechanically consistent game environments. Unlike previous models that focus solely on pixel-level realism, Sta…
-
NVIDIA releases Cosmos 3 Edge, a 4B parameter on-device AI model for robots
NVIDIA has released Cosmos 3 Edge, a 4-billion-parameter open-world model designed for on-device operation in robotics and vision AI. This model enables robots and AI agents to understand their environment, reason in re…
-
New models RxBrain and GTA-VLA advance embodied AI reasoning
Researchers have introduced RxBrain, a novel foundation model for embodied cognition that integrates language and visual reasoning for planning. Unlike existing models that focus on scene understanding or future state p…
-
GigaWorld-Policy-0.5 enhances robot control with faster inference
Researchers have developed GigaWorld-Policy-0.5, an enhanced World Action Model (WAM) designed for more efficient robot control. This model addresses the computational overhead of traditional WAMs by using future visual…
-
NVIDIA Cosmos 3 World Models: Colab-Friendly Miniature Tutorial Released
This tutorial demonstrates how to implement a miniature version of NVIDIA's Cosmos 3 world models using the cosmos-framework, specifically tailored for Google Colab's hardware limitations. It guides users through checki…
-
New AI models advance tactile understanding and visuo-tactile manipulation
Researchers have developed UniTac, a novel unified multimodal model designed for tactile understanding and generation across different sensors. This model captures the physical interaction between sensors and objects th…
-
Mural integrates frozen LLMs into image generation via Mixture-of-Transformers
Researchers have developed a new method called Mural that integrates frozen Large Language Models (LLMs) with diffusion-based image generators. This approach utilizes a Mixture-of-Transformers (MoT) architecture to tran…
-
Symbiotic-MoE framework enhances multimodal AI by merging generation and understanding
Researchers have developed Symbiotic-MoE, a new pre-training framework designed to improve Large Multimodal Models (LMMs) by enabling them to perform both image generation and understanding tasks without catastrophic fo…
-
Robotics research advances world models for action and scene generation · 7 sources tracked
A new tutorial paper clarifies the scope of "world models" in robotics, categorizing them into observation-space and state-space types and introducing "world action models" that link predictions to robot actions. Concur…
-
New frameworks enhance multimodal AI by preserving knowledge and improving generation
Researchers are developing new frameworks to enhance multimodal AI models. Rosetta introduces a composable pretraining approach that preserves core knowledge while adding new modalities non-destructively, using Momentum…
-
Vera layered diffusion model enhances video editing with content preservation
Researchers have introduced Vera, a novel layered diffusion framework designed for content-preserving video editing. Unlike existing methods that regenerate entire videos, Vera focuses on generating an edit layer and an…
-
New AI frameworks enhance video editing with content preservation and real-time capabilities
Researchers have developed new frameworks for video editing, addressing limitations in current automated systems. VideoAgent offers an all-in-one solution for diverse video comprehension and editing tasks, utilizing a m…
-
New AI models tackle long-horizon planning for autonomous driving
Researchers are developing advanced AI models for autonomous driving, focusing on improving trajectory planning and long-horizon decision-making. Several new frameworks, including ParkingTransformer, TerraTransfer, Alig…
-
MaskWAM model unifies masks for enhanced robotic control
Researchers have developed MaskWAM, a novel object-centric world-action model designed to improve robotic control through video prediction. By integrating masks as both inputs and predictions using a Mixture of Transfor…
-
Intel and NVIDIA advance AI hardware and models
Intel is focusing on agentic AI to drive a CPU renaissance and aims to establish a full-stack AI computing platform. Meanwhile, NVIDIA has launched Cosmos 3, an open physical AI model built on a Mixture-of-Transformers …
-
NVIDIA launches Cosmos 3 omnimodal model and Nemotron 3 LLM
NVIDIA has launched Cosmos 3, an omnimodal world model that unifies language, image, video, audio, and action using a Mixture-of-Transformers architecture. This release includes open weights, code, and datasets, with fi…
-
X Square Robot releases 4B VLA model with open code, real-robot tests
X Square Robot has released Wall-OSS-0.5, a 4 billion parameter vision-language-action (VLA) model. The model is built upon a 3 billion parameter vision-language model backbone and incorporates action experts using a Mi…