Mamba2
PulseAugur coverage of Mamba2 — every cluster mentioning Mamba2 across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
RAM-Net introduces sparse state access to improve sequence modeling
Researchers have introduced RAM-Net, a novel sequence modeling approach designed to mitigate inter-token interference in recurrent states. Unlike traditional methods that use a shared, dense state, RAM-Net employs a spa…
-
New analysis explores federated learning for Mamba2 state space models
Researchers have developed new convergence analysis for federated learning algorithms applied to selective state space models (SSMs), such as Mamba2. Existing federated learning methods are largely architecture-agnostic…
-
Univariate Deep Learning Models Show Diminishing Returns for Wave Height Forecasting
A new arXiv paper explores the limitations of univariate deep learning models for forecasting significant wave height (Hs). The study found that while models like DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2 can outp…
-
New AI model AMOR selectively uses attention based on predictive uncertainty
Researchers have introduced AMOR (Adaptive Metacognitive Output Router), a novel hybrid AI architecture that selectively employs attention mechanisms based on predictive uncertainty. This approach augments a recurrent b…
-
New audit method detects causality leaks missed by attention masks
Researchers have developed a new audit method to detect causality violations in sequence models, which goes beyond traditional attention-mask inspections. This lightweight audit, involving two forward passes without tra…
-
NVIDIA releases Nemotron 3.5 Lightning draft models for specialized decoding · 3 sources tracked
NVIDIA has released new draft models under the Nemotron 3.5 Lightning 30B-A3B series, designed for specialized decoding tasks. Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash, with 833 million parameters, accelerates a 30B …
-
New RAD framework enhances long video generation with global memory and local attention
Researchers have introduced a new framework called Recurrent Autoregressive Diffusion (RAD) designed to improve long video generation. RAD integrates temporal recurrent neural network (RNN) layers, specifically LSTM, in…
-
NVIDIA releases Nemotron-Labs-3-Puzzle-75B for Blackwell hardware
NVIDIA has released its Nemotron-Labs-3-Puzzle-75B model, optimized for serving on Blackwell hardware. The model incorporates LatentMoE with Mamba-Interleaving and Multi-Token Prediction (MTP) for enhanced throughput. I…
-
Zyphra releases Zamba2-VL hybrid vision-language models
Zyphra has launched Zamba2-VL, a new family of open-source vision-language models. These models utilize a hybrid architecture combining Mamba2 state-space models with Transformers, offering significantly faster processi…
-
Zamba2-VL models offer faster vision-language processing
Researchers have introduced Zamba2-VL, a new family of vision-language models that leverage a hybrid architecture combining Mamba2 state-space layers with transformer blocks. These models demonstrate strong performance …
-
MambaBEV model uses Mamba2 for improved 3D object detection
Researchers have introduced MambaBEV, a new 3D object detection model for autonomous driving that utilizes the Mamba2 state-space model. This approach enhances global context modeling within the Bird's Eye View (BEV) sp…
-
SurgicalMamba model advances surgical phase recognition accuracy
Researchers have developed SurgicalMamba, a novel model designed for online surgical phase recognition. This model utilizes a dual-path state-space duality (SSD) architecture, inspired by Mamba2, to efficiently process …
-
SurgicalMamba model advances online surgical phase recognition
Researchers have developed SurgicalMamba, a novel model designed for online surgical phase recognition. This model addresses challenges in surgical video analysis, such as long procedure durations and non-uniform time f…
-
NVIDIA launches Nemotron 3 Nano Omni multimodal AI model for agents
NVIDIA has released Nemotron 3 Nano Omni, a multimodal large language model capable of processing vision, audio, video, and text simultaneously. This open model, built on a Mamba2 Transformer Hybrid Mixture of Experts a…