Researchers have developed a new post-training quantization method called CTOAC for Visual State Space Models (VSSD), which are extensions of Mamba architectures for image processing. The study found that activation quantization is a key bottleneck for low-bit VSSD performance, with significant channel-wise magnitude variations and token-localized extremes in representative inputs. CTOAC addresses this by learning per-input-channel clipping bounds to minimize reconstruction loss on linear layer outputs, while other operations maintain original precision. This method demonstrated robustness across VSSD variants on ImageNet-1K, COCO, and ADE20K datasets, preserving accuracy and performance in object detection and segmentation tasks. Furthermore, an optimized deployment on RTX 4090 achieved up to 1.42x speedup compared to FP32. AI
IMPACT Improves efficiency and performance of visual state space models, potentially enabling wider adoption in resource-constrained environments.
RANK_REASON Academic paper detailing a new method for optimizing visual state space models. [lever_c_demoted from research: ic=1 ai=1.0]
- ADE20K
- Channel-wise Token-balanced Output-Aware Clipping
- COCO
- CTOAC
- Mamba
- RTX 4090
- Vim
- Visual State Space Duality
- VmambaSCI: Dynamic Deep Unfolding Network with Mamba for Compressive Spectral Imaging
- VSSD-Base
- VSSD-Small
- VSSD-Tiny
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →