English(EN)Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
新方法增强用于图像生成的视觉自回归模型
作者PulseAugur 编辑部·[4 个来源]·
研究人员开发了新方法来提高视觉自回归模型的效率和性能。一种方法,Shift-and-Sum Quantization(移位求和量化),解决了图像生成任务中注意力值乘积的重建误差和校准数据差异。另一个框架UniAR,使用单一视觉分词器统一多模态理解和生成,通过多级特征融合和比特量化在图像生成和编辑方面取得了最先进的成果。
AI
arXiv:2606.16131v1 Announce Type: cross Abstract: Post-training quantization (PTQ) enables efficient deployment of deep networks using a small set of data. Its application to visual autoregressive models (VAR), however, remains relatively unexplored. We identify two key challenge…
UniAR presents a unified autoregressive framework that uses a single discrete visual tokenizer to bridge visual understanding and generation, achieving state-of-the-art results in image generation and editing through multi-level feature fusion, bitwise quantization, and parallel …
arXiv:2606.18249v1 Announce Type: new Abstract: Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hind…
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a …