PulseAugur
实时 12:36:55
English(EN) UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

新的摊销矩匹配技术增强视觉生成模型

研究人员引入了摊销矩匹配(AMM),一种使用神经网络从数据矩学习分布训练信号的新技术。该方法实例化为摊销弗雷歇距离(AMFD)损失,提供了比精确统计匹配更鲁棒的训练动态,并显著提高了在ImageNet和FDr$^6$等基准测试上的性能。AMM在文本到图像生成方面也显示出潜力,增强了指令遵循能力,并在GenEval基准测试上超越了多步教师模型。 AI

影响 引入了提高视觉生成质量和效率的新技术,可能影响文本到图像和视频合成应用。

排序理由 该集群包含多篇详细介绍视觉生成新方法的论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新的摊销矩匹配技术增强视觉生成模型

报道来源 [5]

  1. arXiv cs.LG TIER_1 English(EN) · Wenze Liu, Xintao Wang, Pengfei Wan, Xiangyu Yue ·

    Amortized Moment Matching for Visual Generation

    arXiv:2607.26860v1 Announce Type: new Abstract: We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals. By casting diffusion denoisers through polynomial projections, we establish a general framework for moment amo…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    UniGen-AR:通过自回归建模统一视觉生成

    Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    闭环:无训练的重访一致性用于自回归生成渲染

    Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive …

  4. arXiv cs.CV TIER_1 English(EN) · Zhipeng Bao, Zhen Zhu, Nupur Kumari, Anurag Bagchi, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert ·

    UniGen-AR:通过自回归建模统一视觉生成

    arXiv:2607.24157v1 Announce Type: new Abstract: Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single mod…

  5. arXiv cs.CV TIER_1 English(EN) · Wenchao Ma, Changran Liu, Sharon X. Huang, Haomiao Jiang ·

    闭环:无训练的重访一致性用于自回归生成渲染

    arXiv:2607.21848v1 Announce Type: new Abstract: Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These ap…