English(EN)UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
新的摊销矩匹配技术增强视觉生成模型
作者PulseAugur 编辑部·[5 个来源]·
研究人员引入了摊销矩匹配(AMM),一种使用神经网络从数据矩学习分布训练信号的新技术。该方法实例化为摊销弗雷歇距离(AMFD)损失,提供了比精确统计匹配更鲁棒的训练动态,并显著提高了在ImageNet和FDr$^6$等基准测试上的性能。AMM在文本到图像生成方面也显示出潜力,增强了指令遵循能力,并在GenEval基准测试上超越了多步教师模型。
AI
arXiv:2607.26860v1 Announce Type: new Abstract: We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals. By casting diffusion denoisers through polynomial projections, we establish a general framework for moment amo…
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through…
Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive …
arXiv:2607.24157v1 Announce Type: new Abstract: Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single mod…
arXiv:2607.21848v1 Announce Type: new Abstract: Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These ap…