PulseAugur
实时 01:00:50
English(EN) Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction

Black Forest Labs发布FLUX 3多模态模型,支持图像、视频和音频

Black Forest Labs推出了FLUX 3,这是一款新颖的多模态基础模型,能够同时处理和生成图像、视频和音频模态的内容。该模型基于Self-Flow方法构建,通过结合流匹配和自监督特征重建,在单一架构内实现了多模态生成和理解的对齐。FLUX 3展示了生成长达20秒并带有同步音频的视频片段的能力,其底层架构还支持能够进行实时动作预测的机器人策略。 AI

影响 为多模态AI树立了新先例,可能加速AI应用和机器人技术中不同数据类型的集成。

排序理由 前沿实验室模型发布,附带系统卡[lever_c_demoted from frontier_release: ic=2 ai=1.0]

在 MarkTechPost 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Black Forest Labs发布FLUX 3多模态模型,支持图像、视频和音频

报道来源 [2]

  1. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    Black Forest Labs 发布 FLUX 3:用于图像、视频、音频和机器人动作预测的多模态流模型

    <p>Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model that learns from images, videos and audio inside a single architecture. It is also the first FLUX model to ship video, audio and action prediction from one set of weights. The Black Forest Labs (BFL) re…

  2. r/singularity TIER_2 English(EN) · /u/elemental-mind ·

    Black Forest Lab 的 Flux 3:图像、视频、音频和动作预测的全模态能力

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1v4osms/black_forest_labs_flux_3_omnimodality_for_image/"> <img alt="Black Forest Lab's Flux 3: Omni-modality for image, video, audio &amp; action prediction" src="https://external-preview.redd.it/bDdpajIxOTB…