PulseAugur
实时 12:36:25
English(EN) Douyin Multimodal Embedding Model Technical Report

抖音发布先进多模态嵌入模型,用于搜索和推荐 · 跟踪2个来源

研究人员开发了抖音多模态嵌入(DME)模型,这是一个为高效、细粒度多模态搜索和推荐设计的两阶段系统。该模型首先进行大规模对比预训练,以建立统一的嵌入空间,然后进行第二阶段,通过基于证据的潜在推理和跨条件重构来增强语义充分性。这种方法使DME在MMEB-v2基准测试中取得了最先进的成果,其9B变体得分为78.4。在抖音的生产环境中,DME在离线评估中显示出2.92%的相对提升,在搜索场景的在线A/B测试中显示出0.1%的生命周期增益。 AI

影响 该模型的双阶段训练方法可能会影响未来的多模态嵌入架构,在效率和大规模应用的细粒度区分之间取得平衡。

排序理由 技术报告,详细介绍了一个新的多模态嵌入模型及其基准测试结果。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

抖音发布先进多模态嵌入模型,用于搜索和推荐 · 跟踪2个来源

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou ·

    抖音多模态嵌入模型技术报告

    arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with compl…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Zhicheng Dou ·

    抖音多模态嵌入模型技术报告

    Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as D…