PulseAugur
中
实时 21:21:42
English(EN) Douyin Multimodal Embedding Model Technical Report

抖音发布先进多模态嵌入模型,用于搜索和推荐 · 跟踪2个来源

研究人员开发了抖音多模态嵌入(DME)模型,这是一个为高效、细粒度多模态搜索和推荐设计的两阶段系统。该模型首先进行大规模对比预训练,以建立统一的嵌入空间,然后进行第二阶段,通过基于证据的潜在推理和跨条件重构来增强语义充分性。这种方法使DME在MMEB-v2基准测试中取得了最先进的成果,其9B变体得分为78.4。在抖音的生产环境中,DME在离线评估中显示出2.92%的相对提升,在搜索场景的在线A/B测试中显示出0.1%的生命周期增益。 AI

影响 该模型的双阶段训练方法可能会影响未来的多模态嵌入架构,在效率和大规模应用的细粒度区分之间取得平衡。

排序理由 技术报告,详细介绍了一个新的多模态嵌入模型及其基准测试结果。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

抖音发布先进多模态嵌入模型,用于搜索和推荐 · 跟踪2个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
技术报告,详细介绍了一个新的多模态嵌入模型及其基准测试结果。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
59 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou ·

    抖音多模态嵌入模型技术报告

    arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with compl…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Zhicheng Dou ·

    抖音多模态嵌入模型技术报告

    Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as D…