PulseAugur
EN
LIVE 04:01:35

Netflix uses multimodal embeddings to personalize content discovery

Netflix has developed and implemented multimodal embeddings to enhance its content personalization systems. By leveraging models like CLIP for image embeddings and a tri-modal foundation model called MediaFM (which fuses visual, audio, and text signals), Netflix has improved its ability to personalize artwork and video previews. This approach has led to better cold-start performance, query-aware search, and superior video preview recommendations, outperforming previous single-modality models in both offline and online tests. The company also established an offline proxy task to accelerate experimentation and productization of these embedding models. AI

IMPACT Enhances content discovery and personalization in streaming services, potentially setting new industry standards for media asset optimization.

RANK_REASON Research paper detailing the application of multimodal embeddings in a production recommender system. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Netflix uses multimodal embeddings to personalize content discovery

COVERAGE [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Ashish Rastogi ·

    Multimedia Asset Personalization via Multimodal Embeddings at Netflix

    Personalized promotional assets, namely artwork images and video preview clips, are critical to content discovery on Netflix. Traditional models for asset selection rely on ID-based interaction history, leaving them blind to asset content and unable to serve newly launched titles…