PulseAugur
EN
LIVE 03:55:52

New Densing Law for User Representation Learning at Billion-Scale Capacity

Researchers have proposed a "User Behavioral Densing Law" to characterize the relationship between data scale and tokenization capacity in large-scale user representation learning. A pilot study using Alipay data demonstrated that tokenization offers sustained performance gains over raw data scaling. The proposed law, derived from theoretical analysis and experiments, suggests an approximately linear relationship between the logarithms of minimum sufficient tokenization capacity and input data size. This law guided the development of ALGN, an adaptive variable-length tokenization method that improves capacity allocation and outperforms existing baselines across various data sources and tasks. AI

IMPACT Provides practical guidance for optimizing tokenization in large-scale user representation learning, potentially improving efficiency and performance.

RANK_REASON Academic paper on a new theoretical law and method for user representation learning. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Densing Law for User Representation Learning at Billion-Scale Capacity

COVERAGE [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Weiqiang Wang ·

    Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

    User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance ex…