A new paper presented at ICML 2026, co-authored by researchers from Meta, Google DeepMind, Cornell University, and NVIDIA, quantifies the information storage capacity of large language model parameters. The study found that each parameter in a Transformer model, using bfloat16 format, can store approximately 3.6 bits of information. This research also sheds light on the phenomenon of "double descent" and has implications for data privacy, suggesting that while average training data is unlikely to be memorized, rare or sensitive information poses a significant risk of leakage. AI
IMPACT Quantifies LLM memory capacity, offering insights into scaling laws and potential privacy risks from memorized training data.
RANK_REASON The cluster reports on a scientific paper detailing new findings about LLM parameter capacity and information theory. [lever_c_demoted from research: ic=1 ai=1.0]
- bfloat16
- Cornell University
- Google DeepMind
- International Conference on Machine Learning
- Meta
- NVIDIA
- single-precision floating-point format
- Transformer++
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →