GPT-4 utilizes a Mixture-of-Experts (MoE) architecture, which means its reported 1.8 trillion parameters are not all active for every computation. Instead, when processing a token, only about 36 billion parameters, or approximately 2%, are engaged. This design choice, where a routing function selects a subset of 'expert' networks for each token, is a deliberate feature with mathematical underpinnings for efficient scaling. AI
IMPACT This architectural detail suggests a path toward more efficient scaling of large language models by selectively activating parameters.
RANK_REASON The item details the architecture of an existing model, GPT-4, explaining its parameter usage and design choices. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →