Tencent's WeChat AI team has developed a new model called WeLM, boasting 617 billion parameters. To overcome GPU shortages, they are employing a "Hidden Decoding" architecture that activates only a fraction of the model's parameters during operation. This approach requires a fourfold increase in training compute but allows for efficient scaling and deployment. AI
IMPACT This novel architecture could offer a path for other AI teams to scale large models efficiently despite hardware constraints.
RANK_REASON The item details a new model architecture and scaling technique for a large language model. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →