PulseAugur
EN
LIVE 01:01:57

Tencent's WeChat AI team scales 617B parameter WeLM model with novel architecture

Tencent's WeChat AI team has developed a new model called WeLM, boasting 617 billion parameters. To overcome GPU shortages, they are employing a "Hidden Decoding" architecture that activates only a fraction of the model's parameters during operation. This approach requires a fourfold increase in training compute but allows for efficient scaling and deployment. AI

IMPACT This novel architecture could offer a path for other AI teams to scale large models efficiently despite hardware constraints.

RANK_REASON The item details a new model architecture and scaling technique for a large language model. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Tencent's WeChat AI team scales 617B parameter WeLM model with novel architecture

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Tencent’s WeChat AI team is fighting the GPU shortage with algorithmic cleverness, scaling its new WeLM model to a massive 617 billion parameters while activati

    Tencent’s WeChat AI team is fighting the GPU shortage with algorithmic cleverness, scaling its new WeLM model to a massive 617 billion parameters while activating just 23 billion of them. Using a novel "Hidden Decoding" architecture, they are trading a 4x increase in training com…