Tencent has released Hy3, an Apache 2.0 licensed Mixture-of-Experts (MoE) model. This model boasts 295 billion total parameters with 21 billion active parameters and utilizes a 3.8 billion parameter MTP layer. Hy3 supports a 256K context window and can be served using vLLM or SGLang with an OpenAI-compatible API. An FP8 quantized variant, Hy3-FP8, is also available, which further reduces VRAM consumption. AI
IMPACT This release offers a large context window and efficient serving options for a powerful MoE model, potentially impacting research and application development.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →