Multi-Token Prediction (MTP)
PulseAugur coverage of Multi-Token Prediction (MTP) — every cluster mentioning Multi-Token Prediction (MTP) across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Tencent releases Hy3, an open 295B MoE model with 256K context
Tencent has released Hy3, an open-source 295 billion parameter Mixture-of-Experts (MoE) model designed for complex reasoning, agentic workflows, and long-context tasks. The model activates only 21 billion parameters per…
-
ai-sage releases GigaChat 3.5 Ultra with 432B parameters
ai-sage has released GigaChat 3.5 Ultra, a 432B parameter Mixture-of-Experts model designed for multilingual tasks, reasoning, and code generation. This new version is approximately 40% more compact than its predecessor…
-
Local LLM Speed Boosted by Gemma 4 MTP and QAT
A recent update to the "Run LLMs Locally" project has introduced Multi-Token-Prediction (MTP) for Gemma models, achieving speed improvements of up to 90% in token generation. This optimization, combined with Quantizatio…