PulseAugur
EN
LIVE 22:09:53

Ornith 1.5 35B A3B model shipped with untrained MTP head, causing slow performance

The Ornith-1.5-35B-A3B large language model is experiencing performance issues because it is being distributed with an untrained Multi Token Prediction (MTP) head. This head, which is crucial for the model's output, was randomly initialized and has not undergone any training. This oversight is the direct cause of the model's unexpectedly slow performance, as reported by users and investigated on HuggingFace. AI

IMPACT This issue with Ornith 1.5 35B A3B highlights the importance of thorough testing and training for all model components, especially heads like MTP, to ensure expected performance.

RANK_REASON The cluster describes a technical flaw in a specific model release that impacts its performance.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Ornith 1.5 35B A3B model shipped with untrained MTP head, causing slow performance

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Max-_-Power ·

    If you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vtu555/if_you_are_wondering_why_ornith_15_35b_a3b_with/"> <img alt="If you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why" src="https://external-preview.redd.it/XlulZLJ08y_-MAoKniu9TOJU…