A user has successfully adapted the TurboFieldfare engine, originally designed for Gemma models, to support Qwen 3.6 35B. This porting effort resulted in the Qwen model requiring less RAM, approximately 1.4 GB compared to Gemma's 2.1 GB, due to its smaller experts and use of linear attention. While the Qwen model is slower in terms of tokens per second on the user's M5 machine, this is attributed to its larger expert files necessitating more frequent SSD reads. AI
IMPACT Enables lower-resource deployment of Qwen models via existing efficient engines.
RANK_REASON User porting an existing engine to a new model.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →