A user has developed a "triple-loop" model architecture, inspired by the Nanbeige 4.5, and applied it to Qwen3.5-9B. This experimental model, trained using distilled logits from Qwen3.8-27B, shows significant improvements in math, long-context tasks, and instruction-following compared to the base Qwen3.5-9B. However, the training was cut short due to resource limitations, resulting in a plateau in performance and a slight decrease in reasoning and translation capabilities. The user believes further gains are possible with a complete training schedule and proper learning rate decay. AI
IMPACT Demonstrates potential for architectural innovations to improve existing models, though further validation is needed.
RANK_REASON User-developed model architecture applied to an existing base model, with performance evaluations. [lever_c_demoted from research: ic=1 ai=1.0]
- ChatGPT
- Claude
- Gemini
- Lordnyx/qwen3.5-9b-triple-loop-fase1
- Modal
- Nanbeige
- Qwen3-0.6B
- Qwen3.5:9b
- Qwen3.8-27B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →