A user is providing updates on their post-training efforts for the Yandex/AliceAI-80B-A3B model. The initial training run encountered a NaN issue, leading to zeroed gradients and a failed run, necessitating a restart. The user has since restarted the training process, monitoring a healthier loss curve and planning to release GGUF versions and a llama.cpp patch upon completion. Previous updates detailed the instruct fine-tuning process using synthetic data and the strategy of training the attention and shared expert layers while keeping base experts frozen. AI
IMPACT This work contributes to the open-source community's efforts in fine-tuning and adapting large language models for specific tasks.
RANK_REASON User-led research and development on an existing model, not a frontier release.
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →