Yandex has released the AliceAI-T5-35B-A0.6B model, a language model featuring an encoder-decoder architecture with sparse Mixture-of-Experts (MoE) layers. This model boasts 34.35 billion unique parameters and utilizes 512 experts per MoE layer, selecting 8 for each token. The release includes instructions and code examples for using the model with the Hugging Face Transformers library, including options for loading the full model or just the encoder, and guidance on fine-tuning with PEFT LoRA. AI
IMPACT Provides a new large language model with a sparse MoE architecture for researchers and developers.
RANK_REASON Release of a new language model with technical details and usage examples. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
- AliceAI-T5
- AliceAI-T5-35B-A0.6B-Base
- AliceAIT5MoEEncoderModel
- Google Colab
- Habr
- Hugging Face
- Kaggle
- transformers
- yandex/AliceAI-T5-35B-A0.6B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →