Thinking Machines has released the full weights for its Inkling model, a Mixture of Experts (MoE) architecture with 975 billion total parameters and 41 billion active parameters. While the weights are freely available, running the model requires significant computational resources, making it impractical for typical home PCs. The model supports text, image, and audio inputs with a context window of up to 1 million tokens, and is positioned as a customizable base for further development rather than a universally superior model. AI
IMPACT The release of Inkling's weights offers researchers and developers a powerful, customizable base model, but highlights the significant infrastructure costs associated with large MoE architectures.
RANK_REASON Release of model weights and technical details by a developer. [lever_c_demoted from research: ic=1 ai=1.0]
- GGUF
- Hacker News
- Inkling
- llama.cpp
- mixture of experts
- Mlx
- r/LocalLLaMA
- SGLang
- Thinking Machines
- transformers
- Unsloth
- vLLM
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →