A new conversion of the Inkling-Small 276B-A12B model, which has approximately 12 billion active parameters, has been optimized to run on consumer-grade hardware with less than 10GB of memory. Benchmarks show the model achieving speeds of around 2.9 tokens per second for generation, though long prompt prefill times are noted as an issue. This development is part of an ongoing effort to make large Mixture-of-Experts (MoE) models accessible on standard hardware. AI
IMPACT Enables running large MoE models on consumer-grade hardware, potentially broadening access and use cases.
RANK_REASON The item describes the optimization and release of a specific model conversion for consumer hardware, which falls under tooling.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →