PulseAugur
EN
LIVE 19:45:41

Thinking Machines releases Inkling-Small, a 276B MoE model runnable on one GPU

Thinking Machines has released Inkling-Small, a 276 billion parameter Mixture-of-Experts model designed to run on a single GPU. This development allows for more accessible deployment of large-scale AI models. The announcement was part of a broader AI briefing that also touched on NVIDIA's work in agentic reinforcement learning and other industry developments. AI

IMPACT Enables wider deployment of large AI models by reducing hardware requirements.

RANK_REASON This is a model release from a company that is not a Tier-1 frontier lab, and the focus is on its accessibility (running on one GPU) rather than a breakthrough performance claim.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Thinking Machines releases Inkling-Small, a 276B MoE model runnable on one GPU

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
This is a model release from a company that is not a Tier-1 frontier lab, and the focus is on its accessibility (running on one GPU) rather than a breakthrough performance claim.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
54 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📰 Latent — Thinking Machines shrinks its flagship model to one GPU • Thinking Machines ships Inkling-Small, a 276B MoE that runs on one GPU • NVIDIA's Molt stri

    📰 Latent — Thinking Machines shrinks its flagship model to one GPU • Thinking Machines ships Inkling-Small, a 276B MoE that runs on one GPU • NVIDIA's Molt strips agentic RL down to ~8.6K lines • Onton claims its search model beats Google Shopping and Amazon by 2.7x • Big Tech si…