PulseAugur
EN
LIVE 20:20:52

Unsloth Studeo achieves 111.3 tokens/sec with Gemma4:e4B model

Unsloth Studeo, running on a laptop, achieved a message timing of 111.3 tokens per second with the Gemma4:e4B model. This performance metric, measured in tokens per second, was described as "crazy" by the user, who noted that the web application's lack of automatic response speaking was a drawback but expressed optimism about potentially coding a solution. AI

IMPACT Demonstrates specific performance benchmarks for local AI model execution.

RANK_REASON User reports performance metrics for a specific model and software combination.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Unsloth Studeo achieves 111.3 tokens/sec with Gemma4:e4B model

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User reports performance metrics for a specific model and software combination.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
95 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    So, Unsloth Studeo, same laptop as Ollama, Gemma4:e4B: (technically gemma-4-E4B-it-qat-GGUF · UD-Q4_K_XL) Message timing 111.3 tok/s. Tokens per second. This is

    So, Unsloth Studeo, same laptop as Ollama, Gemma4:e4B: (technically gemma-4-E4B-it-qat-GGUF · UD-Q4_K_XL) Message timing 111.3 tok/s. Tokens per second. This is just plain crazy. I mean, the web app doesn't automatically speak responses, which sucks, but like I may be able to vib…