PulseAugur
EN
LIVE 11:37:37

Hetzner experiments with OpenAI-compatible LLM inference API

Hetzner is experimenting with offering LLM inference services through an OpenAI-compatible API. This early-stage offering, currently without billing or SLAs, uses the Qwen/Qwen3.6-35B-A3B-FP8 model and is designed to gauge user interest and system scalability. Initial tests indicate fast performance, with a median time to first token of 153 ms and an output speed of 224 tokens per second. AI

IMPACT This experiment could signal a new trend in cloud providers offering direct LLM inference, potentially increasing competition and accessibility.

RANK_REASON Hetzner is experimenting with an LLM inference service, which is a product offering but not a frontier release or significant industry move.

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Hetzner experiments with OpenAI-compatible LLM inference API

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 Deutsch(DE) · Jonas Scholz ·

    Hetzner Inference: First Look

    <p>Hetzner is experimenting with LLM inference.</p> <p>That is not a sentence I expected to write, but I think it is pretty interesting :)</p> <p>Before anyone moves their production AI workloads to Hetzner: this is very much an <strong>experiment</strong>. There is no billing, n…