Hetzner is experimenting with offering LLM inference services through an OpenAI-compatible API. This early-stage offering, currently without billing or SLAs, uses the Qwen/Qwen3.6-35B-A3B-FP8 model and is designed to gauge user interest and system scalability. Initial tests indicate fast performance, with a median time to first token of 153 ms and an output speed of 224 tokens per second. AI
IMPACT This experiment could signal a new trend in cloud providers offering direct LLM inference, potentially increasing competition and accessibility.
RANK_REASON Hetzner is experimenting with an LLM inference service, which is a product offering but not a frontier release or significant industry move.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →