LemonSeed Studio, an iPad-based editor and IDE, has demonstrated on-device inference capabilities using an AMD GPU. The system achieved 159 tokens/sec with the Qwen3.8-27B model on an R9700 GPU, and 64 tokens/sec on a Strix Halo GPU. This engine, LSE, is designed to optimize model performance by fusing operations and generating custom kernels for various GPUs, supporting features like speculative decoding for faster inference. AI
IMPACT Demonstrates improved on-device inference capabilities for LLMs on consumer hardware, potentially enabling more powerful mobile AI applications.
RANK_REASON This is a demonstration of a specific software tool's capability with a particular model and hardware, rather than a new model release or significant industry event.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →