Ben Houston has developed Three-LLM, a WebGPU-based inference engine that allows Large Language Models (LLMs) to run locally within a web browser. This project leverages Three.js and its WebGPU capabilities to execute LLM inference directly on the user's GPU, supporting various models like GPT-2, SmolLM2, Phi, Qwen, and Llama architectures. Significant performance optimizations have been implemented, including reducing command submissions and reusing prompt prefixes, which have led to substantial speed improvements, such as a 4.7x increase in TinyStories decode performance. This advancement highlights the browser's growing potential as a compute platform for AI inference, with implications for privacy, latency, and offline applications. AI
IMPACT Enables privacy-preserving, low-latency AI experiences by running LLMs directly in the browser.
RANK_REASON Demonstrates a novel application of existing web technologies (WebGPU, Three.js) for local LLM inference, rather than a release from a frontier lab or a significant industry-wide event.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →