PulseAugur
EN
LIVE 22:37:49

LFM 2.5 230M model runs in-browser at 1400 tok/s via custom backend

A developer has created a custom backend that allows the LFM 2.5 230M model to run in-browser at speeds of 1400-1500 tokens per second on an RTX 3090 using WebGPU. The system is designed for portability and supports both Nvidia and Apple Silicon hardware, with optimized kernels for each. This technology is being integrated into the Sipp library. AI

IMPACT Enables more accessible and portable local LLM deployments.

RANK_REASON A user-developed tool demonstrating in-browser LLM execution.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LFM 2.5 230M model runs in-browser at 1400 tok/s via custom backend

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/lordhiggsboson ·

    LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v6e0uq/lfm_25_230m_running_at_1440_toks_inbrowser/"> <img alt="LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend" src="https://external-preview.redd.it/Z2E3enFmNHlyZWZoMWNFUNc05jOqDOO_3Mv…