PulseAugur
EN
LIVE 05:47:42

LFM2.5-2.6B model runs on OnePlus 13 CPU at 17 tokens/sec

A user has successfully run the LFM2.5-2.6B language model on a OnePlus 13 smartphone using only the device's CPU. The model, which features 2.69 billion parameters and a 128K context window, achieved a speed of 17 tokens per second. The user developed a custom inference engine, which is only 450kb in size and supports various other model architectures, and is aiming to increase the processing speed to 30 tokens per second. AI

IMPACT Demonstrates the feasibility of running sophisticated LLMs on mobile hardware with custom inference engines.

RANK_REASON User-developed inference engine running a specific LLM on a mobile device.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LFM2.5-2.6B model runs on OnePlus 13 CPU at 17 tokens/sec

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/trikboomie ·

    LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vg8qfv/lfm2526b_on_a_oneplus_13_at_17_toks_pure_cpu/"> <img alt="LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU" src="https://preview.redd.it/oeyr9vuhfkhh1.gif?width=640&amp;crop=smart&amp;s=e37a0906f0628…