A user has successfully run the LFM2.5-2.6B language model on a OnePlus 13 smartphone using only the device's CPU. The model, which features 2.69 billion parameters and a 128K context window, achieved a speed of 17 tokens per second. The user developed a custom inference engine, which is only 450kb in size and supports various other model architectures, and is aiming to increase the processing speed to 30 tokens per second. AI
IMPACT Demonstrates the feasibility of running sophisticated LLMs on mobile hardware with custom inference engines.
RANK_REASON User-developed inference engine running a specific LLM on a mobile device.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →