POCKET is a new on-device LLM designed to run a 35B-parameter model on standard CPUs without requiring a GPU, utilizing the llama.cpp inference stack. This approach aims to overcome the barriers of GPU scarcity and cost for local and edge deployments. While it achieves approximately 3.4x the throughput of the on-device baseline Bonsai, its performance and output quality, particularly for Korean language generation, are heavily dependent on quantization levels, with Q4 or higher recommended for usable results. AI
IMPACT Enables running larger LLMs on consumer-grade hardware, potentially increasing privacy and accessibility for on-device AI applications.
RANK_REASON New product release for running LLMs on consumer hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →