PulseAugur
EN
LIVE 07:31:08

Needle 2: Tiny 45M-parameter tool-calling model runs on smartphones

Cactus Compute has launched Needle 2, a compact 45-million-parameter model designed for tool-calling and structured data extraction. This model is notable for its extremely small footprint, shipping as a 14MB binary and requiring only 28MB of RAM to run a full session, making it suitable for resource-constrained devices like smartphones and wearables. Needle 2 achieves impressive speeds, processing up to 500 tokens per second on a Raspberry Pi 5, and is deployable across a wide range of platforms including iOS, Android, and WebAssembly. AI

IMPACT Enables advanced AI capabilities on extremely low-power devices, expanding the reach of LLMs beyond traditional computing platforms.

RANK_REASON Model release from a new lab with novel architecture and deployment characteristics. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on MarkTechPost →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Needle 2: Tiny 45M-parameter tool-calling model runs on smartphones

COVERAGE [1]

  1. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

    <p>Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no N…