PulseAugur
EN
LIVE 01:27:48

Cactus Compute releases Needle 2, a 14MB tool-calling AI model for edge devices

Cactus Compute has launched Needle 2, a compact 45-million-parameter model designed for tool-calling and structured data extraction. This model is notable for its extremely small footprint, shipping as a 14MB binary and requiring only 28MB of RAM to operate, making it suitable for devices without dedicated AI hardware like GPUs or NPUs. Needle 2 achieves impressive speeds, processing up to 500 tokens per second on a Raspberry Pi 5 and offering efficient performance on various mobile and mixed-reality devices. AI

IMPACT Enables sophisticated AI capabilities on low-power, resource-constrained devices, expanding AI applications beyond traditional cloud or high-end hardware.

RANK_REASON Model release from a specialized AI lab focused on efficient inference. [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Cactus Compute releases Needle 2, a 14MB tool-calling AI model for edge devices

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
Model release from a specialized AI lab focused on efficient inference. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
46 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

    <p>Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no N…

  2. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Cactus Compute has released Needle 2, an open 45M-parameter tool-calling model that runs in a 14MB binary using just 28MB RAM. It targets hardware with no GPU a

    Cactus Compute has released Needle 2, an open 45M-parameter tool-calling model that runs in a 14MB binary using just 28MB RAM. It targets hardware with no GPU and achieves 500 tokens/sec on a Raspberry Pi 5. https://www. marktechpost.com/2026/08/13/ca ctus-compute-needle-2-45m-pa…