Cactus Compute has launched Needle 2, a compact 45-million-parameter model designed for tool-calling and structured data extraction. This model is notable for its extremely small footprint, shipping as a 14MB binary and requiring only 28MB of RAM to operate, making it suitable for devices without dedicated AI hardware like GPUs or NPUs. Needle 2 achieves impressive speeds, processing up to 500 tokens per second on a Raspberry Pi 5 and offering efficient performance on various mobile and mixed-reality devices. AI
IMPACT Enables sophisticated AI capabilities on low-power, resource-constrained devices, expanding AI applications beyond traditional cloud or high-end hardware.
RANK_REASON Model release from a specialized AI lab focused on efficient inference. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Read on Mastodon — mastodon.social →
- A Controlled Study of Attention-Only Transformers
- Apple FM
- Apple Vision Pro
- arXiv
- Cactus Compute
- Cactus Quants
- FunctionGemma
- Meta Quest 3S
- Needle 2
- Raspberry Pi 5
- LFM2.5-230M
- Simple Attention Network
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →