Cactus Compute has released Needle 2, a compact, on-device language model specifically designed for tool-calling tasks. This 14 MB model, quantized to CQ2-bit, requires a fixed 28 MB of RAM and boasts a Simple Attention Network architecture. It aims to enable efficient AI deployment on devices like smartphones and wearables, matching the performance of larger models like FunctionGemma 270M on tool-calling benchmarks. The model is already in production use for voice assistants on wearables and for structured data extraction. AI
IMPACT Enables more capable AI applications on resource-constrained devices, reducing reliance on cloud inference for specific tasks.
RANK_REASON This is a release of a specialized AI tool, not a frontier model release from a major lab.
- Cactus Compute
- Cactus-Compute/needle2
- FunctionGemma
- Henry Ndubuaku
- Hugging Face
- Jakub Mróz
- Karen Mosoyan
- Needle 2
- Pebble Index Ring
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →