Cactus-Compute/needle3

Cactus-Compute releases Needle 3, a sub-30MB foundation model optimized for mobile, wearable, and embedded devices. The architecture uses Laddered Simple Attention with engram memory to handle tool calling, structured extraction, and local text embedding. Developers can install the Python package to run inference on-device, with weights compressed to CQ2-bit for efficiency.

README

Cactus-Compute/needle3 View on Hugging Face

Loading the README from Hugging Face…