Cactus-Compute/needle3
Cactus-Compute releases Needle 3, a sub-30MB foundation model optimized for mobile, wearable, and embedded devices. The architecture uses Laddered Simple Attention with engram memory to handle tool calling, structured extraction, and local text embedding. Developers can install the Python package to run inference on-device, with weights compressed to CQ2-bit for efficiency.
README
Cactus-Compute/needle3 View on Hugging Face
Loading the README from Hugging Face…