Edge0/Edge0-35B-A3B-preview

Edge0-35B-A3B-preview is a 35B sparse MoE model designed for phone-class memory, running in under 3 GiB active RAM. It uses SSD expert offloading and a prerouter to achieve 15 tok/s decode speed. The release includes int4 quantization and LoRA adapters, claiming quality within 3.9 points of the fp16 base. This is an early preview of the edge0 pipeline.

README

Edge0/Edge0-35B-A3B-preview View on Hugging Face

Loading the README from Hugging Face…