empero-ai/Qwen3.8-35B-A3B-Distill

Empero releases Qwen3.8-35B-A3B-Distill, a sparse MoE model distilled from frontier Qwen3.8 teachers. The 35B parameter model activates only 3B per token, enabling single-GPU deployment with 262k context. It targets improved reasoning in math and code via curated teacher traces. Standard runtime compatibility is documented, though independent verification of performance gains is not provided.

README

empero-ai/Qwen3.8-35B-A3B-Distill View on Hugging Face

Loading the README from Hugging Face…