well9472/Nanosaur2-670M
Nanosaur2 is a 670M parameter text-to-image DiT trained for approximately $600. It integrates with ComfyUI using a custom node and specific model files. The architecture combines a frozen Gemma text encoder with a DINOv2-based VAE. The author explicitly states the model is for research purposes to test minimal compute limits. It is not intended to compete with fully trained commercial
This artifact is a 670M parameter text-to-image diffusion transformer designed for research. The author explicitly states the goal is to test minimal compute limits rather than compete with fully trained commercial systems. It generates illustrations using a frozen Gemma text encoder and a custom DINOv2-based VAE. The model was trained for a total cost of approximately $600. Integration requires installing a custom node into the ComfyUI environment and placing three specific model files in designated folders. A provided workflow file allows users to load the configuration directly into the interface. Inference relies on Euler simple sampling with 50 steps and a CLIP guidance scale of 4. Users can employ tag-based prompts or natural language, often using quality tags to adjust output preferences. Potential users should verify that their hardware matches the specific inference requirements documented by the publisher. The author notes that the current version does not match the character knowledge of larger commercial models. There is uncertainty regarding long-term stability since the author states this model was not intended for production use. Further development, including a larger 2B version, depends on securing additional compute funding.
README
well9472/Nanosaur2-670M View on Hugging Face
Loading the README from Hugging Face…