architectds/collabosm

A Python notebook enables running a 125B parameter MoE model on a single A100 GPU in Colab. It uses a pinned runtime, Hugging Face weights, and an OpenAI-compatible endpoint. Reported performance includes 2.8k to 3.9k tokens per second prefill and 90 to 97 tokens per second decode.

README

architectds/collabosm View on GitHub

Loading the README from GitHub…