Model Weight Profiles: Where Do the Parameters Go?
AMD explains how to calculate actual model file sizes by analyzing parameter distribution across embeddings, attention, and dense layers. The post details why quantized checkpoints often exceed naive byte-per-parameter estimates, citing a Llama 3.1 8B example where BF16 embeddings increased size to 9.08 GB. This method helps developers predict storage and KV cache requirements for mixed-precision deployments.