Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Researchers reformulate LLM depth pruning as an Ising spin optimization problem to select transformer blocks for removal. The paper reports a 23-point MMLU improvement over competing methods at 50% compression of Llama-3.3-70B-Instruct. No usable model weights or inference code are provided, leaving the method as a theoretical contribution rather than a deployable tool.

Cover image for Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Researchers framed language model depth reduction as a spin system optimization task. According to the paper, this technique achieves a twenty-three point improvement on benchmarks compared to existing compression strategies when cutting half of Llama-3.3-70B-Instruct. The authors mapped block selection into a constrained binary optimization problem resembling an Ising glass. By computing a second-order Taylor expansion of the model loss, they derived a Hessian matrix capturing pairwise interactions between transformer layers. Evaluating candidate configurations relies on calculating the energy of this spin system rather than running full model evaluations. The publication provides no usable model weights or inference code. Consequently, the approach remains a theoretical framework rather than a deployable tool.