Multiverse Computing has released a paper and open-source code for a new LLM compression method that treats the selection of transformer blocks for removal as a constrained binary optimization (CBO) problem. The method maps the task onto an Ising glass—a disordered spin system—to account for the complex interactions between different blocks.
LLM Block Removal via Ising Optimization Outperforms Baseline Compression Methods
Unlike existing heuristic-based methods that score blocks independently, this approach uses the second-order Taylor expansion of the model's loss to construct a Hessian matrix. This allows the system to account for pairwise couplings between blocks, treating the selection process as a combinatorial optimization problem rather than a simple ranking task.
The practical advantage of this method is its computational efficiency. Once the Hessian is computed using a small calibration dataset, evaluating candidate configurations becomes a low-cost energy calculation that does not require running the full model. For complex cases, the problem can be solved using classical or quantum-inspired solvers such as tabu search.
In practical applications, the method showed significant improvements in deep compression scenarios. For Llama-3.3-70B-Instruct, at a 50% compression rate (removing 40 of 80 blocks), the CBO method maintained an MMLU score near 77, while the strongest baseline fell to the mid-50s. The method also proved effective for hybrid architectures, such as NVIDIA-Nemotron-3-Nano-30B-A3B-FP8, by identifying highly disposable layers within Mixture of Experts (MoE) and attention structures.
Sources
- Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem (Hugging Face Blog, 2026-09-21)
- GitHub