Skip to content

Research

ViT Compression Trade-offs

Four Vision Transformer compression paradigms (token merging, threshold masking, post-training quantisation and entropy-based head pruning) mapped across a 24-combination sweep of the Pareto frontier.

Built
Mar 2026 to Apr 2026
Role
Project co-lead
Sweep combinations
24Sweep combinations
Peak VRAM, best stack
401 to 312 MBPeak VRAM, best stack
Top-1 accuracy preserved
>70%Top-1 accuracy preserved
Backbones benchmarked
3Backbones benchmarked

Overview

Self-attention costs O(N²d) in sequence length, which is why a Vision Transformer that is comfortable on a datacentre GPU is not comfortable on a laptop. There are four well-known ways to make it cheaper, and the literature reports each of them in isolation. The question this study asks is what happens when you combine them.

The answer is a 24-combination interaction sweep across token merging, threshold masking, quantisation and head pruning, benchmarked on ViT-Small, EfficientFormer and EfficientViT against a 50k ImageNet-1K validation subset, and plotted as a Pareto frontier of accuracy against latency.

The headline result is a stack rather than a single technique. INT4 quantisation combined with token merging at r=8 cut peak VRAM from roughly 401 MB to 312 MB while holding Top-1 accuracy above 70 percent, zero-shot, with no fine-tuning to recover the loss.

The most instructive finding is a negative one. Token merging assumes tokens are an unordered set, but hybrid convolutional-transformer models reconstruct a 2D grid between stages, and physically removing tokens destroys it. That constraint is what produced the custom masking method below.

Screens

Master interaction sweep across 24 compression configurations
The master sweep. All 24 combinations, accuracy against latency.

What it does

  • Token Merging (ToMe) via bipartite matching on key-cosine similarity, with proportional attention rescaling so merged tokens are not diluted.
  • A custom L2 threshold-masking method that preserves tensor shape, written specifically for hybrid conv-transformer backbones.
  • Post-training quantisation to INT8 by AbsMax and INT4 by NF4, through bitsandbytes.
  • Zero-shot structural head pruning ranked by the Shannon entropy of each head's attention distribution.
  • A 24-combination master interaction sweep mapping the accuracy against latency Pareto frontier.
  • An identified optimal zero-shot stack, INT4 with ToMe at r=8, holding above 70 percent Top-1 on ImageNet.
  • Result plots for every axis, checked into the repository as evidence rather than described in prose.

How it works

  • Token similarity S_ij = (k_i·k_j)/(‖k_i‖‖k_j‖) drives bipartite matching; attention is rescaled as softmax(QKᵀ/√d_k + log s) so a merged token carries the weight of the tokens it absorbed.
  • Custom masking proxies importance by I_i = ‖x_i‖₂ and zeroes sub-threshold embeddings in place, preserving the (B, N, C) shape that downstream convolutions require.
  • INT8 quantisation by AbsMax with scale S = 127/max|W|; INT4 by NF4, whose 16 levels are the quantiles of a standard normal. That is the right prior, because trained weights are approximately Gaussian.
  • Head pruning ranks heads by H = −Σ A log(A+ε): a high-entropy head attends to everything and therefore selects nothing, so its W_Q, W_K, W_V and projection slices are physically removed.
  • Peak VRAM measured at roughly 401 MB uncompressed against 312 MB for the INT4 plus ToMe r=8 stack, with Top-1 held above 70 percent and no fine-tuning applied.
  • Benchmarked end to end on an RTX 5070 Ti Mobile, so the latency numbers describe consumer-edge hardware rather than a datacentre.

Problems and answers

  • Token merging breaks hybrid models that rebuild a 2D spatial grid between stages.

    Mask rather than delete. Zero the low-L2 embeddings and keep the tensor shape intact.

  • Uniform INT4 quantisation wastes levels on a Gaussian weight distribution.

    Use NF4, whose levels sit at the quantiles of N(0,1), so resolution goes where the mass is.

  • Which attention heads are safe to remove is not obvious from the weights alone.

    Score heads by attention entropy and prune the flattest. An unfocused head is one that is not making a decision.

  • Reporting each technique alone hides how they interact when stacked.

    Run the full 24-combination sweep and publish the frontier instead of four separate best cases.

NewerImage-to-LaTeX Handwriting OCROlderLaTeX Converter