Swap the pruner, keep the gains.
| Pruner | Pruning only | SCOPD+ |
|---|---|---|
| VisionZip | 88.74 | 95.25 |
| DivPrune | 83.47 | 91.23 |
| Random | 82.06 | 88.85 |
| FastV | 81.47 | 86.03 |
Avg6 Ā· Table 3
Pruned VLMs often still have the visual evidence they need.
SCOPD teaches them to use it.
fewer visual tokens
of full-context performance
points over pruning alone
Visual-token pruning makes VLMs cheaper, but accuracy drops sharply at aggressive budgets, usually blamed on lost visual information. That is only part of the story: sampling repeatedly from the same pruned tokens often recovers the correct answer. The evidence survives; the model fails to use it reliably. We call this the representationāutilization gap.
SCOPD closes the gap with on-policy self-distillation: the model reasons from pruned tokens while a full-context copy of itself supervises the same trajectory. SCOPD+ focuses this supervision on the tokens that depend most on visual evidence.
Same pruned tokens, 64 samples: success rises from 53.2% to 79.6%.
One model, two views: a sparse-context student and a full-context teacher.
The student generates its own reasoning from the pruned visual tokens.
An EMA teacher sees all visual tokens, scores the same prefixes, and the student matches it token by token.
Add 1% more tokens and see which predictions shift. Distill only the top 10% most sensitive positions.
Each variant distills 10% of response tokens (dense uses all). Avg6 at 10% visual tokens.
Choosing the least sensitive tokens drops to 89.13, confirming the signal. Table 2.
Qwen2.5-VL-7B with VisionZip on 13 image benchmarks.
Each benchmark divided by the unpruned model's score, then averaged. 100 = full-context performance.
| Benchmark | Base model (all tokens) | Base model | SCOPD | SCOPD+ |
|---|
Trained on images with VisionZip; all numbers below at 10% visual tokens.
| Pruner | Pruning only | SCOPD+ |
|---|---|---|
| VisionZip | 88.74 | 95.25 |
| DivPrune | 83.47 | 91.23 |
| Random | 82.06 | 88.85 |
| FastV | 81.47 | 86.03 |
Avg6 Ā· Table 3
Pruning only ā SCOPD+ (SCOPD: 81.60)
Avg6 Ā· Table 4
Pruning only ā SCOPD+ on five video benchmarks
Avg5 Ā· Table 5