SCOPDSparse-Context On-Policy Self-Distillation
for Efficient Vision-Language Models

      *Equal contribution

      Pruned VLMs often still have the visual evidence they need.
      SCOPD teaches them to use it.

      Visual tokens
      Full 100%Pruned 10%
      Score at 10% tokens (full context = 100)
      Pruning only86.37
      Pruning + SCOPD+92.43

      Normalized mean over 13 image benchmarks Ā· Qwen2.5-VL-7B Ā· VisionZip

      90%

      fewer visual tokens

      92.43%

      of full-context performance

      +6.06

      points over pruning alone

      Overview

      Lost accuracy isn’t always lost evidence.

      Visual-token pruning makes VLMs cheaper, but accuracy drops sharply at aggressive budgets, usually blamed on lost visual information. That is only part of the story: sampling repeatedly from the same pruned tokens often recovers the correct answer. The evidence survives; the model fails to use it reliably. We call this the representation–utilization gap.

      SCOPD closes the gap with on-policy self-distillation: the model reasons from pruned tokens while a full-context copy of itself supervises the same trajectory. SCOPD+ focuses this supervision on the tokens that depend most on visual evidence.

      Key observation

      The evidence is still there.

      Same pruned tokens, 64 samples: success rises from 53.2% to 79.6%.

      Left: a pruned image still yields the correct answer among repeated samples. Middle: with VisionZip at 10% tokens, success rises from 53.2% (greedy) to 79.6% (Pass@64); a language-only control stays between 2.8% and 4.2%. Right: radar plot where SCOPD+ at 10% tokens sits close to the unpruned model on 13 benchmarks.
      500 numerical questions the unpruned model solves greedily; VisionZip keeps 10% of tokens, fixed across all samples. With no image at all, success stays near zero (2.8% → 4.2%), so the recovered answers come from the retained visual evidence.
      Method

      Learn from your own full view.

      One model, two views: a sparse-context student and a full-context teacher.

      1. Reason with less

        The student generates its own reasoning from the pruned visual tokens.

        SCOPD
      2. Supervise with more

        An EMA teacher sees all visual tokens, scores the same prefixes, and the student matches it token by token.

        SCOPD
      3. Focus where vision matters

        Add 1% more tokens and see which predictions shift. Distill only the top 10% most sensitive positions.

        SCOPD+
      (a) High teacher–student KL can reflect language-level disagreement rather than visual dependence. (b) SCOPD+ pipeline: the student at budget b and an intervened student at budget b+Ī“ score the same prefixes; their Jensen–Shannon divergence selects the top ρ% positions, where the full-context teacher's KL loss is applied.
      (a) Large teacher–student disagreement can come from language alone. (b) SCOPD+ compares the student at budgets b and b+Ī“; positions with the largest Jensen–Shannon divergence get the teacher's KL loss (b = 10%, Ī“ = 1%, ρ = 10%).
      • No ground-truth labels
      • No architecture changes
      • No inference overhead
      • +1.9% training time over SCOPD

      Which tokens should be distilled?

      Each variant distills 10% of response tokens (dense uses all). Avg6 at 10% visual tokens.

      Choosing the least sensitive tokens drops to 89.13, confirming the signal. Table 2.

      Results

      The fewer the tokens, the bigger the gain.

      Qwen2.5-VL-7B with VisionZip on 13 image benchmarks.

      Normalized score (Avg13)

      Each benchmark divided by the unpruned model's score, then averaged. 100 = full-context performance.

      Per-benchmark scores

      Scores on 13 image benchmarks
      BenchmarkBase model
      (all tokens)
      Base modelSCOPDSCOPD+
      Generalization

      Train once, transfer broadly.

      Trained on images with VisionZip; all numbers below at 10% visual tokens.

      Unseen pruners

      Swap the pruner, keep the gains.

      Avg6, pruning only versus SCOPD+
      PrunerPruning onlySCOPD+
      VisionZip88.7495.25
      DivPrune83.4791.23
      Random82.0688.85
      FastV81.4786.03

      Avg6 Ā· Table 3

      Another model

      Works on Qwen3-VL-4B.

      75.29 82.92

      Pruning only → SCOPD+ (SCOPD: 81.60)

      Avg6 Ā· Table 4

      Images → video

      No video training needed.

      90.88 96.60

      Pruning only → SCOPD+ on five video benchmarks

      Avg5 Ā· Table 5

      Citation

      BibTeX