MIDL 2026

Adaptive Inference for Medical Vision Transformers:
Token Reduction or Early Exit?

Ji Young Byun*1  HyunSeo Lee*1  Jordan Shuff1,3,4,5  Rengaraj Venkatesh6  Nakul S. Shekhawat†4  Kunal S. Parikh†1,3,4,5  Rama Chellappa†1,2
1 Dept. of Biomedical Engineering, Johns Hopkins University  ·  2 Dept. of Electrical & Computer Engineering, JHU
3,4,5 Wilmer Eye Institute, JHU School of Medicine  ·  6 Aravind Eye Hospital, Pondicherry, India

* Equal contribution  ·  † Corresponding authors

Paper Code Cite

TL;DR

Vision Transformers are powerful but computationally expensive for clinical deployment. We propose a unified adaptive inference framework that combines Token Reduction (TR) and Early Exiting (EE) through dataset-specific profiling, using Jensen-Shannon Divergence to quantify spatial redundancy and a lightweight CNN predictor to select inference strategy per sample. Across five diverse medical imaging datasets, our method achieves 71.4% average FLOPs reduction with only 0.1pp accuracy loss, substantially outperforming individual strategies (EE-only: 55.9%, TR-only: 57.7%) and yielding a 2.07× speedup with 54.3% energy reduction.


Motivation

Why Adaptive Inference for Medical Vision Transformers?

Vision Transformers (ViTs) have achieved state-of-the-art results across medical imaging tasks such as dermatological lesion classification, chest X-ray diagnosis, histopathological analysis, and ophthalmic imaging. Yet their high computational cost creates critical barriers for clinical deployment:

🏥

High-Volume Screening

Diabetic retinopathy and cataract screening programs process thousands of images daily. Even modest per-image latency becomes a serious cumulative burden.

📱

Point-of-Care Devices

Mobile imaging for underserved communities (e.g., Aravind Eye Hospital) operates under severe hardware constraints, making standard ViTs impractical.

🎯

Sample Heterogeneity

Medical images vary widely in complexity. A straightforward case shouldn't use the same compute as an ambiguous one — but uniform strategies don't adapt.

Two prominent efficiency strategies exist: Token Reduction (TR) prunes spatially redundant patches, and Early Exiting (EE) terminates inference when prediction confidence is sufficient. Prior work treats these in isolation. We ask: which strategy is right for which dataset and for which sample?

Through comprehensive dataset profiling, we show that optimal strategies vary dramatically across modalities, and that adaptively combining TR and EE yields gains neither can achieve alone.


Method

A Three-Stage Adaptive Inference Framework

Our framework fine-tunes DeiT-S with early exit heads, profiles each dataset's redundancy and complexity characteristics, then deploys a unified inference pipeline that dynamically selects the optimal path per sample at test time.

STAGE 1 · TRAINING DeiT-S Fine-tuning + EE heads at layers 4, 7, 10 Multi-Exit Loss w_final = 1.0 w₄ = w₇ = w₁₀ = 0.3 STAGE 2 · PROFILING Spatial Redundancy (JSD) Train lightweight CNN predictor Threshold Calibration θ_EE per dataset (EE) θ_R per dataset (TR) STAGE 3 · TEST-TIME INFERENCE ŷ_red > θ_R ? Apply Token Reduction c_k > θ_EE ? Early Exit at Layer k Efficient Prediction 46 avg tokens · exit L6.5 71.4% avg FLOPs reduction
1

Training with Early Exit Heads

Fine-tune DeiT-S with lightweight MLP classifiers at layers 4, 7, and 10. All heads are trained simultaneously using a weighted multi-exit loss.

2

Dataset-Specific Profiling

Profile each dataset's spatial redundancy via Jensen-Shannon Divergence and sample complexity via confidence sweep. Calibrate per-dataset thresholds θ_EE and θ_R.

3

Unified Test-Time Inference

A lightweight CNN predictor estimates redundancy in one pass. Redundant samples use TR; high-confidence samples exit early; complex ones go full depth.

Stage 1: Multi-Exit Loss

We attach lightweight two-layer MLP classifiers (with layer normalization) to the CLS token at layers 4, 7, and 10 of DeiT-S. All heads are trained jointly:

$\mathcal{L}_\text{total} = w_\text{final} \cdot \mathcal{L}_\text{final} + \sum_{k \in \{4,7,10\}} w_k \cdot \mathcal{L}_k$

where $w_4 = w_7 = w_{10} = 0.3$ and $w_\text{final} = 1.0$.

Stage 2: Spatial Redundancy Profiling via JSD

For each validation sample, we compute a ground-truth redundancy score from attention distributions:

$y_\text{red} = 1 - \tfrac{1}{3}\big(\text{JSD}(\mathbf{a}_1, \mathbf{a}_4) + \text{JSD}(\mathbf{a}_4, \mathbf{a}_7) + \text{JSD}(\mathbf{a}_7, \mathbf{a}_{10})\big)$

High $y_\text{red}$ (close to 1) indicates low divergence between attention patterns — i.e., the image is spatially redundant and tokens can be safely pruned. A lightweight CNN is then trained to predict $\hat{y}_\text{red}$ directly from input pixels, avoiding full forward passes at test time.

Dataset redundancy and complexity analysis

Figure 1. Dataset redundancy and complexity analysis across DeiT-S layers. (a) Token-level cosine similarity across layer transitions — RetinaMNIST and ISIC2019 exhibit high initial similarity (~0.8), while INSIGHT, PathMNIST, and PneumoniaMNIST start lower (~0.6). All datasets show monotonically increasing similarity. (b) Sample-wise confidence at EE checkpoints (layers 4, 7, 10, 12) for easy samples (90th pct, lighter) and hard samples (10th pct, darker). PathMNIST and PneumoniaMNIST achieve high early confidence; RetinaMNIST and INSIGHT maintain persistent easy-hard gaps, requiring deeper inference.

Stage 3: Unified Inference Pipeline

At test time, for each input image $\mathbf{x}$:

  1. Score Predictor estimates $\hat{y}_\text{red}$; set use_tr = (ŷ_red > θ_R)
  2. Process ViT layer-by-layer through 12 transformer blocks
  3. At each checkpoint (layers 4, 7, 10): if $c_k > \theta_\text{EE}$ → exit and return prediction; else if TR active → prune tokens before next block
  4. If no early exit, use final layer 12 head

This pipeline ensures: spatially redundant samples use TR → high-confidence samples exit early → complex samples receive full computation.


Results

Superior Efficiency–Accuracy Trade-offs

We evaluate on five medical imaging datasets: ISIC2019 (skin lesions, 9-class), PathMNIST (colon tissue, 9-class), PneumoniaMNIST (pneumonia, binary), RetinaMNIST (diabetic retinopathy, 5-class), and INSIGHT (cataract, real-world clinical data, 4-class).

71.4%
Average FLOPs reduction
across all 5 datasets
0.1pp
Average accuracy loss
(near-zero degradation)
2.07×
Average GPU speedup
over full-model baseline
54.3%
Average energy reduction
at 4.6 mJ (stable)
🏆

Our TR+EE framework outperforms all baselines: EE-only achieves 55.9% FLOPs reduction, TR-only achieves 57.7%, and the state-of-the-art A-ViT achieves only 28.9%. By combining both strategies with dataset-specific thresholds, we exceed these bounds while maintaining diagnostic accuracy within 0.1pp on average.

Unified Framework Performance

Dataset Strategy Accuracy (%) Avg Tokens Avg Exit Layer
ISIC2019Baseline54.219612.0
EE-only56.8 (↑2.6pp)1969.26
TR-only56.8 (↑2.6pp)4012.0
A-ViT56.8 (↑2.6pp)19612.0
TR+EE (Ours)51.720.810.3
PneumoniaMNISTBaseline90.119612.0
EE-only87.81961.86
TR-only92.1 (↑2.0pp)4012.0
A-ViT92.0 (↑1.9pp)19612.0
TR+EE (Ours)88.856.34.65
RetinaMNISTBaseline59.019612.0
EE-only61.0 (↑2.0pp)1965.45
TR-only54.54012.0
A-ViT57.019612.0
TR+EE (Ours)60.8 (↑1.8pp)44.87.7
PathMNISTBaseline94.719612.0
EE-only94.419611.0
TR-only93.54012.0
A-ViT92.019612.0
TR+EE (Ours)96.0 (↑1.3pp)793.0
INSIGHT
(real-world cataract)
Baseline86.119612.0
EE-only87.4 (↑1.3pp)1967.04
TR-only85.54012.0
A-ViT84.919612.0
TR+EE (Ours)86.2 (↑0.1pp)29.26.9
Average (All Datasets) −0.1pp 46.0 6.5

Computation Cost Analysis

Dataset Strategy FLOPs (G) Latency (ms) Energy (mJ) Speedup
ISIC2019Baseline4.611.2019.7761.00×
A-ViT3.742 (↓18.8%)1.43111.7240.84×
TR+EE (Ours)1.569 (↓66.0%)0.6294.8861.91×
PneumoniaMNISTBaseline4.611.1889.4041.00×
A-ViT3.032 (↓34.2%)1.41411.2500.84×
TR+EE (Ours)1.204 (↓73.9%)0.5754.5032.07×
RetinaMNISTBaseline4.611.1899.1161.00×
A-ViT4.320 (↓6.3%)1.420353.2330.84×
TR+EE (Ours)1.398 (↓70.0%)0.6204.8531.92×
PathMNISTBaseline4.611.22512.4011.00×
A-ViT2.386 (↓48.2%)1.46812.6110.83×
TR+EE (Ours)1.050 (↓77.2%)0.4754.2432.58×
INSIGHTBaseline4.611.1969.7761.00×
A-ViT2.902 (↓37.0%)1.42112.5610.84×
TR+EE (Ours)1.394 (↓69.8%)0.6314.5681.90×

Visualizations

Visualization of adaptive framework

Figure 4. Visualization of adaptive framework behavior on INSIGHT and PathMNIST samples. (a) TR-only path: sequential token reduction at 40% keep rate across all checkpoints, preserving informative pupil regions through layer 12. (b)–(c) Combined TR+EE examples: TR activates at the first checkpoint, followed by early exit — clear cataract evidence or salient tissue structure enables confident prediction without deeper layers.

Token reduction strategy comparison

Figure 3. Token reduction strategy comparison across medical imaging datasets. PathMNIST (left) tolerates aggressive reduction — all strategies maintain >99.7% accuracy at 40 tokens. INSIGHT (right) shows greater sensitivity; EViT proves most stable, maintaining 86.3% at 40 tokens (−0.9pp). ToMe collapses below 76 tokens on INSIGHT.


Citation

Cite This Work

If this work is useful in your research, please cite:

@inproceedings{byun2026adaptive,
  title     = {Adaptive Inference for Medical Vision Transformers:
               Token Reduction or Early Exit?},
  author    = {Byun, Ji Young and Lee, HyunSeo and Shuff, Jordan and
               Venkatesh, Rengaraj and Shekhawat, Nakul S. and
               Parikh, Kunal S. and Chellappa, Rama},
  booktitle = {Medical Imaging with Deep Learning (MIDL)},
  year      = {2026}
}
Acknowledgments. We acknowledge support from the National Eye Institute (P30EY001765, R21EY034343), VentureWell Propel Award, Microsoft Acceleration Award, Stephen F Raab and Mariellen Brickley-Raab Rising Professorship in Ophthalmology, Johns Hopkins University, the National Academy of Medicine, and Johns Hopkins University AITC (P30AG073104). Ji Young Byun was supported in part by a discretionary fund at Johns Hopkins University's Whiting School of Engineering.