NeurIPS 2026

ViT-AdaLA: Adapting Vision Transformers with Linear Attention

A three-stage alignment framework that transfers the knowledge of pretrained vision transformers to linear-attention backbones.

Yifan Li1, Seunghyun Yoon2, Viet Dac Lai2, Franck Dernoncourt2, Jason Kuen2, Yu Kong1, Trung Bui2

1 Michigan State University    2 Adobe Research

ViT-AdaLA aligns each linear-attention block to a frozen softmax teacher, then aligns full-network features before downstream fine-tuning.
ViT-AdaLA progressively aligns attention modules and final-layer features before transferring the linearized backbone to downstream tasks.

Linearize a pretrained ViT without discarding its visual prior.

Reuse pretrained VFMs

Begin with an existing softmax-based vision transformer instead of training a linear-attention model from scratch.

Align locally and globally

Match attention blocks first, then recover the complete model’s feature representation with a frozen teacher.

Transfer across tasks

Use the adapted backbone for classification and segmentation; the alignment recipe also works with other attention modules.

Abstract

Vision foundation models built on Vision Transformers perform well across many tasks, but softmax attention becomes expensive for long visual sequences. Existing linear-attention ViTs often require training from scratch, while linearization methods designed for language models do not transfer well to vision. ViT-AdaLA adapts a pretrained ViT in three stages: attention alignment, feature alignment, and supervised fine-tuning. The first stage matches individual linear-attention blocks to the original softmax blocks; the second corrects accumulated errors by matching final-layer features to a frozen teacher. Experiments on classification and segmentation show that this simple, model-agnostic adaptation strategy preserves much of the foundation model’s performance while improving scalability at higher resolution.

Motivation

High-resolution images produce longer visual sequences, making quadratic softmax attention increasingly costly. Can we obtain linear-attention efficiency without discarding the pretrained visual knowledge that makes a VFM useful?

Figure 2(a): training a linearized ViT from scratch versus adapting a pretrained ViT into a linearized model.
(a) From-scratch training vs. linearization
Figure 2(b): a decoder-only LLM generates tokens, whereas a ViT produces spatial features for a separate task head; the diagram contrasts temporal and spatial error effects.
(b) Temporal vs. spatial error propagation

Reuse knowledge, don’t relearn it

Training a new linear-attention backbone from scratch must first acquire its own visual prior. Instead, we start from an existing VFM such as DINOv2 or CLIP and transfer that prior to a linearized model. Adaptation still requires training; the goal is to reuse pretrained knowledge.

Preserve features, not just attention

A ViT supplies spatial features to downstream task heads. Attention replacement can alter patch relationships and intermediate features across layers, so a good local approximation need not preserve the full encoder’s representation. This motivates attention alignment followed by global feature alignment.

Accumulated representation drift is not exclusive to ViTs. Here, the focus is preserving spatial consistency and transferable visual features.

Our goal: linearize the vision foundation models without training from scratch.

Method

The main contribution is a progressive adaptation paradigm, not a new linear-attention architecture. Vanilla linear attention is the default instance.

Stage 1

Attention alignment

Each linear-attention block sees the same input as its frozen softmax counterpart. We tune its Q, K, and V projections to match the attention output.

Local alignment · COCO

Stage 2

Feature alignment

We update the full linearized student to match the frozen teacher’s final features, correcting approximation errors accumulated across blocks.

Global alignment · ImageNet-22K

Stage 3

Downstream fine-tuning

A task head and the adapted backbone are fine-tuned for classification or segmentation. The Stage-2 backbone can also be used frozen.

Task transfer

Experimental Results

We evaluate DINOv2-L, CLIP-L, SigLIP-L, and an ImageNet-1K-pretrained ViT-L. Headline results and detailed baseline comparisons use DINOv2-L. Within each task and backbone, methods share downstream data, task heads, and full fine-tuning budgets; baselines retain their native conversion procedures.

86.0%ImageNet-1K top-1
86.8% softmax teacher
55.55ADE20K mIoU
56.73 softmax teacher
78.73Cityscapes mIoU at 1024²
2.26× softmax throughput

Classification and segmentation (DINOv2-L)

View exact data
MethodImageNet-1K top-1 (%) ↑ADE20K mIoU ↑
Softmax teacher86.856.73
LoLCATS61.617.42
Monarch (native)82.744.95
ViT-AdaLA86.055.55

Transfer across pretrained vision backbones

The same adaptation recipe applies to DINOv2, CLIP, SigLIP, and a supervised ImageNet-1K ViT. Each backbone is compared with its own softmax teacher.

View exact data
BackboneImageNet-1K top-1 (%) ↑ADE20K mIoU ↑Cityscapes mIoU ↑
TeacherViT-AdaLATeacherViT-AdaLATeacherViT-AdaLA
DINOv2-L86.886.056.7355.5580.9878.73
CLIP-L86.485.5————
SigLIP-L86.986.454.4053.1676.5373.33
IN-1K ViT-L——50.8349.7572.2168.25

ImageNet-1K and ADE20K use 512² inputs; Cityscapes uses 1024². These are downstream full fine-tuning results, not frozen-backbone evaluations. Backbones without reported results for the selected dataset are omitted. Classification and ADE20K values follow the main comparison tables.

High-resolution segmentation on Cityscapes

DINOv2-L with the same Mask2Former head and downstream fine-tuning budget. We report both 512 × 512 and 1024 × 1024 input resolutions.

View exact data
Method512² mIoU ↑1024² mIoU ↑1024² throughput (images/s) ↑1024² memory (GB) ↓
Softmax teacher74.8680.987.073.2836
LoLCATS34.7133.6615.381.3823
Nyströmformer (native)48.9747.5614.611.3764
Monarch (native)50.8455.189.741.5469
ViT-AdaLA72.4078.7315.951.3764

At 1024², ViT-AdaLA achieves 78.73 mIoU with 2.26× the softmax teacher’s throughput and 58.1% lower peak memory. Throughput and memory use batch size 1.

Ablation & Key Insights

Keeping the linear-attention architecture fixed reveals two distinct roles: Stage 1 provides local initialization, while Stage 2 recovers end-to-end compatibility with the pretrained encoder.

What does each alignment stage contribute?

View exact data
AlignmentStage 1Stage 2ADE20K mIoU ↑SAM-HQ IoU ↑
No alignment✗✗22.9263.65
Stage 1 only✓✗19.3772.00
Stage 2 only✗✓52.4673.47
Stage 1 + Stage 2✓✓55.5577.48

ADE20K uses DINOv2-L with the same downstream fine-tuning protocol across rows; the softmax reference is 56.73 mIoU. SAM-H is trained on an internal foreground-segmentation dataset and evaluated on SAM-HQ. “No alignment” skips Stages 1–2, not downstream training.

+29.54 mIoU

Stage 2 drives representation recovery

Stage 2 alone raises ADE20K performance from 22.92 to 52.46, closing 87.4% of the gap to the softmax teacher without changing the attention architecture. This supports pretrained-feature compatibility, rather than attention capacity alone, as a key bottleneck.

+3.09 mIoU / +4.01 IoU

Stage 1 is complementary, not sufficient

Adding Stage 1 improves Stage 2 on both ADE20K and SAM-HQ. Its standalone effect depends on the task: ADE20K drops from 22.92 to 19.37, while foreground-segmentation IoU rises from 63.65 to 72.00 (+8.35).

Local → global alignment

Match the student’s actual forward pass

Stage 1 matches attention on shared teacher inputs. Once all blocks are replaced, the student receives altered intermediate features. Stage 2 supervises the final representation on those actual inputs, allowing the full encoder to co-adapt instead of requiring every block to remain identical.

Why adapt more than QKV?

On DINOv2-L, full-network adaptation outperforms QKV-only tuning by 9.6 accuracy points and 6.73 mIoU. This supports adapting the wider encoder to the changed attention computation, rather than only its attention projections.

Tuning scopeImageNet-1K top-1 (%) ↑ADE20K mIoU ↑
QKV only76.448.82
Full network86.055.55

DINOv2-L at 512²; this is a tuning-scope comparison, separate from the Stage 1/2 ablation above.

Stage 1 also accelerates feature alignment

DINOv2-L Stage-2 feature-MSE training curves with and without Stage 1 initialization; the Stage-1-initialized curve decreases faster.
With Stage 1 initialization, Stage-2 feature-MSE decreases faster. Together with the task ablation, this supports Stage 1 as a useful initializer—not an independently reliable adapted model.

The central lesson

Linearizing a VFM is a representation-transfer problem, not just an operator-approximation problem.

Local attention alignment becomes most useful when followed by global feature recovery.

Efficiency and Resolution

Linear attention is not faster in every setting. On one H100 with DINOv2-L at batch size 1, the throughput crossover occurs between 256² and 512² input resolution.

Input resolutionSoftmax (images/s)ViT-AdaLA (images/s)Relative throughput
256²62.8349.280.78×
512²36.5241.561.14×
1024²7.0715.952.26×

At 1024² on Cityscapes, peak memory falls from 3.2836 GB to 1.3764 GB (−58.1%), while mIoU changes from 80.98 to 78.73. All throughput measurements above are end-to-end, batch size 1.

Preserving Transferable Features

After Stage 2, the DINOv2-L backbone can remain frozen while only a task head is trained. This is distinct from the full fine-tuning results above.

PCA-projected feature maps from the softmax DINOv2-L teacher, ViT-AdaLA, and Monarch attention.
PCA projections of final-layer DINOv2-L features. ViT-AdaLA more closely tracks the softmax teacher than Monarch in these examples.
Frozen-backbone evaluationSoftmax teacherViT-AdaLARetained quality
ImageNet-1K linear probe82.6681.3098.4%
Flowers-102 linear probe99.7199.3799.7%

The teacher and linearized student use identical heads; all 304M backbone parameters remain frozen.

BibTeX

Cite this work
@inproceedings{li2026vitadala,
  title={ViT-AdaLA: Adapting Vision Transformers with Linear Attention},
  author={Yifan Li and Seunghyun Yoon and Viet Dac Lai and Franck Dernoncourt and Jason Kuen and Yu Kong and Trung Bui},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}