One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
A recurrent vision Transformer matches full-depth accuracy with one block and about 70% fewer parameters.
reViT shows a single Transformer block applied recurrently can match a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. Depth-specific feed-forward layers are convex combinations of a small shared expert bank, programmed by a continuous normalized-depth coordinate. Trained from scratch on ImageNet-1k, reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters. An eight-expert model distilled only from DINOv2 output features retains nearly all of the teacher's linear-probe accuracy and transfers to classification, segmentation, and depth prediction.
- One recurrent Transformer block matches full-depth accuracy at similar FLOPs.
- A normalized depth coordinate mixes a shared bank of FFN experts.
- reViT-B/16 matches DeiT III with about 70% fewer stored parameters.
- An eight-expert model nearly retains its DINOv2 teacher's linear-probe accuracy.
- Elastic-depth training lets one checkpoint operate at multiple tested depths.
Full article197 words · extracted from arxiv.org · click to collapse
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.12448