Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens

Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens

TMLR Paper5760 Authors

28 Aug 2025 (modified: 08 Sept 2025)Under review for TMLREveryoneRevisionsBibTeXCC BY 4.0

Abstract: Current approaches for segmenting ultra-high-resolution images either slide a window, thereby discarding global context, or downsample and lose fine detail. We propose a simple yet effective method that brings explicit multi-scale reasoning to vision transformers, simultaneously preserving local details and global awareness. Concretely, we process each image in parallel at a local scale (high-resolution, small crops) and a global scale (low-resolution, large crops), and aggregate and propagate features between the two branches with a small set of learnable relay tokens. The design plugs directly into standard transformer backbones (e.g. ViT and Swin) and adds fewer than 2 % parameters. Extensive experiments on three ultra-high-resolution segmentation benchmarks, Archaeoscape, URUR, and Gleason, and on the conventional Cityscapes dataset show consistent gains, with up to 13 % relative mIoU improvement. Code and pretrained models will be released.

Submission Type: Regular submission (no more than 12 pages of main content)

Previous TMLR Submission Url: https://openreview.net/forum?id=sXdSRQsUWK

Changes Since Last Submission: Re-submission after desk reject due to bad font. Changed font.

Assigned Action Editor: ~Adam_W_Harley1

Submission Number: 5760

Loading