Dynamic Hierarchical Token Merging for Vision Transformers

Name: Dynamic Hierarchical Token Merging for Vision Transformers
Keywords: Vision Transformers, [INFO.INFO-AI] Computer Science [cs]/Artificial Intelligence [cs.AI], [INFO.INFO-CV] Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], Neural network compression, Dynamic neural networks, Token merging

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Jan 2025Publisher:SCITEPRESS - Science and Technology PublicationsJournal:Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications

Authors: Haroun, Karim; Allenet, Thibault; Ben Chehida, Karim; Martinet, Jean;

doi: 10.5220/0013284100003912

Dynamic Hierarchical Token Merging for Vision Transformers

- Summary
- Subjects
- Related research
  (1)
- Metrics

Abstract

Vision Transformers (ViTs) have achieved impressive results in computer vision, excelling in tasks such as image classification, segmentation, and object detection. However, their quadratic complexity $O(N^2)$, where $N$ is the token sequence length, poses challenges when deployed on resource-limited devices. To address this issue, dynamic token merging has emerged as an effective strategy, progressively reducing the token count during inference to achieve computational savings. Some strategies consider all tokens in the sequence as merging candidates, without focusing on spatially close tokens. Other strategies either limit token merging to a local window, or constrains it to pairs of adjacent tokens, thus not capturing more complex feature relationships. In this paper, we propose Dynamic Hierarchical Token Merging (DHTM), a novel token merging approach, where we advocate that spatially close tokens share more information than distant tokens and consider all pairs of spatially close candidates instead of imposing fixed windows. Besides, our approach draws on the principles of Hierarchical Agglomerative Clustering (HAC), where we iteratively merge tokens in each layer, fusing a fixed number of selected neighbor token pairs based on their similarity. Our proposed approach is off-the-shelf, i.e., it does not require additional training. We evaluate our approach on the ImageNet-1K dataset for classification, achieving substantial computational savings while minimizing accuracy reduction, surpassing existing token merging methods.

Related Organizations

Commissariat à l’Energie Atomique et aux Energies Alternatives
France
Université Côte d'Azur
France
French National Centre for Scientific Research
France
University of Paris-Saclay
France

Keywords

Vision Transformers, [INFO.INFO-AI] Computer Science [cs]/Artificial Intelligence [cs.AI], [INFO.INFO-CV] Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], Neural network compression, Dynamic neural networks, Token merging

1 Research products, page 1 of 1

fvcore software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Average

Green

Dynamic Hierarchical Token Merging for Vision Transformers

Dynamic Hierarchical Token Merging for Vision Transformers

1 Research products, page 1 of 1

fvcore software on GitHub