TNT: Nested Hierarchical Attention for Fine-Grained Visual Recognition

Transformer in Transformer

Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, Yunhe Wang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Transformer iN Transformer (TNT), a novel nested architecture for vision tasks. It enhances the standard Vision Transformer (ViT) by modeling both global dependencies between image patches and local interactions within those patches, achieving 81.5% top-1 accuracy on ImageNet.

TL;DR

The Transformer iN Transformer (TNT) architecture rethinks the "patch-as-token" paradigm in Vision Transformers. By nesting a sub-transformer within each standard transformer block to process internal "visual words," TNT captures the fine-grained local details that standard ViTs ignore. This approach yields a significant +1.7% accuracy boost on ImageNet over DeiT while maintaining similar FLOPs.

Problem & Motivation: The Semantic Gap in Patch-Based ViT

The breakthrough of the Vision Transformer (ViT) was treating an image as a sequence of patches. However, from an academic perspective, this paradigm has a flaw: it assumes that a single linear projection can capture all relevant information within a patch.

In natural images, details like texture, color gradients, and local shapes are crucial. When we look at a "patch," we often see smaller repeating patterns or specific structures. By treating the patch as an indivisible unit, standard ViTs lose the local inductive bias that made CNNs so successful. The authors of TNT argue that we need to model the attention inside these patches to truly narrow the gap between pixels and semantic labels.

Methodology: The "Word-Sentence" Hierarchy

The core innovation is the TNT Block, which consists of two data flows:

  1. Inner Transformer (Local): Each patch is further subdivided into "visual words." The Inner Transformer calculates self-attention between these words for each patch independently.
  2. Outer Transformer (Global): This operates on the "visual sentences" (patch embeddings).

The bridge between these two flows is a linear projection that aggregates the transformed word embeddings and adds them to the sentence embedding, effectively "enriching" the global representation with local insights.

TNT Framework Overview

Complexity Management

A common concern with nested models is the computational explosion. TNT solves this by using a shared inner transformer for all patches and keeping the word embedding dimension () significantly smaller than the sentence dimension (). For TNT-S, while , resulting in only a 1.14x increase in FLOPs compared to a standard transformer.

Experiments & Results: SOTA Performance

TNT consistently outperforms pure transformer baselines (ViT, DeiT) and competitive CNNs across various scales.

  • ImageNet Accuracy: TNT-S (23.8M params) hits 81.5%, surpassing DeiT-S (79.8%) and ResNet-50 (76.2%).
  • Efficiency: Despite the nested structure, TNT achieves better throughput than PVT (Pyramid Vision Transformer) at similar accuracy levels.

Accuracy vs FLOPs Comparison

Ablation Insight: Position Encodings

TNT uses a dual-level position encoding:

  • Sentence Encodings maintain global spatial layout.
  • Word Encodings preserve local relative positions within a patch, which are shared across all patches. Removing either leads to a drop in performance, proving that spatial context is hierarchical in images.

Visualizing the "Inner" Attention

Visualization reveals that the Inner Transformer focuses on pixels with similar appearances or textures within the patch. In the deeper layers, these local representations become more abstract, capturing the "essence" of the local structure which then informs the global classification token.

Inner Attention Visualization

Deep Insight & Conclusion

The significance of TNT lies in its recognition that tokenization is not a one-step process. By introducing a nested hierarchy, TNT effectively reintroduces a form of "local connectivity" (similar to CNNs) into the Transformer framework without sacrificing the flexibility of self-attention.

Takeaway: For tasks requiring fine-grained detail (like medical imaging or high-res object detection), the TNT architecture provides a superior backbone compared to the "flat" tokenization of vanilla ViTs.

Limitations: While TNT is efficient in terms of FLOPs, the nested loops in implementation can sometimes lead to lower GPU utilization compared to highly optimized flat attention layers, though the authors mitigate this with shared weights across patches.

Find Similar Papers

Try Our Examples

  • Search for recent vision transformer architectures that utilize hierarchical or nested attention mechanisms to improve local feature extraction beyond Transformer iN Transformer (TNT).
  • Which paper first proposed the concept of "visual words" and "visual sentences" in the context of self-attention, and how does TNT's implementation differ from earlier hierarchical ViTs like Swin Transformer?
  • Explore the application of nested transformer architectures in multi-modal tasks like Video-Language pre-training or Medical Image Analysis.
Contents
TNT: Nested Hierarchical Attention for Fine-Grained Visual Recognition
1. TL;DR
2. Problem & Motivation: The Semantic Gap in Patch-Based ViT
3. Methodology: The "Word-Sentence" Hierarchy
3.1. Complexity Management
4. Experiments & Results: SOTA Performance
4.1. Ablation Insight: Position Encodings
5. Visualizing the "Inner" Attention
6. Deep Insight & Conclusion