Abstract
Few-shot fine-grained image classification remains a challenging task due to the subtle inter-class variations and the scarcity of labeled data. Existing few-shot fine-grained methods often struggle to generalize effectively under low-data regimes and fail to discriminate between visually similar samples. To address these challenges, we propose DistillCvT, a two-stage framework to enhance model generalization and feature discrimination through self-supervised learning and self-distillation within prototypical network structure. Convolutional Vision Transformer (CvT) is employed as the feature extractor, which integrates convolutional locality with transformer-based global modeling to capture fine-grained details while preserving global context. DistillCvT mitigates these issues through a two-stage strategy. In the first stage, the model is trained using a prototypical distance-based classification loss, where class probabilities are computed from the distances between query samples and class prototypes, complemented by standard data augmentation to mitigate data scarcity and learn discriminative class representations. In the second stage, an identical student model is trained with a self-supervised jigsaw task to enhance feature discrimination and robustness by aligning support samples with their augmented query counterparts. Simultaneously, the teacher model from stage one guides it through logit-based self-distillation to further improve generalization. Experiments on CUB-200-2011, Stanford Dogs, and Stanford Cars demonstrate that DistillCvT consistently improves performance and achieves competitive results compared with recent state-of-the-art few-shot fine-grained image classification methods.
| Original language | English |
|---|---|
| Article number | 106119 |
| Journal | Image and Vision Computing |
| Volume | 174 |
| DOIs | |
| Publication status | Published - Oct 2026 |
Free Keywords
- Few-shot learning
- Fine-grained image classification
- Prototypical network
- Self-distillation
- Self-supervised learning
ASJC Scopus subject areas
- Signal Processing
- Computer Vision and Pattern Recognition
Fingerprint
Dive into the research topics of 'DistillCvT: Self-supervised distillation of Convolutional Vision Transformers for few-shot fine-grained classification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver