Skip to main navigation Skip to search Skip to main content

DistillCvT: Self-supervised distillation of Convolutional Vision Transformers for few-shot fine-grained classification

Research output: Journal PublicationArticlepeer-review

Abstract

Few-shot fine-grained image classification remains a challenging task due to the subtle inter-class variations and the scarcity of labeled data. Existing few-shot fine-grained methods often struggle to generalize effectively under low-data regimes and fail to discriminate between visually similar samples. To address these challenges, we propose DistillCvT, a two-stage framework to enhance model generalization and feature discrimination through self-supervised learning and self-distillation within prototypical network structure. Convolutional Vision Transformer (CvT) is employed as the feature extractor, which integrates convolutional locality with transformer-based global modeling to capture fine-grained details while preserving global context. DistillCvT mitigates these issues through a two-stage strategy. In the first stage, the model is trained using a prototypical distance-based classification loss, where class probabilities are computed from the distances between query samples and class prototypes, complemented by standard data augmentation to mitigate data scarcity and learn discriminative class representations. In the second stage, an identical student model is trained with a self-supervised jigsaw task to enhance feature discrimination and robustness by aligning support samples with their augmented query counterparts. Simultaneously, the teacher model from stage one guides it through logit-based self-distillation to further improve generalization. Experiments on CUB-200-2011, Stanford Dogs, and Stanford Cars demonstrate that DistillCvT consistently improves performance and achieves competitive results compared with recent state-of-the-art few-shot fine-grained image classification methods.

Original languageEnglish
Article number106119
JournalImage and Vision Computing
Volume174
DOIs
Publication statusPublished - Oct 2026

Free Keywords

  • Few-shot learning
  • Fine-grained image classification
  • Prototypical network
  • Self-distillation
  • Self-supervised learning

ASJC Scopus subject areas

  • Signal Processing
  • Computer Vision and Pattern Recognition

Fingerprint

Dive into the research topics of 'DistillCvT: Self-supervised distillation of Convolutional Vision Transformers for few-shot fine-grained classification'. Together they form a unique fingerprint.

Cite this