TY - GEN
T1 - MedDistilFuse
T2 - 2025 6th International Conference on Computer Vision and Data Mining, ICCVDM 2025
AU - Kuang, Libing
AU - Lim, Kian Ming
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Medical image classification plays an important role in computer-aided diagnosis by enabling automated identification and categorization of disease-related patterns in imaging data. Despite its importance, the field faces significant challenges, including data scarcity, high annotation costs, and poor model generalization. To address these issues, we propose MedDistilFuse, a novel hybrid framework that integrates knowledge distillation with contrastive representation learning. Specifically, MedDistilFuse leverages a pre-trained Contrastive Language- Image Pre-training (CLIP) model as the teacher to transfer rich visual knowledge to a lightweight student network based on Vision Transformers (ViT). To enhance representational capacity with minimal complexity, a Kolmogorov-Arnold Network (KAN)-inspired nonlinear module is incorporated into the student model. Furthermore, a contrastive learning objective is introduced to promote feature discriminability and robustness under sparse supervision. Extensive experiments across five benchmark medical imaging datasets demonstrate the proposed MedDistilFuse consistently achieves promising accuracy and top or competitive AUC scores. GitHub link: the link will be provided upon paper acceptance.
AB - Medical image classification plays an important role in computer-aided diagnosis by enabling automated identification and categorization of disease-related patterns in imaging data. Despite its importance, the field faces significant challenges, including data scarcity, high annotation costs, and poor model generalization. To address these issues, we propose MedDistilFuse, a novel hybrid framework that integrates knowledge distillation with contrastive representation learning. Specifically, MedDistilFuse leverages a pre-trained Contrastive Language- Image Pre-training (CLIP) model as the teacher to transfer rich visual knowledge to a lightweight student network based on Vision Transformers (ViT). To enhance representational capacity with minimal complexity, a Kolmogorov-Arnold Network (KAN)-inspired nonlinear module is incorporated into the student model. Furthermore, a contrastive learning objective is introduced to promote feature discriminability and robustness under sparse supervision. Extensive experiments across five benchmark medical imaging datasets demonstrate the proposed MedDistilFuse consistently achieves promising accuracy and top or competitive AUC scores. GitHub link: the link will be provided upon paper acceptance.
KW - Contrastive learning
KW - Deep learning
KW - Knowledge distillation
KW - Kolmogorov-Arnold Network
KW - Medical image classification
UR - https://www.scopus.com/pages/publications/105031618123
U2 - 10.1109/ICCVDM66874.2025.11290036
DO - 10.1109/ICCVDM66874.2025.11290036
M3 - Conference contribution
AN - SCOPUS:105031618123
T3 - 2025 6th International Conference on Computer Vision and Data Mining, ICCVDM 2025
SP - 6
EP - 11
BT - 2025 6th International Conference on Computer Vision and Data Mining, ICCVDM 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 12 September 2025 through 14 September 2025
ER -