Abstract
Self-supervised learning (SSL) has shown strong potential for learning visual representations from unlabeled data, but transferring these representations effectively to object detection remains challenging. Existing SSL paradigms have complementary strengths and weaknesses: contrastive learning provides strong instance discrimination but lacks spatial sensitivity, while masked image modeling preserves spatial structure but is less effective at distinguishing instances. As a result, current methods often perform well on either classification or detection, but struggle to balance both. To address this problem, we propose Contrastive Masked Histogram-Decoupled Detector (CMHD), a dual-branch SSL framework that combines masked spatial reasoning with gradient-aware contrastive learning. The online branch improves spatial understanding by reconstructing masked regions, while the target branch enhances structural and discriminative representations through gradient-based features. We further introduce a random coordinate attention module to fuse cross-branch features and improve multi-scale representation learning. Extensive experiments on MS COCO, PASCAL VOC, ImageNet-1K, and other benchmarks show that CMHD achieves more balanced performance across classification and detection tasks, outperforming representative pure contrastive learning and masked image modeling baselines. Ablation studies further confirm the effectiveness of each component in the proposed framework.
| Original language | English |
|---|---|
| Article number | 117603 |
| Journal | Signal Processing: Image Communication |
| Volume | 147 |
| DOIs | |
| Publication status | Published - Sept 2026 |
Free Keywords
- Contrastive learning
- Edge-aware feature representations
- Masked image model
- Self-supervised learning
ASJC Scopus subject areas
- Software
- Signal Processing
- Computer Vision and Pattern Recognition
- Electrical and Electronic Engineering
Fingerprint
Dive into the research topics of 'Balancing framework: Enhanced performance through contrastive masked encoders and gradient feature'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver