Abstract
Fine-grained image captioning aims to automatically generate a description with detailed information from a given image. The task poses significant challenges, as it requires image captioning models to accurately capture fine-grained details, effectively differentiate between visually similar yet distinct elements within an image, and generate detailed captions that comprehensively describe the image content. In this paper, we propose a novel framework for fine-grained image captioning that combines reinforcement learning and contrastive learning with specifically designed loss and rewards. Specifically, three image captioning objectives are devised: 1) a novel Keyword Group Contrastive loss for token representation learning by leveraging different groups of keywords matched by visual information; 2) a CLIP Contrastive reward encouraging the generated caption to be more similar to its input image and dissimilar to the other images; 3) a Fine-grained Grammar reward using the grammar ELECTRA discriminator for high-quality caption generation with good grammar. We evaluate the performance of our framework on the FineCapEval benchmark dataset and show that it significantly outperforms the existing state-of-the-art methods in terms of describing fine-grained information from its input images.
| Original language | English |
|---|---|
| Article number | 131405 |
| Number of pages | 11 |
| Journal | Expert Systems with Applications |
| Volume | 312 |
| Early online date | 29 Jan 2026 |
| DOIs | |
| Publication status | Published - May 2026 |
Free Keywords
- Fine-grained image captioning
- Feature representation
- Reinforcement learning
- Contrastive learning
- Multimodal learning
Fingerprint
Dive into the research topics of 'Rewarding fine-grained image captioning with keyword group contrastive'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver