Skip to main navigation Skip to search Skip to main content

Rewarding fine-grained image captioning with keyword group contrastive

Research output: Journal PublicationArticlepeer-review

Abstract

Fine-grained image captioning aims to automatically generate a description with detailed information from a given image. The task poses significant challenges, as it requires image captioning models to accurately capture fine-grained details, effectively differentiate between visually similar yet distinct elements within an image, and generate detailed captions that comprehensively describe the image content. In this paper, we propose a novel framework for fine-grained image captioning that combines reinforcement learning and contrastive learning with specifically designed loss and rewards. Specifically, three image captioning objectives are devised: 1) a novel Keyword Group Contrastive loss for token representation learning by leveraging different groups of keywords matched by visual information; 2) a CLIP Contrastive reward encouraging the generated caption to be more similar to its input image and dissimilar to the other images; 3) a Fine-grained Grammar reward using the grammar ELECTRA discriminator for high-quality caption generation with good grammar. We evaluate the performance of our framework on the FineCapEval benchmark dataset and show that it significantly outperforms the existing state-of-the-art methods in terms of describing fine-grained information from its input images.
Original languageEnglish
Article number131405
Number of pages11
JournalExpert Systems with Applications
Volume312
Early online date29 Jan 2026
DOIs
Publication statusPublished - May 2026

Free Keywords

  • Fine-grained image captioning
  • Feature representation
  • Reinforcement learning
  • Contrastive learning
  • Multimodal learning

Fingerprint

Dive into the research topics of 'Rewarding fine-grained image captioning with keyword group contrastive'. Together they form a unique fingerprint.

Cite this