Abstract
Zero-shot Referring Expression Comprehension (REC) aims to locate the referring object in an image that corresponds to the given natural language description (Referring Expression). This demands the ability of fine-grained instance recognition within complex visual scenes and textual descriptions, as well as comprehension of the relationships between different instances. Existing zero-shot REC methods based on large-scale Visual Language Alignment (VLA) model perform poorly when encountering ambiguous referring expression. To address this issue, we propose a zero-shot referring expression comprehension paradigm guided by Multimodal Large Language Model (MLLM), which introduced a text expansion, and a new activation map aggregation method for VLA model. We leverages the image-to-text generation capabilities of MLLM to refine the referring expression and predict the category of the referring object. In addition, we present an effective activation map aggregation method for VLA model, which extracts higher-quality activation map related to the referring expression from VLA model and, matches the map with the candidate boxes generated by an open-vocabulary detection model to locate the referring object. Extensive experiments demonstrate that our method outperforms the most advanced state-of-the-art methods on the RefCOCO/+/g datasets, with a performance improvement of up to 11.07%.
| Original language | English |
|---|---|
| Article number | 114223 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| Publication status | Published - Dec 2026 |
| Externally published | Yes |
Free Keywords
- GradCAM
- Multimodal Large Language Models
- Referring expression comprehension
- Zero-shot
ASJC Scopus subject areas
- Software
- Signal Processing
- Computer Vision and Pattern Recognition
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Zero-shot referring expression comprehension via guidance of Multimodal Large Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver