Abstract
With the rapid development of Visual Large Language Models (VLLMs), significant improvements have been made in vision tasks such as image description generation and visual question answering. However, for the facial description task, the direct application of VLLMs poses additional challenges due to the lack of fine-grained facial description data and tailored prompts for focusing on facial features. In this work, we first construct a Question–Answer Facial Description dataset (QA-FaceDesc) specifically designed for facial description tasks. Then, to leverage QA-FaceDesc with VLLMs, we propose a Fine-grained Facial Description framework (FFD) that consists of a Facial Detector Module (FDM), a QA-FaceDesc Retrieval Augmented Module (QA-FaceRAM), and a Retrieval Module (RM). In addition, to comprehensively assess performance, we construct three test sets based on existing datasets for fine-grained facial descriptions. Through extensive experiments, our approach outperforms all the baseline models even without VLLMs fine-tuning.
| Original language | English |
|---|---|
| Article number | 133770 |
| Journal | Neurocomputing |
| Volume | 691 |
| DOIs | |
| Publication status | Published - 28 Aug 2026 |
Free Keywords
- Facial description
- Multimodal learning
- Retrieval augmented
- VLLMs
ASJC Scopus subject areas
- Computer Science Applications
- Cognitive Neuroscience
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Fine-grained facial description generation with retrieval augmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver