Skip to main navigation Skip to search Skip to main content

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

  • Zewen Li
  • , Zitong Yu*
  • , Qilang Ye
  • , Weicheng Xie
  • , Wei Zhuo
  • , Linlin Shen*
  • *Corresponding author for this work

Research output: Journal PublicationArticlepeer-review

Abstract

The robust causal capability of multimodal large language models (MLLMs) holds the potential of detecting defective objects in industrial anomaly detection (IAD). However, most traditional IAD methods lack the ability to provide multiturn human–machine dialogs and detailed descriptions, such as the color of objects, the shape of an anomaly, or specific types of anomalies. At the same time, methods based on large pretrained models have not fully stimulated the ability of large models in anomaly detection tasks. In this article, we explore the combination of rich text semantics with both image-level and pixel-level information from images and propose IAD-GPT, a novel paradigm based on MLLMs for IAD. We employ abnormal prompt generator (APG) to generate detailed anomaly prompts for specific objects. These specific prompts from the large language model (LLM) are used to activate the detection and segmentation functions of the pretrained visual-language model (i.e., CLIP). To enhance the visual grounding ability of MLLMs, we propose text-guided enhancer (TGE), wherein image features interact with normal and abnormal text prompts to dynamically select enhancement pathways, which enables language models to focus on the specific aspects of visual data, enhancing their ability to accurately interpret and respond to anomalies within images. Moreover, we design a multimask fusion (MMF) module to incorporate mask as expert knowledge, which enhances the LLM’s perception of pixel-level anomalies. Extensive experiments on MVTec-AD and VisA datasets demonstrate our state-of-the-art performance on self-supervised and few-shot anomaly detection and segmentation tasks, such as MVTec-AD and VisA datasets.

Original languageEnglish
Article number5049512
JournalIEEE Transactions on Instrumentation and Measurement
Volume74
DOIs
Publication statusPublished - 2025
Externally publishedYes

Free Keywords

  • Few-shot anomaly detection
  • multimodal large language model (MLLM)
  • self-supervised anomaly detection

ASJC Scopus subject areas

  • Instrumentation
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection'. Together they form a unique fingerprint.

Cite this