Skip to main navigation Skip to search Skip to main content

Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography

  • Leilei Zeng
  • , Jie Liu
  • , Wenting Chen
  • , Chenyang Lyu
  • , Wenxi Li
  • , Shaonan Liu
  • , Xiande Zhou
  • , Linlin Shen*
  • *Corresponding author for this work

Research output: Journal PublicationArticlepeer-review

1 Citation (Scopus)

Abstract

The exponential growth in computed tomography (CT), coupled with a stagnant number of radiologists, has created an urgent need for autonomous 3D medical image analysis to alleviate the increasing workload in radiology. While medical multi-modal large language models (Med-MLLMs) have emerged as a promising solution, existing approaches predominantly focus on expanding data scale and improving annotation quality. However, these approaches inevitably confront fundamental challenges in medical data acquisition, particularly patient privacy concerns and ethical considerations. To address these challenges, we introduce Body-Prior, a novel framework for 3D CT that shifts the focus from data volume to incorporating structural prior knowledge, thereby enhancing performance without requiring additional data collection. Our approach consists of two key components: (1) the Body-Prior Instruction Dataset (BoiD), which integrates complex anatomical interrelationships and topological features into the instruction dataset, and (2) the Inter-Organ Consistency Constraint (IO-Cons), which enforces part-whole relationships in the feature space to ensure anatomical consistency. Body-Prior processes high-resolution CT scans and generates both textual and visual responses based on given instructions. Extensive experiments across various radiological tasks, including semantic segmentation, referring segmentation, and VQA in both closed and open formats, demonstrate that Body-Prior significantly outperforms existing methods. Our framework offers a promising solution to enhance the efficiency and accuracy of radiological workflows. To facilitate Med-MLLMs research, we will release our data, code and model.

Original languageEnglish
Article number113540
JournalPattern Recognition
Volume179
DOIs
Publication statusPublished - Nov 2026
Externally publishedYes

Free Keywords

  • Abdominal multi-organ segmentation
  • Anatomical prior
  • Multimodal large-scale model

ASJC Scopus subject areas

  • Software
  • Signal Processing
  • Computer Vision and Pattern Recognition
  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography'. Together they form a unique fingerprint.

Cite this