Abstract
The exponential growth in computed tomography (CT), coupled with a stagnant number of radiologists, has created an urgent need for autonomous 3D medical image analysis to alleviate the increasing workload in radiology. While medical multi-modal large language models (Med-MLLMs) have emerged as a promising solution, existing approaches predominantly focus on expanding data scale and improving annotation quality. However, these approaches inevitably confront fundamental challenges in medical data acquisition, particularly patient privacy concerns and ethical considerations. To address these challenges, we introduce Body-Prior, a novel framework for 3D CT that shifts the focus from data volume to incorporating structural prior knowledge, thereby enhancing performance without requiring additional data collection. Our approach consists of two key components: (1) the Body-Prior Instruction Dataset (BoiD), which integrates complex anatomical interrelationships and topological features into the instruction dataset, and (2) the Inter-Organ Consistency Constraint (IO-Cons), which enforces part-whole relationships in the feature space to ensure anatomical consistency. Body-Prior processes high-resolution CT scans and generates both textual and visual responses based on given instructions. Extensive experiments across various radiological tasks, including semantic segmentation, referring segmentation, and VQA in both closed and open formats, demonstrate that Body-Prior significantly outperforms existing methods. Our framework offers a promising solution to enhance the efficiency and accuracy of radiological workflows. To facilitate Med-MLLMs research, we will release our data, code and model.
| Original language | English |
|---|---|
| Article number | 113540 |
| Journal | Pattern Recognition |
| Volume | 179 |
| DOIs | |
| Publication status | Published - Nov 2026 |
| Externally published | Yes |
Free Keywords
- Abdominal multi-organ segmentation
- Anatomical prior
- Multimodal large-scale model
ASJC Scopus subject areas
- Software
- Signal Processing
- Computer Vision and Pattern Recognition
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver