Abstract
Large Vision Models (LVMs), exemplified by the Segment Anything Model (SAM), contain powerful general knowledge from extensive pre-training, yet they often underperform in highly specialized domains. Building large models tailored for each domain is usually impractical due to the substantial cost of data collection and training. Therefore, a key challenge is how to tap into SAM's strong knowledge base and transfer it effectively to new, domain-specific tasks, especially under Cross-Domain or Few-Shot constraints. Previous efforts have leveraged prior knowledge from foundation models for transfer learning; however, they typically target specific tasks and exhibit limited robustness in broader applications. To tackle this issue, we propose a Unified Dense-Guided Semantic Prompting framework (UDG-Prom), a new paradigm for Cross-Domain Few-Shot Segmentation (CD-FSS). First, a Multi-level Adaptation Framework (MAF) is used for integrated feature extraction as prior knowledge. Then, we incorporate a Task-Adaptive Auto Meta Prompt (TA2MP) module to enable the extraction of class-domain-agnostic features and generate high-quality, learnable visual prompts. By combining learnable prompts with a structured model and prototype disentanglement, this method retains SAM's prior knowledge and effectively adapts to CD-FSS through category and domain cues. Extensive experiments on four benchmarks show that our model not only surpasses state-of-the-art CD-FSS approaches but also achieves a remarkable improvement in average accuracy.
| Original language | English |
|---|---|
| Article number | 115207 |
| Journal | Knowledge-Based Systems |
| Volume | 335 |
| DOIs | |
| Publication status | Published - 28 Feb 2026 |
Free Keywords
- Cross-domain
- Few-shot
- Segment anything model
- Semantic segmentation
- Visual prompt
ASJC Scopus subject areas
- Management Information Systems
- Software
- Information Systems and Management
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'UDG-Prom: A unified dense-guided semantic prompting for cross-domain few-shot image segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver