Skip to main navigation Skip to search Skip to main content

Progressive cross-scale semantic alignment for language-guided medical image segmentation

  • Hengzhi Xue
  • , Yin Dai
  • , Qingyong Li
  • , Yudong Yao
  • , Yueyang Teng
  • Northeastern University China

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Language-guided medical image segmentation aims to leverage clinical text as high-level semantic priors to disambiguate visually confusing regions. Yet, its practical adoption is often hindered by two issues: (i) language cues are injected only once and may fade during progressive decoding, and (ii) fusion designs are frequently tailored to specific backbones, limiting portability. To address these challenges, we propose PCSA-Seg, i.e., a backbone-agnostic framework that performs progressive cross-stage semantic alignment for reliable vision-language dense prediction. PCSA-Seg introduces a progressive semantic feature calibration (PSFC) block to propagate and refine semantic priors through hierarchical features, and a lightweight text-guided modulation (TGM) module that combines explicit semantic attention (ESA) with implicit cross-modal alignment (ICA) to provide complementary token-aware guidance and correspondence reinforcement. Our evaluations on the QaTa-COV19 dataset based on X-ray images showed that PCSA-Seg improved the Dice coefficient of the vision-only baseline model from 87.85% to 91.43% through effective multi-stage language conditioning. Experiments on the CT-based MosMedData+ dataset further demonstrated its performance on COVID-19 infection segmentation, with the Dice coefficient reaching 78.60%. Across these two COVID-19 benchmarks, PCSA-Seg delivers competitive segmentation accuracy, while extensive ablations substantiate the contribution of each component and indicate a favorable accuracy-efficiency balance across different visual backbones.

Original languageEnglish
Article number115745
JournalKnowledge-Based Systems
Volume340
DOIs
StatePublished - 12 May 2026

Keywords

  • Language-guided segmentation
  • Medical image segmentation
  • Vision-language model

Fingerprint

Dive into the research topics of 'Progressive cross-scale semantic alignment for language-guided medical image segmentation'. Together they form a unique fingerprint.

Cite this