Skip to main navigation Skip to search Skip to main content

Grounding Before Reporting: Local-to-Global Decoupled Learning for Medical Vision-Language Models

  • Wenzheng Huang
  • , Jiachen Yuan
  • , Qiujin Wang
  • , Zihan Ye
  • , Camilla Fazi
  • , Federica Fontanella
  • , Netzahualcoyotl Hernandez-Cruz*
  • *Corresponding author for this work
  • Xi'an Jiaotong-Liverpool University
  • University of Florence
  • Azienda Ospedaliera Careggi
  • University of Groningen

Research output: Chapter in Book or Report/Conference proceedingConference Proceedingpeer-review

Abstract

Vision-Language Models (VLMs) show promise for automated medical report generation, but conventional joint training often couples fine-grained visual grounding and holistic report generation within a single trainable parameter space. This causes optimization conflicts, where models over-rely on dominant global priors and produce generic reports with degraded anatomical grounding. To address this, we propose Local-to-Global Decoupled Learning (LGDL), a staged framework that preserves local anatomical knowledge during global adaptation. LGDL comprises three components: Local Expert Initialization for organ-level grounding, Decoupled Sequential LoRA Adaptation for separate optimization stages, and Local Knowledge Retention for lightweight replay. Evaluated on two grounded benchmarks: FOCUS (fetal ultrasound) and PadChest-GR (chest X-rays), LGDL improves Global and Local F1-scores by + 0.1237 and + 0.0167 on FOCUS, and + 0.0267 and + 0.0249 on PadChest-GR. These results demonstrate that explicit decoupling effectively mitigates supervision conflicts, yielding more accurate and anatomically grounded reports.

Original languageEnglish
Title of host publicationAdvanced Intelligent Computing Technology and Applications - 22nd International Conference on Intelligent Computing, ICIC 2026, Proceedings
EditorsDe-Shuang Huang, Qinhu Zhang, Yijie Pan, Chuanlei Zhang, Wei Chen, Bo Li, Wenzheng Bao, Prashan Premaratne
PublisherSpringer Science and Business Media Deutschland GmbH
Pages291-302
Number of pages12
ISBN (Print)9789819235377
DOIs
Publication statusPublished - 2027
Event22nd International Conference on Intelligent Computing, ICIC 2026 - Toronto, Canada
Duration: 22 Jul 202626 Jul 2026

Publication series

NameCommunications in Computer and Information Science
Volume3035 CCIS
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference22nd International Conference on Intelligent Computing, ICIC 2026
Country/TerritoryCanada
CityToronto
Period22/07/2626/07/26

Keywords

  • Decoupled Learning
  • Medical Report Generation
  • Parameter-Efficient Fine-Tuning
  • Vision-Language Models

Cite this