TY - GEN
T1 - Grounding Before Reporting
T2 - 22nd International Conference on Intelligent Computing, ICIC 2026
AU - Huang, Wenzheng
AU - Yuan, Jiachen
AU - Wang, Qiujin
AU - Ye, Zihan
AU - Fazi, Camilla
AU - Fontanella, Federica
AU - Hernandez-Cruz, Netzahualcoyotl
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2027.
PY - 2027
Y1 - 2027
N2 - Vision-Language Models (VLMs) show promise for automated medical report generation, but conventional joint training often couples fine-grained visual grounding and holistic report generation within a single trainable parameter space. This causes optimization conflicts, where models over-rely on dominant global priors and produce generic reports with degraded anatomical grounding. To address this, we propose Local-to-Global Decoupled Learning (LGDL), a staged framework that preserves local anatomical knowledge during global adaptation. LGDL comprises three components: Local Expert Initialization for organ-level grounding, Decoupled Sequential LoRA Adaptation for separate optimization stages, and Local Knowledge Retention for lightweight replay. Evaluated on two grounded benchmarks: FOCUS (fetal ultrasound) and PadChest-GR (chest X-rays), LGDL improves Global and Local F1-scores by + 0.1237 and + 0.0167 on FOCUS, and + 0.0267 and + 0.0249 on PadChest-GR. These results demonstrate that explicit decoupling effectively mitigates supervision conflicts, yielding more accurate and anatomically grounded reports.
AB - Vision-Language Models (VLMs) show promise for automated medical report generation, but conventional joint training often couples fine-grained visual grounding and holistic report generation within a single trainable parameter space. This causes optimization conflicts, where models over-rely on dominant global priors and produce generic reports with degraded anatomical grounding. To address this, we propose Local-to-Global Decoupled Learning (LGDL), a staged framework that preserves local anatomical knowledge during global adaptation. LGDL comprises three components: Local Expert Initialization for organ-level grounding, Decoupled Sequential LoRA Adaptation for separate optimization stages, and Local Knowledge Retention for lightweight replay. Evaluated on two grounded benchmarks: FOCUS (fetal ultrasound) and PadChest-GR (chest X-rays), LGDL improves Global and Local F1-scores by + 0.1237 and + 0.0167 on FOCUS, and + 0.0267 and + 0.0249 on PadChest-GR. These results demonstrate that explicit decoupling effectively mitigates supervision conflicts, yielding more accurate and anatomically grounded reports.
KW - Decoupled Learning
KW - Medical Report Generation
KW - Parameter-Efficient Fine-Tuning
KW - Vision-Language Models
UR - https://www.scopus.com/pages/publications/105046309554
U2 - 10.1007/978-981-92-3538-4_25
DO - 10.1007/978-981-92-3538-4_25
M3 - Conference Proceeding
AN - SCOPUS:105046309554
SN - 9789819235377
T3 - Communications in Computer and Information Science
SP - 291
EP - 302
BT - Advanced Intelligent Computing Technology and Applications - 22nd International Conference on Intelligent Computing, ICIC 2026, Proceedings
A2 - Huang, De-Shuang
A2 - Zhang, Qinhu
A2 - Pan, Yijie
A2 - Zhang, Chuanlei
A2 - Chen, Wei
A2 - Li, Bo
A2 - Bao, Wenzheng
A2 - Premaratne, Prashan
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 22 July 2026 through 26 July 2026
ER -