TY - JOUR
T1 - Unleashing the power of optimal head in CLIP and DINO for weakly supervised semantic segmentation
AU - Qiu, Xianglin
AU - Yu, Siyue
AU - Zhang, Bingfeng
AU - Zhang, Zhen
AU - Tillo, Tammam
AU - Xiao, Jimin
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/10
Y1 - 2026/10
N2 - Weakly-supervised semantic segmentation with image-level labels has gained significant attention due to its low annotation cost. Recent methods leverage frozen CLIP and DINO models to create high-quality pseudo labels for supervised training. They typically use CLIP layer attention as affinity to refine class activation maps (CAMs). However, our investigation reveals that, in the multi-heads self-attention (MHSA) module of CLIP, some heads are noisy to precisely describe the feature semantic relationships, leading to the layer attention, which is obtained by averaging the head attentions, being suboptimal. To address it, we propose a class-aware head selection method that directly selects the head best matching the target class to extract affinity for refining the CAM, thus avoiding the influence of noisy heads. We further extend this method to DINO since we found similar noisy heads issue in it, and design a dual-supervision process which leverages the fact that CLIP captures global semantics while DINO excels in local details to harness their synergy through complementary pseudo-labels. In addition, to enhance the dense semantics of CLIP features in the decoder, we align its pixel features with their corresponding text embeddings that serve as category prototypes, thereby improving the final predictions. By integrating the above strategies, our method, termed UPOH, unleashes the power of optimal head in CLIP and DINO to boost the WSSS performance. Experimental results demonstrate that our method achieves new state-of-the-art performance on PASCAL VOC and MS COCO datasets.
AB - Weakly-supervised semantic segmentation with image-level labels has gained significant attention due to its low annotation cost. Recent methods leverage frozen CLIP and DINO models to create high-quality pseudo labels for supervised training. They typically use CLIP layer attention as affinity to refine class activation maps (CAMs). However, our investigation reveals that, in the multi-heads self-attention (MHSA) module of CLIP, some heads are noisy to precisely describe the feature semantic relationships, leading to the layer attention, which is obtained by averaging the head attentions, being suboptimal. To address it, we propose a class-aware head selection method that directly selects the head best matching the target class to extract affinity for refining the CAM, thus avoiding the influence of noisy heads. We further extend this method to DINO since we found similar noisy heads issue in it, and design a dual-supervision process which leverages the fact that CLIP captures global semantics while DINO excels in local details to harness their synergy through complementary pseudo-labels. In addition, to enhance the dense semantics of CLIP features in the decoder, we align its pixel features with their corresponding text embeddings that serve as category prototypes, thereby improving the final predictions. By integrating the above strategies, our method, termed UPOH, unleashes the power of optimal head in CLIP and DINO to boost the WSSS performance. Experimental results demonstrate that our method achieves new state-of-the-art performance on PASCAL VOC and MS COCO datasets.
KW - CLIP
KW - DINO
KW - Multi-heads
KW - Weakly supervised semantic segmentation
UR - https://www.scopus.com/pages/publications/105033236365
U2 - 10.1016/j.patcog.2026.113454
DO - 10.1016/j.patcog.2026.113454
M3 - Article
AN - SCOPUS:105033236365
SN - 0031-3203
VL - 178
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 113454
ER -