TY - GEN
T1 - MiTPose
T2 - 23rd International Conference on Industrial Informatics, INDIN 2025
AU - Wu, Yunfeng
AU - Gao, Qizhong
AU - Liu, Yize
AU - Sun, Jun
AU - Li, Zhuozhi
AU - Jin, Yuhao
AU - Yue, Yong
AU - Zhu, Xiaohui
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Two-dimensional human pose estimation (HPE) has been extensively applied across various domains, including behavioral analysis, identity verification, and automated industrial manufacturing. Compared to convolutional neural networks (CNNs), the Vision Transformer (ViT) has demonstrated impressive results in human pose estimation. However, two main challenges arise: (1) the complexity of image size and parameters grows quadratically, making traditional ViT unsuitable for deployment on edge devices, and (2) the attention mechanism in transformers lacks the ability to capture local fine-grained perception. To address these issues, we propose a novel method Multi-Granularity guided Vision Transformer for human pose estimation (MiTPOSE), which integrates both CNN and transformer for feature encoding. Specifically, we introduce an improved SCConv encoder with Global Response Normalization, which consists of a spatial reconstruction unit and channel reconstruction unit to reduce redundant computations and enhance representative feature learning. Furthermore, we incorporate a novel Multi-Granularity block to address the shortcomings of traditional self-attention mechanisms in capturing local fine-grained details, and a lightweight decoder for keypoint detection. Comprehensive evaluations and tests on the COCO benchmark datasets demonstrate that MiTPose achieves competitive performance in pose estimation compared to state-of-the-art methods.
AB - Two-dimensional human pose estimation (HPE) has been extensively applied across various domains, including behavioral analysis, identity verification, and automated industrial manufacturing. Compared to convolutional neural networks (CNNs), the Vision Transformer (ViT) has demonstrated impressive results in human pose estimation. However, two main challenges arise: (1) the complexity of image size and parameters grows quadratically, making traditional ViT unsuitable for deployment on edge devices, and (2) the attention mechanism in transformers lacks the ability to capture local fine-grained perception. To address these issues, we propose a novel method Multi-Granularity guided Vision Transformer for human pose estimation (MiTPOSE), which integrates both CNN and transformer for feature encoding. Specifically, we introduce an improved SCConv encoder with Global Response Normalization, which consists of a spatial reconstruction unit and channel reconstruction unit to reduce redundant computations and enhance representative feature learning. Furthermore, we incorporate a novel Multi-Granularity block to address the shortcomings of traditional self-attention mechanisms in capturing local fine-grained details, and a lightweight decoder for keypoint detection. Comprehensive evaluations and tests on the COCO benchmark datasets demonstrate that MiTPose achieves competitive performance in pose estimation compared to state-of-the-art methods.
UR - https://www.scopus.com/pages/publications/105032688598
U2 - 10.1109/INDIN64977.2025.11279684
DO - 10.1109/INDIN64977.2025.11279684
M3 - Conference Proceeding
AN - SCOPUS:105032688598
T3 - IEEE International Conference on Industrial Informatics (INDIN)
BT - 2025 IEEE 23rd International Conference on Industrial Informatics, INDIN 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 12 July 2025 through 15 July 2025
ER -