Early-Stage Integration of Attention Mechanisms in OmniPose for Human Pose Estimation
Human pose estimation is one of the most extensively studied problems in computer vision. However, it is noted that a crucial evaluation benchmark is still very difficult and challenging due to large variation in poses, number of persons interacting, and frequent occlusions. Recently, attention modules like SE, CBAM, and ELA have shown great potential for feature extraction improvements particularly for HPE models and more broadly for computer vision tasks. This paper studies the effects that integrating these attention modules at primary levels of feature extraction in the OmniPose model brings about. Three variants of OmniPose with SEBlock, CBAM, and ELA included are described, which were trained under exactly similar conditions to make a fair comparison among them. SE yields the best overall accuracy, ELA balances accuracy and training efficiency at larger batch sizes, while CBAM provides improved stability but sacrifices final accuracy. Such results underscore the fact that the effectiveness of attention mechanisms largely varies with where they are positioned within the architecture of a model, rather than how they are designed internally. This work initiates a shifted mindset for the research community towards laying a foundation for more appropriately designed attention mechanisms and thus opening up their usage to disparate models for HPE at various integration levels.
@inproceedings{phu2026early,
title={Early-Stage Integration of Attention Mechanisms in OmniPose for Human Pose Estimation},
author={Phu, Khac-Anh and Hoang, Van-Dung and Le, Van-Tuong-Lan and Tran, Quang-Khai},
booktitle={Asian Conference on Intelligent Information and Database Systems},
pages={415--430},
year={2026},
organization={Springer, Singapore}
}