Công trìnhPublications A Deformable Hybrid Transformer with Multi-Axis St…
JournalSCIE SCIE Q1 IF 4.2

A Deformable Hybrid Transformer with Multi-Axis Strip Attention for Geometric-Aware Medical Image Diagnosis

IEEE Access
2026 vol. 2026 IEEE Press

Accurate diagnosis in medical imaging is often hampered by two intrinsic factors: the complex, anomalous geometric distortion of anatomical structures and the extreme size variability of pathological lesions. Existing convolutional neural networks (CNNs) struggle with global context, while Vision Transformer (ViT) models are limited by quadratic computational costs or reliance on window-based attention mechanisms that disrupt semantic continuity. To address these limitations, we propose MaxStripViT, a novel hybrid architecture that effectively integrates distortion modeling with multi-axis strip attention mechanisms and Local Position-Aware blocks. Our method introduces three key contributions. (1) A Geometric-Adaptive Stem (GAS) leverages learnable offsets via Deformable Convolutions (DCNv2) to dynamically align the sampling grid with irregular organ boundaries at the earliest feature extraction stage. It effectively mitigates background noise. (2) A Position-Aware Local Block is proposed to enhance the Mobile Inverted Bottleneck (MBConv) with Coordinate Attention. This mechanism explicitly models long-range dependencies along spatial axes, improving the precise localization of subtle lesions. (3) A novel sequential multi-axis strip attention mechanism is proposed to replace shift window-based self-attention. The proposed method is experimented on three dataset benchmarks of CT and ChestX-ray imaging for multi-class lung disease such as CheXtImageNet, IQ-OTH/NCCD, and ChestXray-ImageDataset. The results demonstrated that MaxStripViT achieves robust performance outperforms to standard ViT models and improves performance compared to the most advanced hybrid models currently available, including Swin Transformer, MaxViT, CSWin, ConvNeXt, ConvNeXtv2, and EfficientNet-B7 in terms of classification accuracy and computational efficiency. The proposed method provides a robust solution for health scenarios requiring both geometric flexibility and multi-scale interpretability.