TY - GEN
T1 - BLAH
T2 - 22nd Pacific Rim International Conference on Artificial Intelligence, PRICAI 2025
AU - Chai, Enhui
AU - Cui, Tianxiang
AU - Lin, Ta
AU - Ye, Yujian
AU - Xue, Ning
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - The detection head framework critically influences the balance between classification and localization in small object detection, yet existing designs often neglect task-specific feature interactions, leading to optimization conflicts. To address this, we propose Bi-Level Attention Head (BLAH), a novel framework that harmonizes dual-task learning through structured attention mechanisms and adaptive loss optimization. BLAH introduces two key innovations: (1) Channel Group Self-Attention (CGSA) stacks, which dynamically recalibrate channel-group dependencies to align classification and localization features, resolving spatial-channel decoupling limitations in conventional attention. (2) Dual-Task Attention (DTA), integrating global channel attention for classification robustness (translation invariance) and local spatial attention for precise localization (translation variability), enabling synergistic task interaction without computational overhead. Further, we design a Differentiable Task-Balanced Loss (DTBL) that adaptively modulates gradients between tasks via cosine similarity constraints, ensuring stable optimization without extra parameters. Extensive experiments on MS COCO and VisDrone demonstrate BLAH’s superiority. When integrated with DETR, Deformable DETR, and YOLOv10, BLAH achieves +1.2% mAP on COCO over state-of-the-art detectors (e.g., YOLO-based, DETR-based) while maintaining inference efficiency, and significantly improves small-object detection (e.g., +4.5% APS on YOLOv12). Ablation studies validate each component’s necessity.
AB - The detection head framework critically influences the balance between classification and localization in small object detection, yet existing designs often neglect task-specific feature interactions, leading to optimization conflicts. To address this, we propose Bi-Level Attention Head (BLAH), a novel framework that harmonizes dual-task learning through structured attention mechanisms and adaptive loss optimization. BLAH introduces two key innovations: (1) Channel Group Self-Attention (CGSA) stacks, which dynamically recalibrate channel-group dependencies to align classification and localization features, resolving spatial-channel decoupling limitations in conventional attention. (2) Dual-Task Attention (DTA), integrating global channel attention for classification robustness (translation invariance) and local spatial attention for precise localization (translation variability), enabling synergistic task interaction without computational overhead. Further, we design a Differentiable Task-Balanced Loss (DTBL) that adaptively modulates gradients between tasks via cosine similarity constraints, ensuring stable optimization without extra parameters. Extensive experiments on MS COCO and VisDrone demonstrate BLAH’s superiority. When integrated with DETR, Deformable DETR, and YOLOv10, BLAH achieves +1.2% mAP on COCO over state-of-the-art detectors (e.g., YOLO-based, DETR-based) while maintaining inference efficiency, and significantly improves small-object detection (e.g., +4.5% APS on YOLOv12). Ablation studies validate each component’s necessity.
KW - Global channel-wise attention
KW - Local spatial-wise attention
KW - Object detection
KW - Objective imbalance
KW - Vision transformers
UR - https://www.scopus.com/pages/publications/105040353302
U2 - 10.1007/978-981-95-7084-3_17
DO - 10.1007/978-981-95-7084-3_17
M3 - Conference contribution
AN - SCOPUS:105040353302
SN - 9789819570836
T3 - Lecture Notes in Computer Science
SP - 264
EP - 280
BT - PRICAI 2025
A2 - Mei, Yi
A2 - Qian, Chao
A2 - Bai, Quan
A2 - Xue, Bing
A2 - Khanna, Sankalp
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 17 November 2025 through 21 November 2025
ER -