Abstract
Recognizing actions in RGB-D videos necessitates an in-depth understanding of spatial and temporal information from both RGB and depth modalities, as well as an efficient fusion of these data streams. Existing late fusion approaches often suffer from modality collapse due to high semantic similarity between modalities, where the fusion diminishes the distinct contributions of each modality and compromises overall effectiveness. In this context, we propose a novel End-to-end Cross-Modal Attention Transformer (E-CMAT) model. Our model processes RGB and depth inputs through two distinct expert encoders that utilize factorized spatio-temporal representations to capture dimension-independent features, enhancing motion understanding. The extracted RGB and depth tokens are subsequently fused via our innovative cross-modal attention, which is applied iteratively to ensure a robust integration that accentuates critical patterns and discrepancies between the modalities, thereby facilitating more precise action recognition. Our cross-modal attention utilizes one class token per expert for inter-expert information exchange, focusing attention on salient features and fostering synergistic predictions while reducing computational complexity from quadratic to linear. Furthermore, extensive experiments conducted on widely used benchmark datasets, such as NTU RGB-D 60, NTU RGB-D 120, and THU-READ, demonstrate the favorable performance of E-CMAT compared to state-of-the-art models.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
Free Keywords
- Action recognition in RGB-D videos
- Cross-modal attention
- Multimodal fusion
ASJC Scopus subject areas
- Media Technology
- Electrical and Electronic Engineering
Fingerprint
Dive into the research topics of 'Action Recognition in RGB-D Videos via an End-To-End Cross-Modal Attention Transformer'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver