Skip to main navigation Skip to search Skip to main content

Action Recognition in RGB-D Videos via an End-To-End Cross-Modal Attention Transformer

  • Yujun Ma
  • , Benjia Zhou
  • , Hong Zhang
  • , Xizheng Zhang
  • , Xiangjian He
  • , Ruili Wang*
  • *Corresponding author for this work

Research output: Journal PublicationArticlepeer-review

Abstract

Recognizing actions in RGB-D videos necessitates an in-depth understanding of spatial and temporal information from both RGB and depth modalities, as well as an efficient fusion of these data streams. Existing late fusion approaches often suffer from modality collapse due to high semantic similarity between modalities, where the fusion diminishes the distinct contributions of each modality and compromises overall effectiveness. In this context, we propose a novel End-to-end Cross-Modal Attention Transformer (E-CMAT) model. Our model processes RGB and depth inputs through two distinct expert encoders that utilize factorized spatio-temporal representations to capture dimension-independent features, enhancing motion understanding. The extracted RGB and depth tokens are subsequently fused via our innovative cross-modal attention, which is applied iteratively to ensure a robust integration that accentuates critical patterns and discrepancies between the modalities, thereby facilitating more precise action recognition. Our cross-modal attention utilizes one class token per expert for inter-expert information exchange, focusing attention on salient features and fostering synergistic predictions while reducing computational complexity from quadratic to linear. Furthermore, extensive experiments conducted on widely used benchmark datasets, such as NTU RGB-D 60, NTU RGB-D 120, and THU-READ, demonstrate the favorable performance of E-CMAT compared to state-of-the-art models.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
Publication statusAccepted/In press - 2026

Free Keywords

  • Action recognition in RGB-D videos
  • Cross-modal attention
  • Multimodal fusion

ASJC Scopus subject areas

  • Media Technology
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Action Recognition in RGB-D Videos via an End-To-End Cross-Modal Attention Transformer'. Together they form a unique fingerprint.

Cite this