Orthogonalization-Guided Feature Fusion Network for Multimodal 2D+3D Facial Expression Recognition

Shisong Lin, Mengchao Bai, Feng Liu, Linlin Shen, Yicong Zhou

Research output: Journal PublicationArticlepeer-review

15 Citations (Scopus)


As 2D and 3D data present different views of the same face, the features extracted from them can be both complementary and redundant. In this paper, we present a novel and efficient orthogonalization-guided feature fusion network, namely OGF^2Net, to fuse the features extracted from 2D and 3D faces for facial expression recognition. While 2D texture maps are fed into a 2D feature extraction pipeline (FE2DNet), the attribute maps generated from 3D data are concatenated as input of the 3D feature extraction pipeline (FE3DNet). The two networks are separately trained at the first stage and frozen in the second stage for late feature fusion, which can well address the unavailability of a large number of 3D+2D face pairs. To reduce the redundancies among features extracted from 2D and 3D streams, we design an orthogonal loss-guided feature fusion network to orthogonalize the features before fusing them. Experimental results show that the proposed method significantly outperforms the state-of-the-art algorithms on both the BU-3DFE and Bosphorus databases. While accuracies as high as 89.05% (P1 protocol) and 89.07% (P2 protocol) are achieved on the BU-3DFE database, an accuracy of 89.28% is achieved on the Bosphorus database. The complexity analysis also suggests that our approach achieves a higher processing speed while simultaneously requiring lower memory costs.

Original languageEnglish
Article number9115253
Pages (from-to)1581-1591
Number of pages11
JournalIEEE Transactions on Multimedia
Publication statusPublished - 2021
Externally publishedYes


  • Multimodal facial expression recognition
  • feature fusion

ASJC Scopus subject areas

  • Signal Processing
  • Media Technology
  • Computer Science Applications
  • Electrical and Electronic Engineering


Dive into the research topics of 'Orthogonalization-Guided Feature Fusion Network for Multimodal 2D+3D Facial Expression Recognition'. Together they form a unique fingerprint.

Cite this