All systems operational
Q1 2026

FER-EMFormer: An enhanced mamba-transformer network for facial expression recognition

Weijun Gong · Xusheng Du · Jiaxin Wu · Zhichao Peng
10.1007/s44443-026-00740-4 396 Views 0 Citations
0
Citations
396
Views
Abstract

Abstract

Facial expression recognition (FER) has a variety of applications in advanced intelligent fields such as human–computer interaction, cognitive psychology, and intelligent driving. However, FER in wild scenarios faces multiple challenges, including occlusion, pose variations, and subtle differences, which make current models unable to address these issues effectively. To tackle these challenges, we propose an efficient and robust Enhanced Mamba-Transformer architecture for FER (FER-EMFormer) in complex scenes. The FER-EMFormer primarily consists of two core modules: the Hybrid Enhanced Mamba-Transformer (HEMT) and the Enhanced Cascaded Mamba-Transformer (ECMT). HEMT effectively combines Mamba and Transformer to capture informative global context and spatial dependency features, while integrating detail feature frequency enhancement across multiple views to enable collaborative global–local feature understanding. ECMT uses a cascaded architecture to fuse the optimized global dependencies obtained from Mamba-Transformer, then employs a Transformer with a multi-dimensional aggregation feedforward network to precisely control the network's information flow, yielding high-density discriminative information and further improving the accuracy of facial expression recognition. Extensive experiments show that our FER-EMFormer significantly outperforms current FER models and achieves state-of-the-art performance of 96.06% on RAF-DB, 95.78% on FERPlus, 72.91% on AffectNet-7, and 70.43% on AffectNet-8, while simultaneously demonstrating excellent robustness and generalization capabilities on occlusion and pose variation expression datasets as well as cross-dataset. The code is available at
https://github.com/ferlab08/FER-EMFormer

Cite this Article (APA)
Weijun, G., Xusheng, D., Jiaxin, W., Zhichao, P. (2026). FER-EMFormer: An enhanced mamba-transformer network for facial expression recognition. Journal of King Saud University - Computer and Information Sciences. https://doi.org/10.1007/s44443-026-00740-4
Related Papers
A lightweight model for indoor object detection in unstructured scenes based on joint attention and …
Zhizhong Xing; Leping Li; Ying Yang; Wei Zhou; Guolan Ma; Shaochun Chen; Lechun · 2026
13
cites
424
DDM-YOLO: A lightweight oriented detection model for mature daylily fruits in complex environments
Minqiu Kuang; Xuejie Zou; Fangping Xie; Xiaojian Li; Shang Chen; Dawei Liu; Yuxu · 2026
8
cites
430
Information guided Levy flight for robot search in unknown environments
Weitao Zhao; Zati Hakim Azizul; Xin Lyu; Weijie Kuang · 2026
4
cites
411
3
cites
504
Bridging the gap: A comprehensive survey on AI-driven digital twin networks for future wireless syst…
Yousef Sanjalawe; Salam Fraihat; Salam Al-E’mari; Sharif Naser Makhadmeh · 2026
3
cites
417
Access
View Full Text via DOI
Published in
ISSN 1319-1578
Quartile Q1
AMS Score 100
Field Computer Science & AI
Publisher Elsevier / King Saud University
Country 🇸🇦 Saudi Arabia
View Journal Profile →
Authors
Publication Details
Year 2026
Language English
Added 06 Jul 2026