All systems operational
Q1 2026

VSMatch-Lip: a visual-semantic matching framework for zero-shot lip reading

Jixia Shen · Hao Yuan · Xingyu Zhang · Yakun Zhang · Changyan Zheng · Liang Xie · Zhibin Song · Erwei Yin
10.1007/s44443-026-00481-4 390 Views 0 Citations
0
Citations
390
Views
Abstract

Abstract
Lip reading interprets speech from visual lip movements, offering a vital complement to audio-based recognition in challenging acoustic environments. However, existing models rely on supervised, closed-set classification and fail to recognize out-of-vocabulary words, severely limiting their practical application. To address this zero-shot learning (ZSL) challenge, we propose VSMatch-Lip, a non-generative visual-semantic matching framework. Our approach is grounded in the insight that while lip movements represent a visual manifestation of phonetics, providing a strong physical correlation for generalization, relying solely on this correlation is insufficient due to ambiguities like homophones. Therefore, our core innovation lies in introducing a multi-source fused semantic representation that synergistically integrates lexical meaning with powerful phonetic cues. This design allows the phonetic component to ground the alignment in visual articulation, while the semantic component provides crucial disambiguation, creating a more robust and discriminative target for matching. To effectively optimize this matching process, we design a tailored contrastive learning framework with specialized optimization strategies to tackle the large intra-class variance and training instability. As a key contribution, we also establish the first comprehensive ZSL benchmark on large-scale, in-the-wild datasets. Extensive experiments on this benchmark demonstrate that VSMatch-Lip achieves state-of-the-art performance, consistently outperforming all baselines, including contemporary generative models. Notably, under a 19:1 seen-to-unseen ratio on LRW, it surpasses the strongest generative baseline by nearly 9% in Top-1 unseen accuracy. To the best of our knowledge, this is the first successful and rigorous validation of a non-generative, direct matching ZSL framework on large-scale, in-the-wild lip reading benchmarks.

Cite this Article (APA)
Jixia, S., Hao, Y., Xingyu, Z., Yakun, Z., Changyan, Z., Liang, X., Zhibin, S., Erwei, Y. (2026). VSMatch-Lip: a visual-semantic matching framework for zero-shot lip reading. Journal of King Saud University - Computer and Information Sciences. https://doi.org/10.1007/s44443-026-00481-4
Related Papers
A lightweight model for indoor object detection in unstructured scenes based on joint attention and …
Zhizhong Xing; Leping Li; Ying Yang; Wei Zhou; Guolan Ma; Shaochun Chen; Lechun · 2026
13
cites
424
DDM-YOLO: A lightweight oriented detection model for mature daylily fruits in complex environments
Minqiu Kuang; Xuejie Zou; Fangping Xie; Xiaojian Li; Shang Chen; Dawei Liu; Yuxu · 2026
8
cites
430
Information guided Levy flight for robot search in unknown environments
Weitao Zhao; Zati Hakim Azizul; Xin Lyu; Weijie Kuang · 2026
4
cites
411
3
cites
504
Bridging the gap: A comprehensive survey on AI-driven digital twin networks for future wireless syst…
Yousef Sanjalawe; Salam Fraihat; Salam Al-E’mari; Sharif Naser Makhadmeh · 2026
3
cites
417
Access
View Full Text via DOI
Published in
ISSN 1319-1578
Quartile Q1
AMS Score 100
Field Computer Science & AI
Publisher Elsevier / King Saud University
Country 🇸🇦 Saudi Arabia
View Journal Profile →
Authors
Publication Details
Year 2026
Language English
Added 06 Jul 2026