See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Findings, 2026

Authors: Le Thien Phuc Nguyen, Zhuoran Yu, Samuel Low Yu Hang, Subin An, Jeongik Lee, Yohan Ban, SeungEun Chung, Thanh-Huy Nguyen, JuWan Maeng, Soochahn Lee, Yong Jae Lee

TL;DR: A new benchmark for evaluating how well multimodal large language models understand human speech by jointly processing visual (lip movements, gestures) and auditory (speech audio) signals.


What this paper is about

Multimodal LLMs have made impressive strides in understanding images and text, but their ability to truly understand human speech – combining what they see (facial movements, body language) with what they hear (spoken words, tone) – has not been systematically evaluated.

Key idea

The paper introduces a comprehensive benchmark that tests audiovisual human speech understanding in multimodal LLMs. It evaluates whether these models can integrate visual and auditory cues to answer questions about who is speaking, what they are saying, and how they are communicating, going beyond simple speech recognition to test deeper comprehension.

Why it matters

This benchmark fills a critical gap in evaluating multimodal AI, pushing the field toward models that can understand human communication as holistically as people do.

Figure description