Analysis of Multimodal LLMs

We study the capabilities and limitations of multimodal large language models (MLLMs) across audiovisual, linguistic, and perceptual reasoning tasks. Our work develops benchmarks and diagnostic tools that reveal how well these models integrate signals across modalities, how robust they are in the wild, and whether their reported performance reflects genuine learning or contamination from training data.

Research directions

  • Benchmarking audiovisual understanding. Designing benchmarks that test whether MLLMs can jointly exploit visual cues (lips, gestures, scene context) and auditory speech to answer fine-grained questions about human communication.
  • Universal active speaker detection. Building models that reliably identify who is speaking in cluttered real-world video, handling occlusion, noise, and multi-speaker scenes through multimodal fusion.
  • Contamination and integrity. Detecting benchmark contamination in vision–language models by perturbing inputs in semantically meaningful ways and measuring how performance degrades relative to genuinely learned capabilities.