Skip to main navigation Skip to search Skip to main content

Trust the Eye or the Crowd? Unveiling Conformity in Multimodal Large Language Models

  • Haoran Luo
  • , Yuhan Niu
  • , Hengxian Liu
  • , Yuanfei Sun
  • , Zhihao Yang
  • , Tong Chen
  • , Dexin Liu
  • , Yiyi Miao
  • , Jingshi Zhou
  • , Haiyang Zhang
  • , Jionglong Su
  • , Huixin Zhong
  • , Yanan Liu
  • , Wei Wang
  • , Zimu Wang*
  • , Qi Chen*
  • *Corresponding author for this work
  • Xi'an Jiaotong-Liverpool University
  • School of Smart Governance
  • Renmin University of China
  • Shanghai University

Research output: Contribution to journalConference articlepeer-review

Abstract

Multimodal large language models (MLLMs) have been propelled from passive information processors to active participants in complex social interactions. In such contexts, their ability to preserve independent judgment amidst potentially conflicting social cues becomes crucial. While recent studies have scrutinized conformity in text-only LLMs, the phenomenon remains largely unexplored in multimodal systems. To address this gap, we introduce MM-BenchForm, the first benchmark designed to systematically evaluate conformity behaviors in MLLMs. MM-BenchForm assesses model resilience across multiple cognitive capabilities, ranging from simple visual perception to advanced logic and mathematical reasoning, by leveraging established multimodal datasets such as CLEVR and PlotQA. We further propose five evaluation protocols that simulate different forms of social influence. Using this framework, we conduct extensive experiments on state-of-the-art MLLMs, including GPT-4o mini, DeepSeek-VL2, and GLM-4.5V. Our results reveal pronounced conformity tendencies: models frequently abandon correct visual evidence and instead hallucinate answers that align with induced erroneous opinions. We further analyze the impact of interaction history in shaping conformity and perform a preliminary attention-based analysis to examine how models distribute attention across visual inputs and socially provided information.

Original languageEnglish
Pages (from-to)3797-3802
Number of pages6
JournalProceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD
Issue number2026
DOIs
Publication statusPublished - 2026
Event29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, China
Duration: 13 May 202615 May 2026

Keywords

  • Conformity
  • Multi-modal Reasoning
  • Multimodal Large Language Models
  • Multimodal Perception

Cite this