Abstract
Multimodal large language models (MLLMs) have been propelled from passive information processors to active participants in complex social interactions. In such contexts, their ability to preserve independent judgment amidst potentially conflicting social cues becomes crucial. While recent studies have scrutinized conformity in text-only LLMs, the phenomenon remains largely unexplored in multimodal systems. To address this gap, we introduce MM-BenchForm, the first benchmark designed to systematically evaluate conformity behaviors in MLLMs. MM-BenchForm assesses model resilience across multiple cognitive capabilities, ranging from simple visual perception to advanced logic and mathematical reasoning, by leveraging established multimodal datasets such as CLEVR and PlotQA. We further propose five evaluation protocols that simulate different forms of social influence. Using this framework, we conduct extensive experiments on state-of-the-art MLLMs, including GPT-4o mini, DeepSeek-VL2, and GLM-4.5V. Our results reveal pronounced conformity tendencies: models frequently abandon correct visual evidence and instead hallucinate answers that align with induced erroneous opinions. We further analyze the impact of interaction history in shaping conformity and perform a preliminary attention-based analysis to examine how models distribute attention across visual inputs and socially provided information.
| Original language | English |
|---|---|
| Pages (from-to) | 3797-3802 |
| Number of pages | 6 |
| Journal | Proceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD |
| Issue number | 2026 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | 29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, China Duration: 13 May 2026 → 15 May 2026 |
Keywords
- Conformity
- Multi-modal Reasoning
- Multimodal Large Language Models
- Multimodal Perception
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver