Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
Abstract
Tri-PvP benchmark reveals visual bias and asymmetric evidence-form preferences in omni-modal language models, with early-layer decodable modality effects resistant to surface mitigation.
Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a photograph or recording of a dog) and propositional signals (e.g., the declarative claim "this is a dog"), such that any measured modality bias is inherently confounded with evidence-form bias, precluding clean attribution to either source. To address this, we introduce Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, where vision and audio each take perceptual or propositional form. Evaluating five OLLMs, we find robust visual bias across most models and evidence-type conditions. Crucially, we reveal a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio. Further analyses via layer-wise linear probing and contrastive decoding reveal that modality bias is already linearly decodable from early representation layers and can only be partially mitigated, calling for mitigation strategies beyond surface-level interventions.
Community
Tri-PvP is an 8000-sample tri-modal (image/audio/text) conflict benchmark that controls perceptual vs. propositional evidence in vision and audio. Across five OLLMs, we found that image bias dominates, and models favor perceptual images but propositional (spoken) audio.
Code: https://github.com/MiuLab/Tri-PvP
Dataset: https://huggingface.co/Tri-PvP
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models (2026)
- Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict (2026)
- Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models (2026)
- Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict (2026)
- CAD: Conflict-Aware Decoding to Mitigate Cross-Modal Hallucinations in Omnimodal Large Language Models (2026)
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models (2026)
- Which Source Wins? Task-Dependent Reliance in Vision-Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.06011 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper