Published August 2026 | Version v1
Thesis Embargoed

Affect-Biased Attention in Vision–Language Models

Creators

  • 1. University of Chicago

Contributors

  • 1. University of Chicago

Description

Vision–language models are often expected to form comprehensive and objective representations of visual scenes, yet recent work suggests that models such as CLIP encode some content more strongly than others. Here, we ask whether such selective attention is shaped by the affective properties of visual content. In human attention, emotionally arousing stimuli tend to capture attention more readily than neutral stimuli. We tested whether a similar arousal bias appears in CLIP visual encoders. We presented six ViT-based CLIP models with composite stimuli containing pairs of affective images and quantified representational priority using two complementary metrics. Image-level regressions showed that arousal consistently predicted model representational priority across models, metrics, and image sets, even after controlling for low-level visual features. Valence effects were weaker and less consistent. These results suggest that large-scale multimodal training may transmit not only visual-semantic correspondence, but also human-like affective biases in what visual content receives priority.

Files

Embargoed

The files will be made publicly available on August 25, 2028.

Reason: Journal Submission

Additional details

UChicago Information

Division(s)
Social Sciences Division
Department(s)
Computational Social Sciences (MACSS)