Affect-Biased Attention in Vision–Language Models
Description
Vision–language models are often expected to form comprehensive and objective representations of visual scenes, yet recent work suggests that models such as CLIP encode some content more strongly than others. Here, we ask whether such selective attention is shaped by the affective properties of visual content. In human attention, emotionally arousing stimuli tend to capture attention more readily than neutral stimuli. We tested whether a similar arousal bias appears in CLIP visual encoders. We presented six ViT-based CLIP models with composite stimuli containing pairs of affective images and quantified representational priority using two complementary metrics. Image-level regressions showed that arousal consistently predicted model representational priority across models, metrics, and image sets, even after controlling for low-level visual features. Valence effects were weaker and less consistent. These results suggest that large-scale multimodal training may transmit not only visual-semantic correspondence, but also human-like affective biases in what visual content receives priority.