Skip to main navigation Skip to search Skip to main content

What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in comprehending complex visual content. However, the mechanisms underlying how VLMs process visual information remain largely unexplored. In this paper, we conduct a thorough empirical analysis, focusing on the attention modules across layers, by which we reveal several key insights about how these models process visual data: (i) the internal representation of the query tokens (e.g., representations of "describe the image"), is utilized by the model to store global image information; we demonstrate that the model generates surprisingly descriptive responses solely from these tokens, without direct access to image tokens. (ii) Cross-modal information flow is predominantly influenced by the middle layers (approximately 25% of all layers), while early and late layers contribute only marginally. (iii) Fine-grained visual attributes and object details are directly extracted from image tokens in a spatially localized manner, i.e., the generated tokens associated with specific object or attribute attend strongly to their corresponding regions in the image. We propose novel quantitative evaluation to validate our observations, leveraging real-world complex visual scenes. Finally, we demonstrate the potential of our findings in facilitating efficient visual processing in state-of-the-art VLMs.
Original languageEnglish
Title of host publication2025 IEEE/CVF Conference On Computer Vision And Pattern Recognition (CVPR)
PublisherIEEE Computer Society
Pages14549-14558
Number of pages10
ISBN (Electronic)979-8-3315-4364-8
ISBN (Print)979-8-3315-4365-5
DOIs
StatePublished - Aug 2025
Event2025 Conference on Computer Vision and Pattern Recognition-CVPR-Annual - Nashville, Tunisia
Duration: 10 Jun 202517 Jun 2025

Publication series

NameIeee Conference On Computer Vision And Pattern Recognition

Conference

Conference2025 Conference on Computer Vision and Pattern Recognition-CVPR-Annual
Country/TerritoryTunisia
CityNashville
Period10/06/2517/06/25

Fingerprint

Dive into the research topics of 'What's in the Image? A Deep-Dive into the Vision of Vision Language Models'. Together they form a unique fingerprint.

Cite this