DiaVLo: Diagnosing Behaviours of Vision-Language Models
AI Digest - ArXiv AI
DiaVLo: Diagnosing Behaviours of Vision-Language Models
Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignmen
Source: ArXiv AI