DiaVLo: Diagnosing Behaviours of Vision-Language Models
Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.