What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.