PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference
Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two fundamental limitations. First, most approaches operate exclusively post-vision encoder, leaving the substantial latency of the visual encoding phase unoptimized. Second, under…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.