Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.