Urgent.News

What's breaking now, across thousands of outlets.

AI

UniEvo-VL: On-Policy Self-Distillation for Multimodal Image Generation

UniEvo-VL trains a unified multimodal model to internalize its own image critiques by matching a critique-conditioned EMA teacher along denoising trajectories sampled from the student’s current policy. From inference-time reflection to learned behavior A multimodal model that can both generate and understand images has a useful feedback loop available: generate an image, compare it with the…

UniEvo-VL is a novel approach that trains a unified multimodal model to improve its own image generation skills through on-policy self-distillation. This innovative method allows the model to learn from its own critiques by constructing a teacher prompt based on the critiques, without requiring additional fine-tuning.

The training process of UniEvo-VL involves two main steps: generating and assessing. First, the model generates an image from a given prompt and then compares the generated image with the prompt. If the image is rejected due to discrepancies, the model constructs a revised prompt by combining the original prompt with the discrepancy critique. The revised prompt is then used as privileged information for the teacher role during distillation.

During distillation, the student model generates a denoising trajectory conditioned only on the original prompt. At various states sampled from this trajectory, the student's denoising distribution is compared with a teacher distribution conditioned on the revised prompt. This design ensures that the student learns to reproduce the teacher's denoising behavior while receiving the corrective feedback during training, rather than relying on it during inference.

UniEvo-VL maintains separate student parameters and an exponential-moving-average teacher, with the EMA decay set to 0.999. The teacher receives the revised prompt containing the critique, making the correction a privileged piece of information for the student to learn from. By using on-policy matching, the teacher is evaluated at the states the student actually visits, including challenging or difficult states generated by the student's current policy.

The paper reports significant improvements in the GenEval benchmark, with scores increasing from 0.747 to 0.808. Additionally, there are improvements in the GenEval2 Soft-TIFA benchmark, with scores rising from 32.97 to 35.53. However, the paper also highlights the importance of post-revision verification and the effectiveness of the external critic, as shown by the stronger performance when using GPT5.6-Luna as the external critic.

The results demonstrate that trajectory-level distillation provides a more comprehensive learning signal compared to training on revised final images alone.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →