Urgent.News

What's breaking now, across thousands of outlets.

AI

Vision-in-the-Loop: When the AI Rewrites Its Own Prompts from the Generated Frame

On the AI video ad platform I work on, every scene goes through the same painful loop: write a prompt, send it to an AI video model provider, wait two minutes, open the result, squint at the frame, and decide what went wrong. Camera too wide. Product missing from the hero shot. Color palette drifted warm when the brand brief says cool neutrals. Avatar looks like a different person than scene…

The AI video ad platform I work on experiences a tedious manual loop when creating scenes. First, an operator writes a prompt, then sends it to an AI video model provider to generate a scene. After waiting two minutes to view the result, they compare the generated frame with the reference ad. If any issues arise, such as a product missing from the hero shot or color palette drifting, the operator must rewrite the prompt to address the problem.

This process, which repeats across twelve scenes and three iterations, consumes an hour of human attention that could be better spent on brand strategy.

The new vision-in-the-loop approach aims to eliminate this inefficiency. After the first frame of each scene is generated, a per-scene feedback loop runs immediately. The vision model receives the generated frame alongside scene metadata like shot type, product placement rules, and color palette constraints. It then returns a structured critique, identifying issues like framing too wide, product absent, palette drift, or identity mismatch.

The rewrite logic targets these issues, adjusting the visual prompt with minimal changes to prevent losing the brand voice or swapping the protagonist.

In production testing, the vision critique caught issues that text-only plan QA missed, such as framing drift, product absence, palette mismatch, and motion mismatch. The rewrite step is constrained, ensuring only the flagged visual prompt fields are adjusted. For person scenes, vision-in-the-loop sets motion-transfer as the default generation mode, ensuring more lifelike results.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Tuesday 25 August →