I set filmmaker traps for AI "directors." The models fell for one. My rubric fell for three.
I set filmmaker traps for AI "directors." The models fell for one. My rubric fell for three. I QC frames for photorealism and anatomy before anything becomes video. I regenerate bad frames; I do not argue with them. I also direct: corporate events, mythology films, horror shorts, stock shoots. If AI models are going to help on a real set, "usually right" is not good enough. The bar is three…
The author created a set of adversarial tasks to evaluate AI video directors called DirectorBench. The tasks included writing prompts, planning shots, and judging frames for visual consistency. The author tested two AI models, Anthropic Opus 5.5 and Grok 4.6, on these tasks. Opus performed well on all tasks, while Grok struggled with certain constraints, demonstrating potential bias towards keyword matching over contextual understanding.
The author also found that their scoring rubric unintentionally penalized natural phrasing and directness, impacting the model's performance. The pilot aimed to develop a benchmark that could distinguish between models that break rules versus those that explain their reasoning, as well as differentiate between bad answers and well-expressed good answers.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.