Building a Focused AI Video Workflow for One or Two Portraits
When an AI video tool is built around a known reference scene, the input form is part of the product. Users should not have to write a prompt that describes a choreography the system already knows. They should only need to answer one question: which person should appear in each role? I built Rumpelstiltskin AI Video around that idea. It creates a personalized version of the Rumpelstiltskin tiptoe…
Creating a concise AI video workflow focused on a single or dual portrait involves a few key steps. First, the user interface presents two distinct modes: "Dancer Only" for a single portrait replacing the lead dancer, and "Dancer & Maiden" for two portraits replacing both characters. This explicit character mapping contrasts with a generic prompt box. The backend maintains a simple input model, where photoMode is either one or two, and a second file is mandatory only in the two-photo scenario.
The API supports JPG, PNG, and WebP files up to 20 MB each. It validates the reference video duration, desired resolution, and aspect ratio before processing. To ensure reliability, each generation request is made idempotent. The system assigns an idempotency key to each attempt, which the server verifies before reserving credits. This prevents accidental double charges from repeated clicks.
The video generation process follows a staged workflow. The server first uploads the portrait files and the silent reference video, then conducts safety checks before sending the task to the video provider. The audio processing occurs separately, merging the original reference audio into the generated video. This separation simplifies failure tracking, as each stage has its own status and error handling.
Credit management is treated as a reservation system. A generation costs 100 credits at 480P resolution or 170 credits at 720P resolution. The server checks the user's balance and credits before deducting the amount and before initiating the provider request. If any stage fails, the failed generation path refunds the used credits rather than issuing a cash refund. The user interface clearly communicates this credit return process, aligning billing behavior with database actions.
Finally, the product's promise remains narrow and straightforward. It focuses solely on transforming a portrait into a character within the original tiptoe dance, offering users a one-time credit pack, a choice of resolution (480P or 720P), and an MP4 download with the original reference audio upon successful processing. This clear, single-purpose approach influences the entire implementation, resulting in a short form, strict validation, visible request states, and straightforward pricing that avoids the illusion of a free or unlimited service.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.