How We Built AI Rap Duo: Two Photos In, Lip-Synced Rap Video Out
You've probably seen the trend: two people in an orange booth, a mic hanging between them, both lip-syncing a rap verse in perfect sync — sometimes it's two friends, sometimes it's two cats. We built a generator for exactly that format. It's called AI Rap Duo : you upload two photos, type a one-line topic, and get back a vertical rap duo video where both faces perform an original AI-written verse…
AI Rap Duo is an online AI rap duo video generator that allows users to upload two photos, type a one-line topic, and generate a vertical video of the two performers lip-syncing an original AI-written rap verse and beat. This post provides a technical walkthrough of the pipeline, discussing challenges and solutions encountered during development.
Key aspects include:
1. Identity preservation across two faces: Lip-syncing two faces in one frame is more challenging than single-face animation. To address this, the model uses strong reference-image conditioning, ensuring both identities survive the entire duration of the video.
2. Lip-sync on generated audio: Unlike traditional photo-to-video tools that animate first and then add audio, AI Rap Duo generates video and audio simultaneously. This joint generation process results in more natural lip-syncing and vocal tracks.
3. Crafting the verse: The pipeline starts with words, not video. Users input a single-line topic, and the lyric generator creates an original rap verse with explicit structure and voice assignments for each performer. This approach ensures the verse is relevant to the topic and suitable for a rap duo performance.
The pipeline consists of several phases:
- Upload and consent: Users provide the two photos and confirm they have permission to use the likenesses. This non-negotiable step is crucial for building trust and complying with privacy regulations.
- Topic → lyrics: Users input a topic, and the lyric generator creates an original rap verse tailored to the topic and the two performers. This step guarantees that the lyrics are relevant, engaging, and suitable for a rap performance.
- Reference conditioning: The two uploaded photos are fed to the video model as identity references. High-quality photos with visible eyes and good lighting help the model preserve the identities of the performers. Normalizing the references ensures consistent lighting and color across the generated video.
By focusing on these key aspects, AI Rap Duo creates unique, fully generated rap duo videos without the need for filming, editing, or borrowed audio. This innovative approach allows users to create engaging, copyright-safe content for platforms like TikTok, Reels, and Shorts.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.