An AI rap duo video works best when viewers can tell who is doing what. Establish performer A and performer B, keep their positions readable, and give each shot one main action. Two people gesturing in a frame is a performance visual; it does not establish that either person is delivering the words in your recording.
AI Genjutsu prepares silent visual clips. Exact alternating verses need an external audio and lip-sync workflow. Use this guide to prepare your reference and shot briefs, and check current studio availability before rendering.
Establish the pair before directing motion
Write a small continuity card: A stands on the left in a rust jacket; B stands on the right in a charcoal shirt. They are separated by one shoulder-width of empty space, with a plain wall behind them. Keep these descriptors consistent across shot briefs rather than adding new wardrobe details each time.
For an image-led shot, the single starting image should already show both people clearly. Avoid overlapping faces, extreme profile angles, crossed arms in front of another person, or a crowd behind them. Leave enough room around the frame for modest hand movements.
Two separate portrait uploads are not a supported joint-identity input here. You can photograph the pair together or prepare a permitted composition in an external image editor; check that the finished reference looks coherent before treating it as one starting image. A reference can guide appearance, but it does not lock identity throughout a clip.
Use four short shots with different jobs
Plan these as separate four-second clips and assemble a sixteen-second edit externally. Keep the same wall, light direction, outfits, and A/B positions.
| Shot | Visual brief | Editing purpose |
|---|---|---|
| 1: establish | Waist-up two-shot; A and B hold position; a restrained push-in. | Introduce the pair before the first cut. |
| 2: A leads | A makes one open-hand gesture; B listens with hands low. | Support A's section without competing movement. |
| 3: B leads | B makes a small forward hand gesture; A remains still. | Make the change of visual lead readable. |
| 4: shared finish | Both give one subtle head nod; camera stays locked. | Land on the final accent without a crossing move. |
The edit can suggest a handover. It cannot create accurate mouth timing from silent, unsynchronized footage. For exact delivery, record the performers against your final audio or use a separately verified audio-driven lip-sync process, then check each speaker's segment in the editor.
An original shot prompt
Use this as a brief for shot 2, not as a verified output:
Four-second waist-up two-shot of two original performers beside a textured concrete wall. Performer A stays on the left wearing a rust jacket and makes one open-hand gesture near chest height. Performer B stays on the right wearing a charcoal shirt, listening with hands lowered. Both faces remain visible and separated. Locked camera, soft side lighting, restrained performance energy, no extra people, no position exchange, no readable text.
Change the leading performer for shot 3 while preserving the rest. Instructions such as “no position exchange” express a goal; inspect the result rather than assuming compliance.
Diagnose the frame, then simplify
| Problem | Next change |
|---|---|
| Faces appear to swap | Use more distinct clothing and more space between the heads; reduce turning. |
| Hands or bodies intersect | Remove simultaneous gestures and keep one performer still. |
| An extra person appears | Simplify the background and explicitly describe exactly two visible performers. |
| The visual lead is unclear | Cut to the next shot at the audio handover instead of directing both to move. |
| Mouths do not match the recording | Use a verified external lip-sync workflow or replace the section with non-mouth B-roll. |
Questions before export
Can the duo rap my lyrics inside AI Genjutsu?
No. This site's output is silent, and it does not accept an audio track or provide exact lyric-driven mouth animation.
Does Seedance's audio capability change that?
ByteDance describes Seedance 1.5 Pro's joint audio-video capabilities. This site's configuration disables audio; model-level features are not an enabled site feature.
Can I use a famous artist's face or voice?
Build around original performers and material you have permission to use. A familiar name in a prompt does not establish permission or authenticity. AI Genjutsu is an independent site.
Continue with audio-first rap video planning, Migos AI source and workflow checks, or orange-stage scene planning.