Guides

Creator field notes

AI Rap Duo: Plan a Two-Person Performance Video

Build a readable two-person rap scene with clear positions, shared reference framing, a four-shot plan, and an honest audio-sync workflow.

Updated 2026-10-115 min read
Planning diagram03
AB
Performer ASeparate silhouettesPerformer B
Illustrative storyboard, not generated output.

An AI rap duo video works best when viewers can tell who is doing what. Establish performer A and performer B, keep their positions readable, and give each shot one main action. Two people gesturing in a frame is a performance visual; it does not establish that either person is delivering the words in your recording.

AI Genjutsu prepares silent visual clips. Exact alternating verses need an external audio and lip-sync workflow. Use this guide to prepare your reference and shot briefs, and check current studio availability before rendering.

Establish the pair before directing motion

Write a small continuity card: A stands on the left in a rust jacket; B stands on the right in a charcoal shirt. They are separated by one shoulder-width of empty space, with a plain wall behind them. Keep these descriptors consistent across shot briefs rather than adding new wardrobe details each time.

For an image-led shot, the single starting image should already show both people clearly. Avoid overlapping faces, extreme profile angles, crossed arms in front of another person, or a crowd behind them. Leave enough room around the frame for modest hand movements.

Two separate portrait uploads are not a supported joint-identity input here. You can photograph the pair together or prepare a permitted composition in an external image editor; check that the finished reference looks coherent before treating it as one starting image. A reference can guide appearance, but it does not lock identity throughout a clip.

Use four short shots with different jobs

Plan these as separate four-second clips and assemble a sixteen-second edit externally. Keep the same wall, light direction, outfits, and A/B positions.

ShotVisual briefEditing purpose
1: establishWaist-up two-shot; A and B hold position; a restrained push-in.Introduce the pair before the first cut.
2: A leadsA makes one open-hand gesture; B listens with hands low.Support A's section without competing movement.
3: B leadsB makes a small forward hand gesture; A remains still.Make the change of visual lead readable.
4: shared finishBoth give one subtle head nod; camera stays locked.Land on the final accent without a crossing move.

The edit can suggest a handover. It cannot create accurate mouth timing from silent, unsynchronized footage. For exact delivery, record the performers against your final audio or use a separately verified audio-driven lip-sync process, then check each speaker's segment in the editor.

An original shot prompt

Use this as a brief for shot 2, not as a verified output:

Four-second waist-up two-shot of two original performers beside a textured concrete wall. Performer A stays on the left wearing a rust jacket and makes one open-hand gesture near chest height. Performer B stays on the right wearing a charcoal shirt, listening with hands lowered. Both faces remain visible and separated. Locked camera, soft side lighting, restrained performance energy, no extra people, no position exchange, no readable text.

Change the leading performer for shot 3 while preserving the rest. Instructions such as “no position exchange” express a goal; inspect the result rather than assuming compliance.

Diagnose the frame, then simplify

ProblemNext change
Faces appear to swap

Use more distinct clothing and more space between the heads; reduce turning.

Hands or bodies intersectRemove simultaneous gestures and keep one performer still.
An extra person appears

Simplify the background and explicitly describe exactly two visible performers.

The visual lead is unclear

Cut to the next shot at the audio handover instead of directing both to move.

Mouths do not match the recording

Use a verified external lip-sync workflow or replace the section with non-mouth B-roll.

Questions before export

Can the duo rap my lyrics inside AI Genjutsu?

No. This site's output is silent, and it does not accept an audio track or provide exact lyric-driven mouth animation.

Does Seedance's audio capability change that?

ByteDance describes Seedance 1.5 Pro's joint audio-video capabilities. This site's configuration disables audio; model-level features are not an enabled site feature.

Can I use a famous artist's face or voice?

Build around original performers and material you have permission to use. A familiar name in a prompt does not establish permission or authenticity. AI Genjutsu is an independent site.

Continue with audio-first rap video planning, Migos AI source and workflow checks, or orange-stage scene planning.

Choose your next step