How to Make a Photo Sing (and Look Real): What Six Test Renders Taught Us
Three things decide whether a singing-photo video looks real, in this order. First, what drives the mouth: feed the renderer the isolated voice, not the full song. In our test that alone moved the mouth-to-voice sync score from 0.04 to 0.39 on the same engine and the same photo. Second, direction: without a one-line instruction the singer walked towards the camera and the framing drifted; with it, the person stayed put and sang. Third, the photo: a head-and-shoulders crop gave the mouth 151 pixels of width against 88 for the uncropped shot of the same person, so the articulation is simply easier to see. The engine mattered less than all three.
We were not happy with our first music videos: the person in the photo wandered about, the mouth seemed to follow the band rather than the singer, and the sound was mono. Instead of guessing, we set up a small bench test. One photo, one chorus, the same ten seconds, rendered six different ways, each one measured the same way. This is the write-up, with the clips, the numbers, and what we changed on Tunely the same day.
Want it now? Make a music video from a photo — free, in seconds.
The test: one photo, one chorus, six renders
The photo is an AI-generated portrait that belongs to us, so nobody's real face sits on a test bench. The song is a birthday song made on Tunely, and every render uses the same ten seconds, starting one second before the chorus.
For each render we measured four things. The sync score is the correlation between how far the mouth is open in each frame and how loud the isolated voice is at that instant: 1.0 would be a mouth that opens exactly with the voice, 0 is a mouth moving at random. The mouth lag is the time shift at which that correlation peaks. The mouth width is how many pixels the mouth gets in the finished frame. And we noted what the file itself was: size, frame rate, mono or stereo.
Two clips to start with. On the left, the way our first version worked. On the right, the way Tunely makes them now.
About the sync score: it needs a clip with pauses in it, because it works by comparing loud moments with quiet ones. On a passage where the singer never stops, every render scores low. Compare scores only between renders of the same clip, which is what every table below does. We have no human baseline in this test, so read the numbers against each other.
Finding 1: what drives the mouth matters more than the engine
A finished song is a voice plus drums, bass, guitars and everything else. Hand all of that to a renderer and the mouth tries to follow all of it. So we separated the voice from the song first and let the bare voice drive the mouth, then put the full stereo song back under the picture at the end.
Same engine, same photo, same ten seconds. The only difference is what the renderer listened to.
| What drove the mouth | Sync score | Mouth opening, loud moments | Mouth opening, quiet moments |
|---|---|---|---|
| What drove the mouthThe full song | Sync score0.04 | Mouth opening, loud moments0.115 | Mouth opening, quiet moments0.107 |
| What drove the mouthThe isolated voice | Sync score0.39 | Mouth opening, loud moments0.213 | Mouth opening, quiet moments0.102 |
Look at the last two columns. Driven by the full song, the mouth was open about as much when the singer was silent as when she was belting: it was keeping time with the band. Driven by the voice, it opened twice as wide on the loud syllables as in the gaps. This was the single biggest improvement in the whole test, and it has nothing to do with which engine you use.
Finding 2: give the singer a direction, or the singer improvises
Our first version sent a photo and a sound file and nothing else. Left to itself, the renderer invents a performance: in the "before" clip the woman walks towards the camera, the framing tightens, a hand goes to the chest. It is lively. It is also not the photo you uploaded any more, and over twenty seconds the face can drift away from the person you know.
One sentence fixes it. Every render now carries the same direction: sing to the camera, stay in place, natural head and shoulder movement, the camera does not move. Here is the same uncropped photo rendered the old way and the way we do it now. The direction is what keeps her in place: Engine A further down is the same engine as the "before" clip, and once it had the direction it stayed put too.
Timing improved as well. In the "before" render the mouth ran about 200 milliseconds ahead of the voice, which is the uncanny feeling of a dubbed film. The current pipeline peaks within one frame of the voice.
Finding 3: the photo's framing decides how much singing you can see
We rendered the same person twice on the same engine: once from the uncropped three-quarter photo, once from a head-and-shoulders crop of that very photo.
| Photo | Mouth width in the finished video | Sync score |
|---|---|---|
| PhotoUncropped, three-quarter length | Mouth width in the finished video88 px | Sync score0.41 |
| PhotoHead-and-shoulders crop of the same photo | Mouth width in the finished video151 px | Sync score0.39 |
| PhotoFor reference: the "before" clip | Mouth width in the finished video68 px | Sync score0.26 |
The sync is equally good in both. What changes is how much of it you can see: the crop gives the mouth 1.7 times the width, so consonants and held vowels read clearly on a phone screen. If the photo you love is a full-body shot, crop it to the chest before you upload it. You lose the shoes and gain the performance.
The engines, given identical inputs
With the inputs fixed (isolated voice, the same direction, the head-and-shoulders crop) we ran four engines. We letter them rather than name them: engines change every few months, and the three findings above held on all of them.
| Engine | Output | Sound as delivered | Sync score | Mouth lag | Render time, 10 s clip |
|---|---|---|---|---|---|
| EngineA, the one we used before | Output832×1120, 25 fps | Sound as deliveredmono | Sync score0.25 | Mouth lag160 ms early | Render time, 10 s clip4 min 1 s |
| EngineB, the one Tunely uses now | Output1168×1760, 30 fps | Sound as deliveredstereo after our remix step | Sync score0.39 | Mouth lagwithin one frame | Render time, 10 s clip4 min 36 s to 5 min 48 s |
| EngineC, a lighter tier of B | Output784×1168, 30 fps | Sound as deliveredstereo | Sync score0.26 | Mouth lagwithin one frame | Render time, 10 s clip3 min 33 s |
| EngineD | Outputno output | Sound as deliveredn/a | Sync scoren/a | Mouth lagn/a | Render time, 10 s clipfailed twice on this photo |
Engine D has rendered other photos for us without trouble. It refused this one twice with an internal error, which is a data point of its own: a video you cannot get is worse than a slower one.
What we changed on Tunely the same day
Everything above went into the product on 17 September 2026. When you make a music video on Tunely now:
- The voice is separated from your song and drives the mouth. Your full stereo song goes back under the picture afterwards.
- Every render carries the same direction: stay in place, sing to the camera, the camera does not move.
- The 20-second clip opens on the chorus instead of the intro. A person standing silently through four bars of guitar is not a music video.
- The lyrics are burned in as captions, timed to the song, in the right script for the language.
- The finished file is up to 1168×1760 at 30 frames per second (its shape follows your photo), with no watermark.
- It takes about ten minutes. You can close the page: an email brings you back when it is ready.
How to make yours: a five-minute checklist
Make a song first on the create page, or open one in your library and tap Music video. Then:
- One person, looking at the camera. Group photos make the renderer guess who is singing.
- Face large in the frame. Head and shoulders beats full body. Crop before you upload if you need to.
- Mouth visible, light on the face. No sunglasses, mask, microphone or hand in front of the mouth. Side profiles rarely work.
- Pick the part that sings. The slider opens on the chorus; press Preview and make sure a voice is singing for most of the 20 seconds.
- Match the voice to the face if you care about realism. A woman's photo on a baritone vocal is funny, which is fine if funny is the goal.
- Then leave it alone for ten minutes. The email will find you.
What still goes wrong
We would rather tell you than have you find out. Very long held notes can look like a frozen open mouth. Hands near the face sometimes grow an extra finger. Teeth are better than they were and still not perfect in every frame. Passages with heavy vocal effects separate less cleanly, so the mouth follows them less precisely. And the sync score we used is a rough instrument: it rewards mouths that open when the voice is loud, it cannot hear whether an "oo" looks like an "oo". We checked every clip by eye as well, and so should you.
Frequently asked questions
How long does it take to make a singing-photo video?
About ten minutes for a 20-second clip. In our end-to-end run through the live site it was 10 minutes 2 seconds from pressing the button to the video playing on the page, of which 1 minute 40 seconds was separating the voice from the song. You do not have to wait on the page: an email arrives when it is done.
What kind of photo works best?
One person, facing the camera, face large in the frame, mouth clearly visible, decent light. In our test a head-and-shoulders crop gave the mouth 151 pixels of width against 88 for the uncropped three-quarter photo of the same person, and the difference is obvious on a phone.
Why does the mouth in AI lip-sync videos look off?
Usually because the renderer was given the whole song. With drums and guitars in the signal the mouth follows the band. When we drove the same engine with the isolated voice instead, the sync score went from 0.04 to 0.39 and the mouth stopped moving through the silent gaps.
Can I choose which part of the song the person sings?
Yes. A slider picks the 20 seconds, and it opens on the chorus by default. Use the Preview button and choose a stretch where a voice is singing most of the time.
Do I need a subscription?
No. A music video is a one-time purchase, separate from the plans, and the file is yours to download and keep, with no watermark.
Does it work in other languages?
Yes. The mouth follows the sound of the voice, not the spelling, and the burned-in captions use the right script for Spanish, Portuguese, Arabic, Hindi, Chinese, Japanese and the other languages Tunely sings in.