Blog

How to Make a Photo Sing (and Look Real): What Six Test Renders Taught Us

The Tunely Team · 2026-09-17 · 9 min read
Short answer

Three things decide whether a singing-photo video looks real, in this order. First, what drives the mouth: feed the renderer the isolated voice, not the full song. In our test that alone moved the mouth-to-voice sync score from 0.04 to 0.39 on the same engine and the same photo. Second, direction: without a one-line instruction the singer walked towards the camera and the framing drifted; with it, the person stayed put and sang. Third, the photo: a head-and-shoulders crop gave the mouth 151 pixels of width against 88 for the uncropped shot of the same person, so the articulation is simply easier to see. The engine mattered less than all three.

We were not happy with our first music videos: the person in the photo wandered about, the mouth seemed to follow the band rather than the singer, and the sound was mono. Instead of guessing, we set up a small bench test. One photo, one chorus, the same ten seconds, rendered six different ways, each one measured the same way. This is the write-up, with the clips, the numbers, and what we changed on Tunely the same day.

Want it now? Make a music video from a photo — free, in seconds.

The test: one photo, one chorus, six renders

The photo is an AI-generated portrait that belongs to us, so nobody's real face sits on a test bench. The song is a birthday song made on Tunely, and every render uses the same ten seconds, starting one second before the chorus.

For each render we measured four things. The sync score is the correlation between how far the mouth is open in each frame and how loud the isolated voice is at that instant: 1.0 would be a mouth that opens exactly with the voice, 0 is a mouth moving at random. The mouth lag is the time shift at which that correlation peaks. The mouth width is how many pixels the mouth gets in the finished frame. And we noted what the file itself was: size, frame rate, mono or stereo.

Two clips to start with. On the left, the way our first version worked. On the right, the way Tunely makes them now.

Before: full song as the driver, no direction, uncropped photo
After: isolated voice, one line of direction, head-and-shoulders crop

About the sync score: it needs a clip with pauses in it, because it works by comparing loud moments with quiet ones. On a passage where the singer never stops, every render scores low. Compare scores only between renders of the same clip, which is what every table below does. We have no human baseline in this test, so read the numbers against each other.

Finding 1: what drives the mouth matters more than the engine

A finished song is a voice plus drums, bass, guitars and everything else. Hand all of that to a renderer and the mouth tries to follow all of it. So we separated the voice from the song first and let the bare voice drive the mouth, then put the full stereo song back under the picture at the end.

Same engine, same photo, same ten seconds. The only difference is what the renderer listened to.

Driven by the full song: the mouth keeps moving through the gaps
Driven by the isolated voice: it opens with the words
What drove the mouthThe full songSync score0.04Mouth opening, loud moments0.115Mouth opening, quiet moments0.107
What drove the mouthThe isolated voiceSync score0.39Mouth opening, loud moments0.213Mouth opening, quiet moments0.102

Look at the last two columns. Driven by the full song, the mouth was open about as much when the singer was silent as when she was belting: it was keeping time with the band. Driven by the voice, it opened twice as wide on the loud syllables as in the gaps. This was the single biggest improvement in the whole test, and it has nothing to do with which engine you use.

Finding 2: give the singer a direction, or the singer improvises

Our first version sent a photo and a sound file and nothing else. Left to itself, the renderer invents a performance: in the "before" clip the woman walks towards the camera, the framing tightens, a hand goes to the chest. It is lively. It is also not the photo you uploaded any more, and over twenty seconds the face can drift away from the person you know.

One sentence fixes it. Every render now carries the same direction: sing to the camera, stay in place, natural head and shoulder movement, the camera does not move. Here is the same uncropped photo rendered the old way and the way we do it now. The direction is what keeps her in place: Engine A further down is the same engine as the "before" clip, and once it had the direction it stayed put too.

No direction: she walks in, the framing drifts
Same photo, the current pipeline: she stays and sings

Timing improved as well. In the "before" render the mouth ran about 200 milliseconds ahead of the voice, which is the uncanny feeling of a dubbed film. The current pipeline peaks within one frame of the voice.

Finding 3: the photo's framing decides how much singing you can see

We rendered the same person twice on the same engine: once from the uncropped three-quarter photo, once from a head-and-shoulders crop of that very photo.

PhotoUncropped, three-quarter lengthMouth width in the finished video88 pxSync score0.41
PhotoHead-and-shoulders crop of the same photoMouth width in the finished video151 pxSync score0.39
PhotoFor reference: the "before" clipMouth width in the finished video68 pxSync score0.26

The sync is equally good in both. What changes is how much of it you can see: the crop gives the mouth 1.7 times the width, so consonants and held vowels read clearly on a phone screen. If the photo you love is a full-body shot, crop it to the chest before you upload it. You lose the shoes and gain the performance.

The engines, given identical inputs

With the inputs fixed (isolated voice, the same direction, the head-and-shoulders crop) we ran four engines. We letter them rather than name them: engines change every few months, and the three findings above held on all of them.

Engine A with the new inputs
Engine C with the new inputs
EngineA, the one we used beforeOutput832×1120, 25 fpsSound as deliveredmonoSync score0.25Mouth lag160 ms earlyRender time, 10 s clip4 min 1 s
EngineB, the one Tunely uses nowOutput1168×1760, 30 fpsSound as deliveredstereo after our remix stepSync score0.39Mouth lagwithin one frameRender time, 10 s clip4 min 36 s to 5 min 48 s
EngineC, a lighter tier of BOutput784×1168, 30 fpsSound as deliveredstereoSync score0.26Mouth lagwithin one frameRender time, 10 s clip3 min 33 s
EngineDOutputno outputSound as deliveredn/aSync scoren/aMouth lagn/aRender time, 10 s clipfailed twice on this photo

Engine D has rendered other photos for us without trouble. It refused this one twice with an internal error, which is a data point of its own: a video you cannot get is worse than a slower one.

What we changed on Tunely the same day

Everything above went into the product on 17 September 2026. When you make a music video on Tunely now:

  • The voice is separated from your song and drives the mouth. Your full stereo song goes back under the picture afterwards.
  • Every render carries the same direction: stay in place, sing to the camera, the camera does not move.
  • The 20-second clip opens on the chorus instead of the intro. A person standing silently through four bars of guitar is not a music video.
  • The lyrics are burned in as captions, timed to the song, in the right script for the language.
  • The finished file is up to 1168×1760 at 30 frames per second (its shape follows your photo), with no watermark.
  • It takes about ten minutes. You can close the page: an email brings you back when it is ready.
A full 20-second video made through the live site, start to finish in 10 minutes 2 seconds

How to make yours: a five-minute checklist

Make a song first on the create page, or open one in your library and tap Music video. Then:

  • One person, looking at the camera. Group photos make the renderer guess who is singing.
  • Face large in the frame. Head and shoulders beats full body. Crop before you upload if you need to.
  • Mouth visible, light on the face. No sunglasses, mask, microphone or hand in front of the mouth. Side profiles rarely work.
  • Pick the part that sings. The slider opens on the chorus; press Preview and make sure a voice is singing for most of the 20 seconds.
  • Match the voice to the face if you care about realism. A woman's photo on a baritone vocal is funny, which is fine if funny is the goal.
  • Then leave it alone for ten minutes. The email will find you.

What still goes wrong

We would rather tell you than have you find out. Very long held notes can look like a frozen open mouth. Hands near the face sometimes grow an extra finger. Teeth are better than they were and still not perfect in every frame. Passages with heavy vocal effects separate less cleanly, so the mouth follows them less precisely. And the sync score we used is a rough instrument: it rewards mouths that open when the voice is loud, it cannot hear whether an "oo" looks like an "oo". We checked every clip by eye as well, and so should you.

Frequently asked questions

How long does it take to make a singing-photo video?

About ten minutes for a 20-second clip. In our end-to-end run through the live site it was 10 minutes 2 seconds from pressing the button to the video playing on the page, of which 1 minute 40 seconds was separating the voice from the song. You do not have to wait on the page: an email arrives when it is done.

What kind of photo works best?

One person, facing the camera, face large in the frame, mouth clearly visible, decent light. In our test a head-and-shoulders crop gave the mouth 151 pixels of width against 88 for the uncropped three-quarter photo of the same person, and the difference is obvious on a phone.

Why does the mouth in AI lip-sync videos look off?

Usually because the renderer was given the whole song. With drums and guitars in the signal the mouth follows the band. When we drove the same engine with the isolated voice instead, the sync score went from 0.04 to 0.39 and the mouth stopped moving through the silent gaps.

Can I choose which part of the song the person sings?

Yes. A slider picks the 20 seconds, and it opens on the chorus by default. Use the Preview button and choose a stretch where a voice is singing most of the time.

Do I need a subscription?

No. A music video is a one-time purchase, separate from the plans, and the file is yours to download and keep, with no watermark.

Does it work in other languages?

Yes. The mouth follows the sound of the voice, not the spelling, and the burned-in captions use the right script for Spanish, Portuguese, Arabic, Hindi, Chinese, Japanese and the other languages Tunely sings in.

More from the blog