SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
芋出し画像

I tried making an interactive lip-sync video with the video generation AI, Dzine 🎙✚

Hello! I'm Erica, a weekend vibe coder. 😊

I usually spend my time writing code as a systems engineer, but on weekends, I live to tinker with AI tools. I am particularly passionate about lip-sync videos—those magical videos where characters move their mouths in time with audio.

I've tried various tools in the past, but my ultimate goal is to use AI to completely recreate the scene of a man and a woman naturally hosting a podcast.

This time, I'm sharing a record of my attempt to see how close I can get to that ideal using Dzine, a video generation AI that's currently the talk of the town.



🎙 Why I'm obsessed with interactive videos

I've made lip-sync videos before, but creating a scene where two people are talking actually requires a significant amount of brute force.

  1. Create a background image that fits the scene

  2. Write a script with Gemini, convert it to audio with VOICEPEAK, and split the conversation audio into separate tracks for male and female using SpeakerSplit

  3. Create a male lip-sync video with HeyGen, remove the background, and place it

  4. Create a female lip-sync video with HeyGen, remove the background, and place it

  5. Edit with video editing software (Filmora)


Honestly, it's quite a heavy lift for a weekend hobby (lol). Plus, when you composite videos that were exported separately, the feeling of them being in the same space inevitably fades. The way shadows fall and the eye movements end up looking somewhat disjointed.

That's why I was researching Kling 3.0 and ElevenLabs V3 this time. Kling 3.0 is expected to support multi-speaker dialogue, and ElevenLabs V3 has become surprisingly natural in its Japanese intonation.

In the midst of this, I found that a tool called Dzine is receiving very high praise in overseas communities. I thought, could this allow me to generate the natural dialogue I wanted from a single image? With that thought, I created an account with high expectations.


🛠 Production flow for this project (Roadmap)

Referring to YouTube and existing articles, I decided to proceed with the following four steps this time.

  1. Base image creation (Gemini / Nanobanana Pro): Create a high-quality image that will serve as the core of the podcast studio.

  2. Script creation (Gemini): Generate a natural back-and-forth script based on my note draft.

  3. Audio generation (ElevenLabs V3): Use the latest model to export expressive male and female voices. (I used the VOICEPEAK audio I created previously for this).

  4. Video generation (Dzine): Integrate the image and audio to have it turned into a video.


🎚 Step 1: Creating prompts while practicing English

Since I usually post content themed around English learning and AI utilization, I also create the prompts I feed into AI in English. Rather than just using translation software, my personal style is to use voice input (in this case, Super Whisper) as much as possible to double as speaking practice.

First, I use Gemini (Nanobanana Pro) to create the base studio scene.

Prompt: I'm making a podcast video with a man and a woman on Dzine, so please create a podcast studio-style background and create the attached image of a large horizontal sofa with the man sitting on the left and the woman on the right in high resolution (1920 x 1080px, 16:9).

At first, the sofa was too small or the two were too close together, so I couldn't quite get the ideal composition. So, I gave additional instructions.

The sofa should be a two-seater, and the man and woman should be larger and sitting a little farther apart.

Interacting with AI feels like being a director giving instructions to a cameraman, which is fun. After persistently adjusting things like 'Make the man a little bigger!' or 'Change the logo in the back!', I completed the perfect base image with the logo 'Erica and Kenny’s Podcast' included.

Completed base image

🎬 Step 2: Working in Dzine and the paywall

Now for the main dish: it's time for Dzine. The interface is very modern and easy to understand; you just select the 'Lip Sync' menu to get started.

  1. Upload the base image

  2. Specify the speakers

  3. Upload the audio files previously created with VOICEPEAK for each speaker

The amazing thing about Dzine here is that you can individually select which face in the image corresponds to which audio. Since this is a dialogue, I specify the tracks corresponding to the man's face and the woman's face respectively.

Alright, all that's left is to press the Generate button!

Just as I thought that, the screen stopped. When I logged in for free with my Google account, I had 50 credits. However, the cost required for this video generation was... 720 credits! Even if I pay, it's 3000 credits per month!

Ah, I realized this is the kind of thing that requires an adult response if you're going to do it seriously (lol). It's not priced for playing around with a free trial, but as a full-scale video production tool. But considering the effort of compositing up until now and the benefit of being able to use the Kling model, it's a small price to pay. So, I accidentally ended up paying for it.💞


⏳ Step 3: Is the AI wait time 1-2 hours!?

When I pressed the generate button, a surprising message appeared.

“Waiting for 1-2h”

At first, I doubted my own eyes. I thought it was a mistake for 1-2 minutes. But it was truly hours.

With famous tools like HeyGen, it only crops and processes the person's mouth, so it finishes in a few minutes.
However, Dzine (especially when using models like Kling) seems to try to recalculate the entire image to maintain consistency, including not just the mouth, but also the swaying of the body and the background lighting while speaking.
Since the other person is also moving their face and hands while speaking, I can imagine it is processing a massive amount of data.

I told myself this was like waiting for a movie to render, and decided to do other tasks (preparing dinner) while waiting. About an hour later, I got a notification and checked the completed video, and...

I received feedback that the audio was too quiet, so I am currently remaking it. I will replace it once it's done, so please check back later!

The audio correction didn't go well, so I created it using a different method in a separate video. Please refer to the end of the text.

🎉 Complete! The feeling that a person is really there

Watching the finished video gave me goosebumps. It is clearly on a different level from the synthetic lip-syncing we've seen so far.

・Spatial consistency: The two people are in the same lighting, and there are subtle shoulder movements and changes in gaze that match the tempo of the conversation.
・Lip-sync accuracy: The fluent Japanese and the mouth movements are perfectly synchronized.
・Reality: There is no awkwardness that says 'this is synthesized'; it feels like watching a podcast that was actually filmed with a camera.


🌟 Finally: The future of AI creativity

What I felt during this experiment is that AI video production is evolving from combining parts to integrating spaces.

Until now, we had to create backgrounds, people, and voices separately and have humans work hard to piece them together. But with tools like Dzine, the AI understands the situation itself—that these two people are talking in this space—and outputs it.

Of course, there are still hurdles, such as high credit costs and long wait times. But just a year ago, we were laughing at unnatural mouth movements, and now we are experiencing 1-hour render waits. The speed of this evolution is beyond exciting; it's even a bit awe-inspiring.

For future improvements, I'm thinking of displaying slides on a background monitor or whiteboard, linking them to the conversation, or incorporating more complex gestures.

I hope today's article is helpful for those interested in AI video generation. If you thought, 'I want to try making an interactive video too!', please give it a shot.

If you liked this article, please 'Like' and follow me; it encourages my weekend production! (lol) 💃

Well then, see you in the next AI utilization report! This was Erica, your weekend vibe-coder!👋


Special Thanks & Reference:
・Reference for starting procedure: Levelma [Generative AI Information]
Thank you for teaching me with such detailed steps.



・Reference for lip-sync technology: The Zinny Studio
Although it's anime-style, it was very helpful.



[Planned Addition]
It was so much fun that I got carried away and am currently making a vertical short video (9:16).
It currently says "Waiting for 2-3h" for the production time.
I created the audio with Elevenlabs V3, split it into male and female tracks using SpeakerSplit, and am now rendering it in Dzine. I'll post it here once it's finished!

Something that looks pretty good has been created!
The separation of the male and female voices failed (I overlooked that SpeakerSplit had failed), but it took exactly 2 hours.
It takes time, but if you create a base image for the podcast version, you can just prepare the audio and make lots of podcast videos!

*I've been told the sound is too quiet, so I'm currently remaking it. I'll swap it out once it's done, so please come back and check again!

The audio correction didn't go well, so I created it in a different way in a separate video.

I've remade it here.

Thank you for reading and watching!

いいなず思ったら応揎しよう

゚リカ@フォロバ100 | 孊び続けるAI゚ヌゞェント ぜひ応揎お願いしたす掻動費にあおおいい蚘事を぀くっおいきたすコメントも倧歓迎