Fully Automated Translation & Dubbing for English Videos - Voice-Cloning Auto-Dubbing Tool Update Ver2.2 - Python x Whisper x Gemini
Hello, this is Rcat.
This is the second update for the tool introduced here.
For those new to this, I will explain briefly: this tool transcribes the narration in a video using AI, then uses that to perform text-to-speech using reading software, and furthermore inserts subtitles all fully automatically.
This time, I have added a feature to transcribe in English, automatically translate it, and perform Japanese text-to-speech and subtitle insertion.
There are several other improvements as well, so please take a look.
Introduction
Terms of Service
Please check the terms of service in advance when using information or works.
Regarding Comments
Please check the terms of service guidelines before commenting.
English Translation and Dubbing Feature
Overview
A user who has been using my tools for a long time suggested an idea while continuing to use this tool. That is, "Can you automatically dub English?"
I often thought about things like the content I distribute, but I didn't have the idea of converting existing content for my own use, so I decided to adopt it.
This feature is realized by modifying the specifications of the text proofreading stage using generative AI, which is positioned as Step 2 of this project. Briefly explained, it is as follows.
Transcribe from an audio file
-
Pass the transcribed data to generative AI for proofreading
After text composition, insert a flow to translate into Japanese
Change to output the Japanese result
Read aloud using reading software using the transcribed text
Merge audio and video again to complete
How to use
Since the AI prompt has changed with this update, please recreate the chatbot in Dify and set a new API key.
First, switch the operation mode of the tool.
A box that was previously hidden will now be displayed. Switch this to English dubbing mode. Then, please execute Step 1.

Next is text proofreading. In the previous voice-cloning dubbing mode, this was an item that could be skipped, but since translation uses generative AI, this step must be taken.
If you select English dubbing mode at the beginning, the subsequent processing will automatically switch, so if you have the AI settings ready, just executing the step is enough.

The subsequent processing is the same as before.
Therefore, to summarize briefly: make sure to set up Step 2's text proofreading so it can be used properly. And, at the very beginning, select English translation mode. This is the only difference from before.
Use Case
Since I am terrible at English, I verified that it was working correctly using the following method.
1. Have ChatGPT generate simple English sentences.

2. Input them into Google Translate, press the play button, and record the screen.
By doing this, I create a pseudo-English video.

3. Run it through this tool.
Here are the results of the fully automated translation feature (using Gemini 1.5 Flash).
The original source is here.
Local File Search Feature
Overview
Until now, the method was to upload the video to be used, but for a tool intended for local use, uploading is just a waste, isn't it?
Therefore, I have changed it so that if the source material is in a specific folder locally, you can search by name and use it without having to upload it.


Specifically, if you create a folder named "data" inside the tool's folder, its contents will be searchable.
Then, when you try to upload a file, it checks if a file with the same name exists within that folder. If the names match, the upload is not performed, and the data from the folder is selected instead.
By doing this, you can prevent unnecessary duplication of relatively large video data.
By the way, the contents of the work folder are also searchable.
Therefore, if you want to modify and overwrite a transcription report, please change its name. Otherwise, it will not be updated because there is one with the same name locally.
Addition of Pause Feature
Overview
This is a feature that temporarily stops processing before step 3 when in all-step batch execution mode.
When using this tool, I believe the points where users want to check are after the transcription is finished or after the text composition is finished.
Therefore, I have made it possible to insert a pause at those points so that you can check or modify the files and then resume the reading and video creation process.

Specifically, an item for pausing before reading has been added when executing all steps.
If this item is set to 'Yes', the button will change once the transcription or text composition is finished.
In this state, you can check the transcription or report, choose to upload or do nothing, and then press the resume button to continue the process.
Distribution Information
Please check the distribution information from the article introduced at the top.
いいなと思ったら応援しよう!
情報が役に立ったと思えば、僅かでも投げ銭していただけるとありがたいです。