Creating Audio Problems with AI | I want to make a YouTube animation!
The Importance of Audio on YouTube
I have created several YouTube channels in the past, but I believe the biggest issue when creating an animation channel is the 'audio' aspect.
Not limited to animation, many people on YouTube are particular about sound, and some even listen to it like a radio.
Actually, some of the channels I run have narration using my own voice, but it's just so hard to hear.
My articulation is just too poor...
Plus, there are lip noises like 'smack' or 'pop' that bother me way too much...
The hell of editing while listening to that...
In fact, that channel initially had no audio, just text scrolling on the screen.
And then, it caused a miracle where I achieved monetization in 20 days! But for some reason, I decided to try adding 'live voice narration' halfway through, started slacking on updates, a year passed just like that, and I lost my monetization—the worst possible pattern.
I've sworn to myself that when I revive it someday, I'll stick to just text, lol.
I'm digressing, but since this is an animation, I can't just use text.
Audio is important to convey personality, after all!
Two Problems with AI Audio
AI is evolving rapidly, but I think there are two problems with AI audio.
One is that it cannot convey emotion.
The other is that it lacks consistency.
I have used two audio software programs so far.
Both allow you to easily create high-quality audio just by inputting text or words.
In both cases, while one issue is resolved, the other remains...
'VOICEPEAK' lacks emotion
One is called 'VOICEPEAK'.
If you've ever watched YouTube, you've definitely heard this voice.
Basically, it's a set that includes 3 male voice types, 3 female voice types, and a girl's voice, called the 'VOICEPEAK Commercial Use 6 Narrator Set'.
I've been using this audio for about three years, and at the time of its release, VOICEPEAK was the only one that could be used this effectively, and I was moved by the overwhelming difference in quality compared to other products.
At the time, it was the golden age of monotone types like 'Zundamon' and 'Yukkuri'.
The fact that it could add intonation to Japanese was truly at a moving level, and although the price was a bit high at just over 20,000 yen, I bought it immediately (it's now just under 30,000 yen).
Using this audio, I even recorded my personal best of 700,000 views.
However, you can't really add emotion to it.
While there are soft and excellent text-to-speech options, the emotion doesn't come across.
This is fatal for making an anime, isn't it?
Google's "speech to text" lacks consistency
In the midst of this, Google's speech to text system appeared last year.
This is truly at an impressive level!!
It's like this.
The parts where it speaks while laughing are just too impressive.
So, I thought, "Time to switch immediately!" and tried it out, but when I used it, I discovered something fatal.
The voice types are relatively abundant, and you choose from them to make it speak, but please take a listen.
Don't you think it sounds like a different person?
Actually, I'm having it speak with the same character settings.
Moreover, while you can make it laugh by inserting (laughs), there are times when it just reads out "warai" literally even when you input the exact same text.
And, it seems there are words that are difficult for the AI, so it doesn't speak them well.
With VOICEPEAK, you can finely adjust only the parts that aren't working, but with speech to text, which is basically about leaving everything to the AI, you can't just rephrase a part of it.
Because the voice will change.
Therefore, you end up having to redo the entire submitted text, but it's a common occurrence that parts that were readable before are no longer readable, and it just doesn't work out no matter what you do.
I searched for new AI voice software
This time, it's animation.
Neither VOICEPEAK, which lacks emotion, nor Google's inconsistent "speech to text" can be used.
So, I thought that the task of finding an AI that can produce the voice I imagine would be the key to success this time.
And so, in parallel with creating the concept, I searched extensively for voice apps.
There were no AI voices that were good enough, and at one point, I even seriously considered hiring a voice actor...
But, it's definitely not cheap...
2.5 yen per characterand such, so it's quite difficult to have them continue for a long time.
That's why I searched everywhere.
And finally, thanks to my hard work, I succeeded in finding one!
First, please listen.
This time, I recorded it in three parts.
What do you think?
Of course, there is some fluctuation, but it has a reasonable amount of consistency, and you can make some adjustments to the way things are said.
It's not 100% perfect, and I do think it might be better to hire a voice actor eventually, but since I don't intend to spend money on it right now, I'm very satisfied with being able to do this much.
This app is called ElevenLabs.
Actually, I found this app early on, but after listening to a few samples, I couldn't really imagine how to use it effectively, so I didn't put it on my list of candidates.
This is because you have to give instructions "in English" to add emotion (though some Japanese is okay).
Also, because there are instructions that can't be realized depending on the voice type, I had a preconceived notion that it was tricky to use.
However, the fact that you have to give instructions means that you don't add unnecessary emotion.
I decided to consider it a strength, as it means there are no arbitrary performances decided by AI, unlike Google's "speech to text."
Moreover, since you can start for zero yen, if I find better voice synthesis software before monetization, I can switch to that.
This is a pretty good point.
By the way, ElevenLabs also allows you to use your own voice as a sample to make it speak.
In other words, it's possible to make your own voice speak without lip noise, so you don't have to go through the hell of editing while reflecting on your own voice.
That said, I won't be speaking, though lol
