Re-evaluating 'AivisSpeech' that I had shelved: Comparison with VOICEPEAK and how to use them for business
Hello.
Which speech synthesis software do you all use for video production and narration creation?
I usually use 'VOICEPEAK' as my main tool because of its overwhelming reading quality.
However, there was actually another piece of software that I had left dormant on my PC for a long time. That is the free speech synthesis software 'AivisSpeech'.
Although I installed it once in the past, due to the 'Anneli issue' that caused a stir in the community in the summer of 2025, I ended up leaving it alone without using it.
For details, I would appreciate it if you could refer to the video below by Yupro.
However, I recently had a sudden thought and checked the official website for the first time in a while, and I was surprised. Although Anneli was gone, a large lineup of attractive alternative voice models was available.
'With the current AivisSpeech, couldn't I use it extensively for practical work...?'
With that in mind, I actually had it read long texts and made adjustments assuming business use, and thoroughly compared it with the VOICEPEAK I usually use.
In this article, I will talk about my impressions after using AivisSpeech again, having shelved it once before.
1. AivisSpeech, I actually 'shelved' it once
AivisSpeech appeared as software capable of high-performance speech synthesis for free.
When I first learned of its existence, I also expected, 'This could be a powerful weapon for video production!', and immediately installed it on my PC.
However, as a result, I did not use it in earnest at that time and quietly 'shelved' it.
The biggest reason for that was the 'Anneli issue' that was a hot topic in the community at the time. As I touched on briefly at the beginning, the AivisSpeech of that time had various controversies regarding Anneli, who was the default character, so I somehow left AivisSpeech alone without using it.
Some time passed since then, and the other day, due to a sudden impulse, I checked the official website, wondering, 'By the way, what is happening with AivisSpeech now?'
Then, to my surprise, Anneli was gone, and various alternative voice models were being distributed on a platform called 'AivisHub'.

When I looked at the site, a wide variety of voice models created by creators were lined up, and an environment was in place where I could download my favorite voices and use them immediately on AivisSpeech. I was completely betrayed in a good way, thinking, 'Wait, can I choose from so many voices now?'
Feeling that 'with the current AivisSpeech, it is worth trying it out properly once more!', I launched the software again.
I compared it with VOICEPEAK, which I currently use as my main tool.
2. 'VOICEPEAK' for quality, 'AivisSpeech' for ease of use
I immediately poured text into it and generated various voices. How did it compare to 'VOICEPEAK', which I use extensively?
As a test, I had each software read the opening text of this article, and I have posted the results in an unadjusted state (I have only corrected some of the pronunciations of the characters).
What do you think?
This is just my personal opinion, but I have the impression that VOICEPEAK allows you to thoroughly craft quality, while AivisSpeech is for quickly and easily getting things into shape, and if VOICEPEAK's level of perfection is 90 points, then AivisSpeech is about 75 points, I suppose.
You might think, "Wait, isn't 75 points a bit mediocre?" but that is definitely not the case.
If you include the surprise that it speaks this naturally for free, it's a 75-point score that easily clears the practical threshold. However, VOICEPEAK's performance is just too good.
So, where does this "15-point difference" come from?
It comes from the difference in "fine-tuning capability."
When creating video narration, you always end up with specific preferences, such as "I want to change the intonation of just this one word," "I want to make the pause between this phrase and that one just a fraction of a second longer," or "I want to lower the tone of the ending just a little bit."
VOICEPEAK completely addresses these "itchy spots."
The sample I mentioned earlier had no adjustments made for the sake of a default comparison, but you can control the details exactly as you wish, such as the accent of each word and the length of the speech.

On the other hand, AivisSpeech has no adjustment items other than accent, so it is not suited for the kind of precise voice adjustment that VOICEPEAK offers.
There are inevitably times when you have to compromise at a certain point, thinking, "I just can't shake the feeling that this intonation is off!"

Some might think, "Then isn't VOICEPEAK the only choice?"
However, what I realized after actually using it is that AivisSpeech has a weapon called "overwhelming ease of use" that more than makes up for this 15-point difference.
3. Two reasons why I want to incorporate AivisSpeech into video production
I wrote in the previous chapter that "fine-tuning is VOICEPEAK's exclusive domain," but conversely, AivisSpeech has the advantage that you can use it quickly without spending time on fine-tuning.
While actually working, I realized two reasons why I can incorporate this into my video production workflow.
Reason 1: It creates a voice that is "perfectly listenable" without any adjustments.
When using speech synthesis software, the task that consumes the most time is "voice adjustment." If you spend time fixing awkward intonation and adjusting pauses, time passes in the blink of an eye.
However, even when you just "paste" the text (simply inputting it), AivisSpeech generates audio at a level that is sufficiently "listenable" without any accent adjustment.
Of course, it's not perfect, but it's like the "75-point quality" I mentioned earlier is output instantly the moment you paste the text. For YouTube explanatory videos or content that conveys information with good tempo, this speed of "75-point raw input" is an overwhelming weapon. It is also perfect for uses like "I just want to quickly apply audio to check the overall picture of the video."
Reason 2: It is overwhelmingly easier for reading long sentences.
Personally, I felt this was the biggest advantage of AivisSpeech. The work is just so easy when you want to have it read long lines or long manuscripts all at once.
When you try to process long sentences with VOICEPEAK, you inevitably start worrying about the details and tend to get stuck in a swamp of "fix this, fix that..." Also, the longer the text, the more the effort required to maintain and adjust the tone throughout the whole thing skyrockets.
In that respect, AivisSpeech is good at "reading the whole thing naturally in a rough way" in a good sense. It is a pleasure to have it return a continuous audio track with few awkward parts overall, even when you dump in a long narration manuscript that lasts for several minutes.
"VOICEPEAK for fine-tuning details," and "AivisSpeech for processing long texts quickly and shaping them." I have now developed a clear vision for how to use them differently.
4. Disadvantages of AivisSpeech (and potential solutions)
While AivisSpeech is powerful and convenient for long texts, I found a few points that felt "almost there" when trying to incorporate it into practical work, especially for business-related video production.
Here, I will introduce those disadvantages and my own solutions.
Disadvantage 1: AivisHub lacks formal voices suitable for business narration
AivisHub offers a wide variety of voice models, which is fun just to browse, but the overall trend is that voices with strong character traits (anime-style, cute, or energetic voices) are the mainstream.
Therefore, when trying to find "calm and formal narration voices" that can be used as-is for corporate service introduction videos, serious explanatory videos, or internal training materials, it is a bit of a struggle due to the limited options.
[Solution] Lowering the pitch of the "Mai" voice model
Thinking, "If there are few business-like voices, I'll try adjusting one to create it," I decided to adjust the voice model called Mai.
By default, the voice sounds a bit young and cute, but when I lowered the "pitch" in the settings, the center of gravity of the voice dropped, and it changed to a much calmer, more intellectual tone. With this one tweak, I was able to create a persuasive narration that is perfectly suitable for business use.
(I have posted the adjusted voice below.)
Disadvantage 2: Overall performance is a bit sluggish (*Could it be my Mac environment?)
A bottleneck I felt while actually working was the "sluggishness of the app itself."
The response when entering text, tweaking various settings, or generating audio felt sluggish overall, with a sense of being one beat behind.
It might be a compatibility issue with the Mac environment I am using, but when I want to proceed with work quickly and rhythmically, this sluggishness felt slightly stressful.
5. Conclusion
Although I had once "shelved" AivisSpeech, after pulling it out again for the first time in a while, I realized it has evolved into a more practical tool than I imagined.
AivisSpeech has this much potential for free, so if you haven't touched it yet, or if you are like me and closed it away in the past, please experience the current version (and the richness of AivisHub).
It will surely become a new weapon for your video production!
Thank you for reading this far!
