[AI Entertainment] One line of text, instant video. 2026 AI video, MVs, and avatars are truly insane
2:00 AM. I was testing new features of an AI video generator while letting anime play in the background on Abema.
I enter a prompt and hit the 'Generate' button. I wait 10 minutes. Then, the scene that was fictional just a moment ago appears on the screen. This would have been unthinkable in the past.
Now in 2026, the intersection of tech and entertainment is evolving at an insane pace. The landscape of fandom is changing.
Creating 'lip-sync videos' in an instant with Suno + DomoAI
A workflow has now been established where songs created by the audio generation AI 'Suno' are converted directly into 'singing videos'.
First, generate a song from text using Suno (supports English and Japanese). Next, feed that vocal track and lyrics into an avatar video tool called 'DomoAI'. This produces a video of an AI avatar (looking just like a human) singing the song in perfect lip-sync.
Lip-sync technology has become truly natural. The mouth moves, expressions change, and the head turns. While the limb movements of digital avatars are still a bit subtle, it's hard to tell the difference around the face—it's at a 'is this live-action?' level.
In terms of cost, with Suno (monthly subscription with a free tier) and DomoAI (offers a free trial), it ranges from 0 yen to about 3,000 yen per month. Tasks that used to cost hundreds of thousands of yen at a video production studio can now be completed on a laptop.
Thinking about applications for fandom: a workflow like 'recreating an idol's vocals with AI to make an original music video.' This is already starting in underground circles. While there are complex issues like copyright and ethics, the distance between 'technical feasibility' and 'people actually using it' has shrunk terrifyingly fast.
Google Veo 3.1 released for free. The 'democratization of video generation' has truly happened
Google released a video generation AI called Veo 3.1 to all personal accounts for free (April 2026).
You can create up to 10 videos per month, each up to 8 seconds long. No credit card required. Just account registration.
8 seconds? Isn't that short? You might think that at first, but it fits perfectly with the formats for TikTok, Instagram Reels, and YouTube Shorts. It also comes with plenty of presets, so you can generate a 'decent-looking' video just by choosing from a template.
Speaking of technical standards, as of April 2026, it has reached a level where 'short videos of about 10 to 30 seconds are indistinguishable from AI even to professionals.' In other words, they are indistinguishable from videos made by video production studios.
However, the story about Sora is interesting. OpenAI ended its provision of Sora in March. The reason was 'to concentrate resources on agent development.' In other words, they decided that having AI autonomously reason and act is more important than perfecting the quality of a single video. The tech industry's priorities have shifted from 'visualizing entertainment' to 'general-purpose AI'.
Automatic generation of '1 song = 1 MV' with SunoMV
There is another approach. A service called SunoMV automatically generates a full music video when you upload a song created in Suno.
It analyzes the content of the lyrics and generates 'matching scenes' for each phrase. If it's a 3 to 5-minute song, the whole thing comes out as a story-driven music video.
Camera angle switching, color tones, and the like are handled automatically by the AI 'reading the meaning' without a technician having to 'adjust it by hand'.
For creators like the Vtuber community or doujin music circles who want to drop original songs every month, this is a revolutionary tool. The production cost per song drops dramatically.
2026's Suno MV supports 13 visual styles, allowing for optimized output in vertical format for TikTok and Instagram, and horizontal format for YouTube.
VTuber and digital avatar technology is overtaking humans in terms of 'human-likeness'.
The evolution of digital avatars is, personally, the biggest shock.
By 2026, AI VTubers have reached a level where conversation timing, emotional vocal fluctuations, and eye movements are almost indistinguishable from humans. Many streamers don't want to show their faces, but now they can completely hide them with high-quality digital avatars.
A cross-tool workflow has been established: Midjourney (image generation) → Photoshop/Stable Diffusion (precision editing) → Dream Machine/Gen-3/Kling (video conversion) → After Effects (final adjustments).
The realization that 'generative AI is the raw material, and editing determines the final form' has permeated the production side. In other words, we have reached a stage where AI doesn't fully automate everything, but rather 'AI provides the materials, and the craftsmanship begins from there.'
In terms of fandom, the flow of turning 2D characters into 3D, then into VTuber avatars, and then streaming, can now be done by individuals. However, this is still complicated by copyright, personality rights, and character usage licensing. The technology is there, but social rules haven't caught up.
Runway Gen-4.5, Pika 2.5, Luma Dream Machine. There are too many choices, and it's confusing.
The competition is fierce.
Runway Gen-4.5 automatically builds 'cinematic continuity' from text prompts. In other words, it generates narratives like 'a scene of the sun rising → a scene of the city waking up → people starting their activities' on its own from a single line of text.
Pika 2.5 specializes in anime-style video generation. It properly delivers the 'quality of animation' that otaku are looking for.
Luma Dream Machine has natural camera movements. It has a high level of sophistication in cinematic expression, such as 'cinematic diagonal shots' or 'zoom-ins'.
The criteria for choosing depends on 'what you want to make.' If it's an MV for your favorite character, color tone is important, so use Runway. If it's anime-style, use Pika. If it's cinematic quality, use Luma.
But honestly, for beginners, Google Veo 3.1 is enough for anyone. It's free, the quality is high enough, and the satisfaction level is high.
Since the technology has reached a professional level, 'what to express' has become more important now.
What I realize in 2026 is that 'the democratization of video production has already happened.'
In the past, making videos required skills, software, time, and money. Now, if you have a text prompt and an internet connection, anyone can output professional-level quality.
When that happens, the issue becomes 'imagination, not technology.' What do you want to make? What kind of emotions do you want to express? How do you want to film your favorite character?
In that sense, people with strong 'obsessions'—otaku, fans, and subculture enthusiasts—are actually in the strongest position. The reason is that their discernment is clear. Their criteria for 'this is good' or 'this is bad' do not waver.
The tech side is racing to 'update tools,' but it's possible that the 'creativity' on the user side is running ahead now.
Otaku fan activities have become the cutting edge of video entertainment.
Until recently, the cutting edge of entertainment was created by major media outlets and film production companies—those with the money and manpower.
Now, individual VTubers, doujin music circles, fan enthusiasts, and anime fans are using AI to attempt to 'visualize the ideal image of their favorite characters.'
Create a song with Suno -> Create an MV with Suno MV -> Post to TikTok and Twitter. Within 24 hours.
DeepL (translation) -> Create a multilingual version with Suno. Global expansion.
Individuals can now do all of that.
In the end, the intersection of tech and entertainment was more complex than we first imagined.
The good side: Individual creativity is being visualized. Fan activities are becoming video content. Applications that ignore copyright and portrait rights are certainly happening, too.
The complex side: The line between 'AI generation' and 'human creation' is already blurred. How do we handle copyright fees? How do we handle character usage permissions? What is the legal status of digital avatars?
The interesting side: Even so, someone uses it. Someone tries it. New culture is born through ingenuity.
Late at night, posts on Twitter saying 'I made this MV' are increasing. Videos that weren't even imagined 24 hours ago now exist.
Watching that scene, I really feel that 'the speed of technological evolution is faster than human ethical judgment.'
But that's not all bad. A new degree of freedom in expression is being born. How to use it is left up to us humans.
いいなと思ったら応援しよう!
あなたの心に、花束を。
チップありがとうございます。あなたのこと、ちゃんと覚えています。
