5-Second AI Videos Don't Need Text | Beginners Should Use 'Light' Instead of 'Text'
When creating a 5-second video with AI,
you might think, "Since I'm at it, I want to include words like 'Thank you' or 'Like' in the video."
I also used to feel that it was easier to convey what I wanted to say if there was text included.
However, after reading about the creative experiences of others, I realized something new.
But I have realized something again after reading about other people's creative experiences.
Including text in AI videos is surprisingly difficult for beginners.
And you can convey your feelings well enough in a 5-second video without forcing text into it.
In fact, for beginners, it might be easier to create beautiful videos by using 'light' and 'movement,' which AI excels at, rather than text.

It is difficult to include text in AI videos
Image generation AI and video generation AI are not very good at displaying text accurately.
Especially since Japanese characters have complex shapes,
the text gets distorted,
it turns into different kanji,
parts of it become unreadable,
the position of the text shifts,
or the shape changes halfway through,
these kinds of things happen easily.
Even in still images, it is difficult to display text accurately, and in videos, you also have to maintain that text for several seconds.
The background moves, too.
People move, too.
Light and cameras move, too.
Within that, you need to keep displaying text accurately and in a readable state.
For a beginner, this is a fairly high-difficulty instruction.
In 5 seconds, you also need time to read the text.
A 5-second video is truly short.
If you include text, you also need time for the viewer to notice the text and read its content.
For example,
even if you include text like 'Thank you for the like',
if the display time is short, the video will end before they can read it.
Conversely, if you try to display the text for a long time, the time available for character movement and production effects decreases accordingly.
If you try to fit all of the following into 5 seconds, the amount of information becomes too high:
character movement
changes in expression
background effects
text display
time to read the text
the amount of information becomes too high.
I think it is easier to succeed as a beginner if you don't make the video take on too many roles.
You can add text later.
If you really want to add text to your video, you don't need to have the AI video generator create the text for you.
First, create a video without text, and then after it's finished, you can overlay text using a video editing app like CapCut.
This method has the following advantages:
You can use the correct text
You can choose your favorite font
You can adjust the display position
You can decide how long it is displayed
You can edit it as many times as you like
These are the benefits.
Have the AI create the visuals, and add the text using an editing app.
It is easier to achieve a stable result by separating the roles.
But is text really necessary?
However, I felt something else this time as well.
In the first place, text is not essential for a 5-second video.
Even without writing 'Thank you', you can show:
A smile
Receiving a heart
Holding something dearly against the chest
Being enveloped in light
Fireworks turning into a heart
Waving towards the camera
With just movements like this, gratitude and joy are conveyed sufficiently.
You can create videos where the viewer can receive emotions without needing to explain them with text.
In fact, for a short time like 5 seconds, it might be better to have the viewer feel the atmosphere rather than explaining it with words.
After all, you can put text in the message field for 'like' thank-yous.
AI is better at 'light' than 'text'
Video generation AI is not good at text, but it is relatively good at lighting effects.
For example,
Particles of light
Sparkles
Petals
Wind
Waves
Stars
Fireworks
Glowing hearts
Soft backlighting
These kinds of expressions leave an impression even in short videos.
Moreover, they don't need to be in the 'correct shape' like text does.
Even if the shape changes slightly, if it's sparkles or light, it looks natural as an effect.
Rather than forcing AI to do things it's not good at, it's easier to create beautiful videos by using expressions it excels at.
Why 'difficulty level' is necessary in prompts for beginners
Previously, I created and published a prompt to have AI generate ideas for 5-second videos.
I included instructions in that prompt to have it write the 'difficulty level' for each idea.
At first, I included it as an item to make it easier for beginners to choose.
But this time, I realized that difficulty level has another important meaning.
That is,
for beginners to avoid ideas that contain elements AI is not good at.
Even if an idea looks simple,
displaying accurate text,
performing multiple actions simultaneously,
switching scenes multiple times in a short period,
having it hold small objects accurately,
the difficulty level rises sharply when instructions like these are included.
For ideas aimed at beginners, it's not just about being 'cute' or 'interesting,' but also,
whether the content is something the video generation AI can easily succeed at,
is also important.
Perhaps there is a reason why the sample video has no text either
The sample video for the '2nd Ski Video Contest: Summer Battle' hosted by Naoyuki-san also used no text, featuring an effect of drawing a heart with fireworks.
At first, I thought it was just a purely cute and easy-to-understand effect.
But thinking about it now, that video makes a lot of sense as an example for beginners.
Fireworks, light, hearts.
All of these are easy for generative AI to express, and the meaning is conveyed even in a short time.
Even without words like "Like" or "Thank you," you can understand the sentiment just by watching.
It might have even been designed with ease of success in mind as a sample video.
(Naoyuki-san, please point it out if this part is strange > <)
I also make my 5-second videos without text
I didn't include any text in the 5-second video I made last time.
A heart of blue light emerges from the waves, a character catches it, and holds it carefully against their chest.
There is no "Thank you" text.
But I believe the feeling of gratitude was sufficiently expressed through the expressions, movements, and lighting effects.
It's not that it won't be understood because there's no text; it's precisely because there's no text that you can focus on the visuals themselves.
A 5-second video is not something to explain, but something to be felt.
Thinking of it that way, you can definitely make an emotional video even without text.
Beginners should start with "videos that make you feel"
For those who are about to make their first 5-second AI video, I recommend starting with a video without text.
Instead of writing the words you want to convey directly on the screen,
try thinking,
"How would I express these words through movement or light?"
If it's "Thank you",
Offering a heart
Receiving light
Bowing deeply
Breaking into a big smile
A flower blooming
If it's a 'like',
hearts fly out
stars gather
fireworks turn into hearts
a character turns around happily
In this way, you replace words with visuals.
Instead of making them read text, have them feel it through the video.
That might make for a more 'AI video-like' creation.
↓Click here for an article for PolloAI beginners↓
(I wrote this on the day I first registered for PolloAI)
Summary
Adding text to AI videos is quite difficult for beginners.
The text can get distorted, become unreadable, or there might not be enough time to read it within a short video.
If you want to add text, the most stable method is to add it using an editing app after creating the video.
However, for 5-second videos, you don't necessarily need to include text.
Things that AI excels at,
Light
Heart
Fireworks
Expression
Movement
Sparkle
If you use these, you can convey emotions without words.
Beginners should start by not including text and using expressions that AI excels at.
And instead of making videos that explain, make videos that make people feel.
I think this is one of the tricks to making a successful 5-second AI video.

Bonus | I tried making this video this time

The prompt I released last time is a suggestion taught to me by Gemini after putting it into practice.
1st Place: ⑯ Summer vacation jumping out of a trunkReason for selection:
It is a perfect composition that instantly conveys the theme of 'travel' and the presence of 'Rai-chan'.
The movement of 'trunk opens -> jumps out' fits perfectly into the 5-second timeline, giving the viewer visual surprise and good pacing.
Even with generative AI, the motion of 'an object opening and a character coming out' is easy to generate, and its strength is that it is easy to improve the quality.
It said the difficulty level is medium.
I wanted it to be a closed trunk opening and something jumping out from inside, but it was a bit impossible to express that with 5 credits~😅
There are issues like the straw hat appearing and disappearing,
but...
If I try too hard, I'll run out of credits, so I'm ending it here for now.
My goal is to mass-produce 5-second videos with low credits, so I'll save my credits🤭
I won't be entering this work.
I'll keep making more🤭
☕Afterword
I realized that making 5-second videos is quite a battle against frustration.
Everyone is working so hard... yes, that frustration will become the fuel for your growth.
Best of luck to you as you go through trial and error✨✨
Fight on!
#AIVideo
#VideoGenerationAI
#GenerativeAI
#AIUtilization
#ForBeginners
#noteSkiVideoContest
#PolloAI
いいなと思ったら応援しよう!
いつも温かい応援をありがとうございます。
いただいたチップは、記事づくりや画像・動画制作など、今後の創作活動に大切に活用させていただきます。
これからも楽しんでいただける記事を作っていきます🦁✨
