Gemini Omni Flash is the video version of Nano Banana
A single selfie video taken at Oguchi Beach in Itoshima turned into the Shibuya Scramble Crossing, then a snowy mountain, and finally a snowboarder when processed by Gemini Omni Flash. AI that creates videos from scratch is no longer rare. But honestly, to use a phrase in the style of Makoto Shiina, I was "blown away" by a video AI that can "edit".

In a short video set in an office, a boss puts his hand on a female employee's shoulder in the style of Go Nagai, and she brushes it off. When I followed up with, "Depict the scene where she tears her clothes and transforms with great impact," a video appeared showing the man in the suit transforming into a muscular figure. The source material was just a single character sheet. The instructions were just Japanese conversation. I felt the scent of a future where the very process of creating storyboards and then opening editing software would become unnecessary.
You can use Omni Flash not only in Gemini, but also in flow and higgsfield.
https://labs.google/fx/ja/tools/flow

Depict the scene where she tears her clothes and transforms from attachment 1 to attachment 2 with great impact. Go Nagai style. The background is an office late at night.
I was standing alone on Oguchi Beach in Itoshima.
Wearing a T-shirt and a black cap, with the green mountains in the background reminiscent of my days as a summer strawberry farm manager. I turned my smartphone and took a ten-second selfie video. It was just that, an unremarkable piece of footage.
However, when I threw this video into Google's Gemini Omni Flash and typed, "Please change the background to the Shibuya Scramble Crossing in Tokyo," the next moment, I was standing right in the middle of Shibuya. The ocean turned into buildings, and the sound of waves turned into the bustle of the crowd. The person's pose and expression remained completely intact, while only the background was replaced entirely. I don't understand the logic at all. I don't understand it, but it worked. That is all that matters.


Getting carried away, I continued. I typed, "Change the outfit to a snowboarder's attire." Then, the same person, in the same standing pose, was standing at the Shibuya crossing dressed in yellow snow gear. When I added, "Change the background to a snowy mountain. The sky is bright blue," it transformed into footage of me holding a snowboard on a real snowy mountain. The clincher was the phrase, "Make it look like I'm snowboarding on a snowy mountain." What came out next was a video of me gliding down the slopes, kicking up snow. From a beach selfie to snowboarding on a snowy mountain. It only took four rounds of conversation.

I will be honest here. Gemini Omni Flash is a brand-new model announced at Google I/O on May 19, 2026, and public preview began on June 30 via the Gemini API and Google AI Studio. The cost is $0.10 per second of video output. Currently, the maximum length that can be generated is ten seconds, and it does not yet support audio editing or video extension. Google positions this as the "video version of Nano Banana." They are trying to do with video the same thing that Nano Banana did by allowing images to be edited through dialogue.
Conventional video generation AI was, in short, a one-shot deal. Even if you just wanted to change the lighting to evening, rewriting the prompt would change the composition and the character's facial features entirely. You had no choice but to keep rolling the gacha until you were satisfied. However, Omni Flash uses a mechanism called the "Interactions API" that allows you to pinpoint and fix only the parts you want to change. Jumping from the beach video to Shibuya, transforming from Shibuya to a snowboarder, and flying to a snowy mountain. All of this proceeds while maintaining the original person's movements, expressions, and camera work. This is the true nature of "conversational editing."
Until now, I had been preoccupied with "creating 1 from 0" in AI manga, Kindle, and video editing. But what I realized while tinkering with Omni Flash is that what really eats up time is not 0 to 1, but the work of fixing 1 to 99. Changing the background. Changing the costume. Fixing the expression. Retaking, redrawing, and re-rendering every time—that effort is now finished in a single round of conversation. My, what an era we have come to.
By the way, the generated video is embedded with a watermark called SynthID, and there is also a mechanism to verify that it is AI-generated via the Gemini app or search. I think the fact that convenience and proof of origin are advancing simultaneously is a sense of balance typical of the current AI industry.
And so, I traveled from the coast of Itoshima to Shibuya and even to snowy mountains without moving a single step. It is certainly a convenient era we live in, but the fact that I can feel like I have 'been there' without breaking a sweat or getting a sunburn is a bit lonely, and perhaps a little frightening as well.
Thank you for reading. Next time, I plan to write a practical guide on how far this technology can be applied to the production of manga backgrounds.
This is the student's note.
Please follow them.
