SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

[July 2026 Edition] What are Nano Banana 2 Lite and Gemini Omni Flash? A new production style for mass-producing images quickly and cheaply to create videos


Hello, kazu here.

On July 1, 2026, Google released two new models for images and video at once: Nano Banana 2 Lite and Gemini Omni Flash.

The former is for images, and the latter is for video.

Since I regularly use AI for design and video production, I thought this would lower production costs for individuals and small to medium-sized businesses even further!

They are provided for the Gemini API and Google AI Studio.

From Google AI Studio Playground
From Google AI Studio Playground (Omni Flash Preview)

I will start with the conclusion.

  • Nano Banana 2 Lite is the fastest and cheapest image model in the Nano Banana family. It can generate images from text in 4 seconds for approximately $0.034 per image.

  • Gemini Omni Flash is a video model that creates videos from text, image, and video inputs, and allows for editing while conversing in natural language. It costs $0.10 per second.

  • By chaining these two together, you can create an end-to-end production workflow where you mass-produce images at high speed and then turn those images directly into video.

In short, a production method where you iterate through a large number of prototypes and take the best ones all the way to video has become realistically priced.

For those who are busy

I will summarize the key points regarding Nano Banana 2 Lite and Gemini Omni Flash in bullet points.

Overview

  • Google released the image generation model 'Nano Banana 2 Lite' and the video generation/editing model 'Gemini Omni Flash' today.

  • By combining both models, you can build an end-to-end multimedia production workflow from high-speed image generation to video creation.

About Nano Banana 2 Lite

  • It is positioned as the fastest and lowest-cost image model within the Nano Banana family.

  • With a low latency of just 4 seconds from text to image generation, it is suitable for interactive prototyping and high-speed visual drafting.

  • The price is set at $0.034 per 1000px resolution image, with a design that prioritizes cost efficiency.

  • While prioritizing speed, it maintains prompt fidelity, character consistency, and readability of text within images.

  • It is recommended as the upgrade path for current Nano Banana (gemini-2.5-flash-image) users.

  • Available starting today in Google AI Studio, Gemini API, and Gemini Enterprise Agent Platform. It will also be rolled out sequentially to services for general users such as Search AI mode and the Gemini app.

Positioning of the Nano Banana Family

  • Nano Banana 2 Lite: Speed-focused, for high-frequency workflows requiring ultra-low latency.

  • Nano Banana 2: General-purpose model. The "workhorse" with the best balance of performance and cost.

  • Nano Banana Pro: For complex and professional use cases. Prioritizes accuracy and advanced reasoning over speed.

  • Nano Banana (Legacy version): Legacy model. Migration to Nano Banana 2 Lite is recommended.

About Gemini Omni Flash

  • Announced at Google I/O, this model integrates Gemini's multimodal reasoning with video generation and editing.

  • It natively enables high-quality video generation and conversational editing from inputs combining text, images, and video.

  • The price is $0.10 per second of video output, which is on par with Veo 3.1 Fast.

  • Key strengths: Interactive video editing using natural language, scene control by combining multiple modalities as references, video composition leveraging real-world knowledge such as history and biology, and synchronization of text/graphics with video actions.

  • Current limitations: Video generation is limited to 10 seconds (to be extended in the future), audio reference uploads and scene extensions are not yet supported, video references under 3 seconds are accepted by the API schema but not processed correctly, and there are some limitations on character consistency during scene transitions and panning.

  • Available in public preview starting today in Google AI Studio and Gemini API. It is also available in the Gemini app and Google Flow.

Combined Usage

  • By generating images at high speed with Nano Banana 2 Lite and passing those images to Gemini Omni Flash as references, you can animate them into high-quality videos.

  • By using the Interactions API, you can perform up to three consecutive edits while maintaining session history and context.

  • Demo apps released include "Anywhere" (instantly teleporting selfies to famous world locations and turning them into video), "Space Lift" (generating multiple design proposals from room photos and experiencing them in video), and "Omni product studio" (converting still images into cinematic videos for e-commerce).

Safety and Transparency

  • Both models adopt Google's SynthID digital watermarking technology, allowing AI-generated content to be verified in Gemini apps, Chrome, Search, and more.

Below, we will look at how this impacts practical work.

What has changed with this release?

The point is that images and videos have not only improved separately, but they are now connected.

Until now, generative AI production was divided: image tools for images, and video tools for videos.
The two models Google released this time are designed with the premise of a workflow where what is created in images is passed to the video side to be animated. A model that can mass-produce images quickly and cheaply, and a model that can turn those images into video, arrived on the same day. This is the key.

Nano Banana 2 Lite (Fast and Cheap Image Generation)

Nano Banana 2 Lite (model ID: gemini-3.1-flash-lite-image) is an image model that prioritizes speed and cost above all else.

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/

In terms of numbers, it takes 4 seconds to generate an image from text, and it costs about $0.034 per 1000-pixel class image.
This is effective for use cases where you 'run dozens of prototypes.' Since it costs only a few yen per image, you can easily afford to generate 20 candidates for a social media image and choose the best one.

Speed-oriented models often sacrifice quality, but Nano Banana 2 Lite is said to maintain practical aspects such as prompt adherence, character consistency, and the rendering of text within images. Google is advising developers using the old Nano Banana (gemini-2.5-flash-image) to switch to this 2 Lite immediately to improve quality, speed, and cost.

It is available via developer-focused Google AI Studio, Gemini API, and Gemini Enterprise Agent Platform. In addition, it is expanding to consumer-facing areas such as Search's AI Mode, Gemini apps, NotebookLM, Google Photos, Stitch, Google Flow, and Google Ads. In other words, even people who don't write code will be able to encounter it within Gemini apps and Search.

How to use the Nano Banana family

What is important this time is that Nano Banana is not just one model, but a family chosen based on the use case. Understanding this will reduce wasted costs.

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/
  • Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image): Specialized for speed. For situations requiring near real-time mass generation and ultra-low latency.

  • Nano Banana 2 (Gemini 3.1 Flash Image): The general-purpose flagship. Delivers high quality with low latency, offering the best balance of performance and cost.

  • Nano Banana Pro (Gemini 3 Pro Image): For complex, professional use cases. When you need the strongest control and advanced reasoning, and accuracy is more important than speed.

  • Old Nano Banana (Gemini 2.5 Flash Image): Legacy. Switching to 2 Lite is recommended.

This way of thinking is actually very similar to the effort discussion I wrote about regarding Claude Sonnet 5. Don't try to do everything with the highest performance; assign light tasks to fast and cheap models, and reserve the high-precision models only for the critical moments. Image generation is the same: it is smart to use Lite for mass-produced drafts, and Pro or P2 for the final finishing touch.

Gemini Omni Flash (Video Generation and Conversational Editing)

Gemini Omni Flash (model ID: gemini-omni-flash-preview) is a model that combines Gemini's multimodal reasoning with video generation and editing.

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/

What's interesting is that you can input not only text but also images and videos, and you can edit videos while conversing in natural language. You can make edits with instructions like 'change the color of this person's clothes' or 'pan a little slower.' The price is $0.10 per second of video output, which is on par with Veo 3.1 Fast.

Four strengths are listed: conversational video editing, multimodal referencing that maintains consistency by combining images, text, and video, composition using real-world knowledge such as history, biology, and narrative logic, and a feature that synchronizes video movement with text and diagrams.

It is available starting today in public preview via Google AI Studio and the Gemini API. It can also be used in the Gemini app and Google Flow.

The chain from image to video is the highlight this time

Using the two separately is fine, but the real value emerges when you chain them together.

The flow is as follows: Generate images at high speed with Nano Banana 2 Lite, then pass those images as references to Gemini Omni Flash to turn them into video. Furthermore, by using the Interactions API, you can perform up to three consecutive edits while maintaining session history and context.

Google has also released demo apps where you can experience this chain. There are three: 'Anywhere,' which turns selfies into videos of you in world-famous locations; 'Space Lift,' which creates interior design proposals from photos of a room and shows them in video; and 'Omni product studio,' which turns still images into cinematic product videos. This production method is a perfect fit for product pages, property listings, and travel-related content.

How will production costs for individuals and small businesses change?

This is not an official announcement, but my own perspective from the field.

I run AI Plus Works, an internal AI organization, using Claude Code, and I build design and video production on the premise of using AI. What I strongly feel there is that the biggest bottleneck in production is not 'skill,' but 'the number of trials and the connection between post-processing steps.'

In the era when one image cost several dozen yen, I was afraid to produce many candidates. As for video, there was the hassle of having to recreate it using different tools than those used for images. These two tools address both of these bottlenecks simultaneously. You can try dozens of images for a few yen each, and you can pass the good images directly to video. From SNS images, product images, and LP assets to advertising banners and short-form videos, the range that individuals and small companies can handle on their own expands significantly.

Points to note when using

I will also write about the honest limitations. If you expect too much without knowing these, you will be disappointed.

Gemini Omni Flash is currently limited to up to 10 seconds for video. Longer durations will be supported soon.

In the API, audio referencing and scene extension are not yet available, and while 3-second video references are accepted in the schema, they are not currently processed correctly. It is also explicitly stated that character consistency may break when changing scenes or panning the camera. It is correct to approach this understanding that it is still in the public preview stage.

One more thing: generated images and videos will include a digital watermark called SynthID. This is a mechanism that allows you to check whether content is AI-generated via the Gemini app, Gemini in Chrome, or search. If you are using it for work, it is safer to operate on the premise that it is AI-generated.

Summary

Here is a 3-line summary of Google's latest release.

  • Nano Banana 2 Lite is the fastest and cheapest image model, capable of mass-producing images at approximately $0.034 per image in 4 seconds.

  • Gemini Omni Flash is a video model that allows you to create videos while conversing, using images or videos as input ($0.10 per second).

  • By chaining these two together, you can go from mass-producing images to video conversion in one go. This lowers production costs for individuals and small to medium-sized businesses.

You don't need to follow every single new model that comes out.

What matters is deciding which part of your production flow to assign to which model and at what quality level.
Mass-produce drafts with the fast Lite model, and increase precision only where necessary. Once you decide on that allocation, this becomes a highly cost-effective production environment.

FAQ

Q. Where can I use Nano Banana 2 Lite?
A. For developers, it is available via Google AI Studio, Gemini API, and Gemini Enterprise Agent Platform. Additionally, it is being rolled out to consumer-facing surfaces such as AI Mode in Search, the Gemini app, NotebookLM, Google Photos, Stitch, Google Flow, and Google Ads.

Q. How much does it cost?
A. Nano Banana 2 Lite costs approximately $0.034 per 1000-pixel class image. Gemini Omni Flash costs $0.10 per second of video output, which is at the same level as Veo 3.1 Fast.

Q. Should I switch from the old Nano Banana?
A. Google recommends switching to Nano Banana 2 Lite if you are using the old Nano Banana (gemini-2.5-flash-image). It is said to improve quality, speed, and cost.

Q. Can I create long videos with Gemini Omni Flash?
A. Currently, it is limited to 10 seconds. Longer generation is announced to be supported soon. There are also some limitations at the public preview stage, such as audio referencing and scene extension not yet being available via API.

Q. Can you tell if images or videos are AI-generated?
A. Yes. A digital watermark called SynthID is included. You can check whether content is AI-generated via the Gemini app, Gemini in Chrome, and Search.

Click here if you want to learn about Generative AI and Dify through videos

In the Generative AI course I teach, I also explain how to build applications using Generative AI and Dify.

Generative AI Course (Basic)
Generative AI Course (Practical)
AI Talent Course (From the basics of machine learning and deep learning to AI agent development with Python)

Click here for project consultations

I support the construction of business efficiency agents utilizing AI, including Dify and n8n. I offer a free initial consultation, so if you are a corporation looking to build AI agents in-house, please feel free to contact us via the link below.

For inquiries, click here.

X (formerly Twitter) is also available for inquiries.

Linktree

YouTube Channel

We also offer an AI consulting service (for individuals and corporations, 3 months of one-on-one support) to provide ongoing, hands-on support for AI utilization. You can start with a 30-minute free consultation.Click here for details on our AI consulting service

Threads

If you found this article helpful, please give it a 'Like'! 🙇

Introduction to our One-Coin Membership

We deliver technical verification logs for ChatGPT, Dify, etc., as well as ready-to-use materials like configuration files (YAML) and prompts at least once a month. You can join for just one coin, so please consider signing up.

Generative AI Lab (Membership)

#GenerativeAI #AI #GoogleGemini #ImageGeneration #VideoGeneration #NanoBanana #TriedWithAI #BusinessEfficiency #ContentCreation #Prompt #gemini #flash


いいなと思ったら応援しよう!