Home Pricing Google Gemini

Google AI Pricing (August 2026)

Last updated:

Latest update (August 8): Google DeepMind open-sourced WeatherNext 2 and WeatherNext Cyclones. The release has no per-token or per-forecast Google model rate: Mini can run on a free Colab TPU, while full deployment and forecast-data access can incur infrastructure and platform costs. Read the WeatherNext cost guide →

All Google Gemini Models — Price per 1M Tokens

Showing 4 current models. .
Showing 4 grouped models from 4 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

WeatherNext pricing and access

WeatherNext is not a Gemini API model. Google has released WeatherNext 2, WeatherNext Cyclones, and a smaller Mini checkpoint as research code and model weights. Code and notebooks use Apache 2.0, while other repository materials use CC BY 4.0. There is no published Google per-token, per-run, or per-forecast model price, so WeatherNext does not appear as a zero-dollar row in our canonical pricing API.

The official Mini notebook recommends a free Colab v5e-1 runtime. Larger checkpoints require a v5p TPU or sufficient GPU memory; Google says non-Mini models need H100-class memory on GPU. Operational costs can also include initial-condition data, storage, transfer, orchestration, and forecast verification. Daily WeatherNext outputs are available through BigQuery, Earth Engine, and Cloud Storage, where normal platform usage charges may apply.

WeatherNext is also outside the current 49-task language-model Labs harness. A valid comparison needs pinned atmospheric inputs, forecast verification data, cyclone-specific skill metrics, ensemble-size controls, and accelerator-hour logs. See the WeatherNext pricing and cyclone-model analysis for the full cost boundary and explicit Labs blocker.

Managed Agents pricing

Managed Agents in the Gemini API now use Gemini 3.6 Flash by default. Google bills all input, output, intermediate input, and reasoning tokens at the selected model's standard rates, with normal tool fees on top. CPU, memory, and sandbox execution are free during the preview period. Free-tier projects can also try Managed Agents within Google's quotas.

The live pricing table on this page shows the current Gemini 3.6 Flash rates. Autonomous loops can generate many turns, so use the new token ceiling and estimate cost per completed workflow. See the Managed Agents pricing and hooks guide for rollout advice and the exact billing boundary.

Antigravity CLI also appears in a community omp setup as a separate search tool behind DeepSeek V4 Flash and GPT-5.6 Luna. Read the oh-my-pi cost and setup guide for the quota boundary and end-to-end Labs blocker.

Google's Gemini API stands out in July 2026 for a large production context window, competitive model pricing, and a managed-agent preview with no separate environment-compute fee. Gemini 3.6 Flash is now the default for the Antigravity agent, while Gemini 3.1 Pro remains the long-context flagship. The April 1 free-tier reshuffle removed free Pro-model access but kept Flash models available for prototyping.

Gemini 3.x (Current Production)

Google's latest generation is now GA and the recommended production tier:

  • Gemini 3.6 Flash ($1.50 input / $7.50 output) — Current Flash leader and default for the Antigravity managed agent.
  • Gemini 3.5 Flash-Lite ($0.15 input / $0.60 output) — Lowest-cost model supported by Managed Agents.
  • Gemini 3.1 Pro ($2.00 / $12.00 per 1M under 200K; $4.00 / $18.00 above) — Current flagship. Paid-only as of April 1, 2026. 2M-token context window.
  • Gemini 3 Pro ($2.00 input / $12.00 output) — Stable flagship alternative. Paid-only.
  • Gemini 3 Flash ($0.50 input / $3.00 output) — Earlier Flash generation. Retains free tier with reduced quota.
  • Gemini 3.1 Flash-Lite ($0.25 input / $1.50 output) — Cheapest Tier-1 budget model. Retains free tier with reduced quota.

Gemini 2.5 (Legacy — Paid Tier Only)

The 2.5 family remains available but is moving to legacy status. All 2.5 models are now paid-tier only.

  • Gemini 2.5 Pro ($1.25 input / $10.00 output) — Legacy flagship. Still competitive on price but superseded by 3.x.
  • Gemini 2.5 Flash ($0.30 input / $2.50 output) — Legacy mid-tier. Paid-only after April 1, 2026.

Free Tier: What Changed April 1, 2026

Google tightened the free tier significantly on April 1, 2026. The new state:

  • Pro-tier models (3.1 Pro, 3 Pro, 2.5 Pro): free tier removed. Paid-only.
  • Gemini 3 Flash: free tier retained with reduced daily quota.
  • Gemini 3.1 Flash-Lite: free tier retained with reduced daily quota.

Google still has the most generous free tier overall — Flash access at scale is unmatched — but the "free Pro" era is over. For flagship-class access without payment, your only remaining option is Claude's $5 trial credits.

Full breakdown: What changed April 1 and three ways to keep prototyping cheaply →

Context Caching: 75-90% Savings

Gemini's context caching is exceptionally aggressive. Gemini 2.5 Pro cached input drops from $1.25 to $0.13/M — a 90% discount. For applications with repeated system prompts or reference documents, this can slash your bill dramatically.

How Gemini Compares

At the flagship tier, Gemini 3.1 Pro ($2.00/M input) matches GPT-5.4 on price and undercuts Claude Opus 4.8 ($5.00/M) by 60%. Gemini still wins on context window — 2M tokens vs Opus 4.8's 1M and GPT-5.4's 270K — and on vision resolution.

At the budget tier, Gemini 3.1 Flash-Lite ($0.25/M input) is the cheapest mainstream model from a Tier-1 provider in 2026. Only DeepSeek's cached pricing is lower, at the cost of non-Tier-1 operational maturity.

For detailed head-to-head benchmarks and cost scenarios, see our flagship comparison guide.

Price History

Only models with a recorded price change are charted here.

Gemini 3.6 Flash

Price history tracking started April 2026. Flat model charts stay hidden until a price change is detected.
View pricing changelog →

Frequently asked questions

How much does Google WeatherNext cost?

Google has not published a WeatherNext per-token, per-run, or per-forecast model rate. WeatherNext code and notebooks are Apache 2.0, other repository materials are CC BY 4.0, and the Mini checkpoint can be tested on a free Colab v5e-1 runtime. Self-hosting and Google Cloud data access can still create compute, storage, query, and transfer charges.

How much do Gemini API Managed Agents cost?

Managed Agents use standard Gemini model rates for input, output, intermediate input, and reasoning tokens, plus normal tool fees. During preview, Google does not separately bill environment compute such as CPU, memory, or sandbox execution. Gemini 3.6 Flash is the default model, and free-tier projects can now experiment within Google's quotas.

How much does Gemini 3.6 Flash cost?

Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, including thinking tokens. Cached input costs $0.15 per 1M tokens, plus the published cache-storage charge.

How much does Gemini 3.1 Pro cost per token?

Gemini 3.1 Pro costs $2.00 per 1M input tokens and $12.00 per 1M output tokens for contexts up to 200K. Above 200K context, pricing jumps to $4.00 input / $18.00 output. Cached input drops to $0.20 per 1M — a 90% discount. Gemini 3.1 Pro is priced head-to-head with GPT-5.4.

Does Google Gemini have a free tier?

Yes, for eligible Flash models and now for Managed Agents on projects without active billing. Free-tier quotas vary, and free-tier data may be used to improve Google products. Gemini Pro models remain paid-only.

What changed on April 1, 2026?

Google removed all Pro-tier models from the free API tier — Gemini 3.1 Pro, Gemini 3 Pro, and Gemini 2.5 Pro now require billing to be enabled. Flash and Flash-Lite remain free but with tightened daily quotas. Google did not publish a formal changelog; the change surfaced through API error messages and updated pricing documentation.

How does Gemini compare to OpenAI and Claude?

At the flagship tier, Gemini 3.1 Pro ($2.00/$12.00) matches GPT-5.4 on pricing and undercuts Claude Opus 4.8 ($5.00/$25.00) by roughly 60%. Gemini still wins on context window (2M tokens vs 1M/270K) and vision resolution. For budget volume, Flash-Lite at $0.25/$1.50 is the cheapest Tier-1 budget model on the market.

What Gemini models does Google offer?

Google offers Gemini 3.1 Pro (flagship, tiered pricing), Gemini 3.1 Flash-Lite (budget), Gemini 3 Pro, Gemini 3 Flash (new default), and legacy Gemini 2.5 Pro / Flash / Flash-Lite models. The 3.x family is now the recommended production tier; 2.5 remains available but is moving to legacy status.

What is Gemini 3.1 Pro's context window?

Gemini 3.1 Pro supports a 2-million-token context window — the largest in production among Tier-1 providers. This is 2x Claude Opus 4.8's 1M-token Claude API window and ~7x GPT-5.4's 270K window. Pricing tiers at 200K: $2/$12 per 1M under, $4/$18 per 1M over.

How much does Gemini context caching save?

Google offers 90% discounts on cached context. Gemini 3.1 Pro drops from $2.00 to $0.20 per 1M input tokens on cache hits. For applications with large system prompts or reference documents, caching can reduce total input cost by 80-90%.

Methodology

Pricing sourced from https://ai.google.dev/pricing on . All prices in USD per 1 million tokens. Raw data: /api/pricing.json. API docs.

Compare All Providers

See how Google compares to OpenAI, Anthropic, DeepSeek, and more.