Back
Google DeepMindModelGoogle DeepMind2026-08-11

Gemini 2.5 Pro Gets 30% Price Cut and Doubled Rate Limits

Google DeepMind has cut Gemini 2.5 Pro API pricing by 30% across input and output tokens and doubled rate limits for paid tiers, effective immediately on both Google AI Studio and Vertex AI.

Original source

Google DeepMind announced a 30% reduction in Gemini 2.5 Pro pricing across both input and output tokens, alongside doubled rate limits for paid-tier users. The changes take effect immediately on Google AI Studio and Vertex AI, requiring no migration or configuration changes for existing users.

The pricing reduction follows a broader industry pattern of frontier model costs declining as inference infrastructure matures and competition intensifies. For developers running production workloads on Gemini 2.5 Pro, the combined effect of lower per-token costs and higher throughput ceilings meaningfully changes the unit economics of applications that were previously constrained by either cost or rate limits.

Gemini 2.5 Pro sits in the premium tier of Google's model lineup, positioned for complex reasoning, long-context tasks, and multimodal workloads. The rate limit doubling is particularly relevant for teams building agentic pipelines or batch-processing workflows that were hitting ceilings at previous limits. Both changes apply across the board — no new pricing tiers, no tiered rollout, no waitlist.

Panel Takes

The Builder

The Builder

Developer Perspective

A 30% price cut with no API changes, no migration, no new environment variables — this is exactly how you improve a platform. The rate limit doubling is the more interesting number for anyone building pipelines that were throttling; cost-per-token matters less than throughput when you're waiting on retries. The DX bet here is doing nothing visible and making everything cheaper, which is the correct bet.

The Founder

The Founder

Business & Market

This is margin compression as competitive strategy — Google can absorb price cuts that would gut a startup building on top of a single model API, which is exactly the point. The doubling of rate limits signals they want production workloads, not just prototype traffic, and that's the revenue-quality shift that matters for their cloud business. Any wrapper company that built a pricing model assuming today's input token costs is now either repricing or watching their margin argument evaporate.

The Skeptic

The Skeptic

Reality Check

A price cut is real, but let's be precise about what's driving it: inference costs are dropping industry-wide, and this is Google keeping pace with Anthropic and OpenAI rather than pulling ahead. The rate limit doubling is the more meaningful signal — it suggests previous limits were a genuine constraint on adoption, not a safety floor, and removing them is an admission, not a gift. What kills this in 12 months isn't competition; it's whether Gemini 2.5 Pro's reasoning quality justifies the premium tier against models that will keep getting cheaper across the board.

The Futurist

The Futurist

Big Picture

The thesis here is that frontier model pricing will converge toward commodity within 18 months, and Google is betting they win on distribution and infrastructure — Vertex AI integration, enterprise contracts, the full GCP stack — not on model differentiation alone. The second-order effect is that every product roadmap currently gated behind 'too expensive to run at scale' just had its unlock condition move closer. The trend this is riding is inference deflation, and Google is on-time, not early — but being on-time with the largest enterprise sales motion in the market is not a bad position.

Bookmarks

Loading bookmarks...

No bookmarks yet

Bookmark tools to save them for later