Google ships cheaper Gemini models, but Gemini 3.5 Pro is still stuck in testing
The new Flash lineup targets lower token costs for AI agents, while Gemini 3.5 Pro remains delayed and Gemini 4 starts pre-training.

Google announced three new Gemini models designed to be faster and cheaper to run AI agents, while also saying Gemini 3.5 Pro is still in testing. It also began a major pre-training run for Gemini 4, even as 3.5 Pro, initially promised for June, is not yet ready.
Google is rolling out new Gemini models meant for one thing: making AI agents cheaper to run, faster to iterate, and less likely to melt your token budget. But there is a catch that matters to anyone betting on Google’s next “frontier” jump: Gemini 3.5 Pro, initially promised for June, is still in testing.
That tension sets the agenda. On Tuesday, Google said it is launching 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, positioning them as a “sweet spot” between efficiency and power for agent workloads. Meanwhile, Google said Gemini 3.5 Pro is being tested with partners and will ship “as soon as it’s ready.” In other words, Google is shipping improvements now, but asking customers to wait for the bigger upgrade they originally penciled into early summer.
Let’s break down what’s actually arriving this week, because these details tell you what Google thinks the market will reward. First, 3.6 Flash is described as a “workhorse” that is better than 3.5 Flash at tasks such as coding. Google also claims 3.6 Flash reduces token usage by up to 17% versus its predecessor, based on analysis from Artificial Analysis Index, which benchmarks AI tools across different capabilities. Token usage is the practical billing unit in many AI deployments, so a 17% reduction is not a cosmetic change. It can translate into meaningfully lower inference costs or more work per dollar, depending on your architecture.
Second, Google unveiled 3.5 Flash-Lite, which it says is its fastest and “most cost-effective” version of its 3.5 model yet. That is a direct signal to builders who need high throughput and predictable latency more than maximal quality. Third, there is 3.5 Flash Cyber, designed to find and fix cybersecurity vulnerabilities with a performance level described as similar to rivals, but at a “lower price per token than larger models.” Tulsee Doshi, senior director of product management in Google’s Gemini group, said that in a blog post.
Why push these models now, when the headline model many people want is still not ready? Because the economics of AI usage are being reined in. The core reality is that using lots of tokens does not always deliver proportionally better outcomes, and more companies are already capping spend. Google CEO Sundar Pichai put a point on this earlier this year, saying companies are “already blowing through their annual token budgets, and it’s only May,” and arguing that a mix of Flash and other frontier models could save money. That context matters: Google is not just launching models, it is shaping procurement behavior by giving teams options that look good on both performance and cost.
And yes, there is competitive pressure in the cybersecurity lane. Anthropic and OpenAI have both rolled out cybersecurity models in recent months, so Google is moving into a category where buyers will ask: which tool does the job at the lowest effective cost? In this week’s lineup, Google is explicitly claiming a lower price per token than larger models for Flash Cyber. That is the kind of positioning that can win budget holders who are skeptical of “bigger is better” when budgets are already getting squeezed.
Meanwhile, the part investors and model strategists care about most is what’s not shipping yet. Gemini 3.5 Pro is still in testing with partners, with Google saying it will ship “as soon as it’s ready.” Business Insider previously reported that Google pushed the targeted release date for 3.5 Pro to July, and Bloomberg has recently reported the model may have been further delayed. Google also said as of Tuesday it has not placed any model among the top 10 ranking in the Artificial Analysis leaderboard, which uses a composite benchmark scoring models across areas including maths and reasoning. Even if you do not treat leaderboards as scripture, being outside the top 10 is a public signal that the “frontier moment” is not arriving on schedule.
Then there is Gemini 4, which Google is not just teasing, it is starting to build in earnest. Google said it has begun its “most ambitious pre-training run yet” for Gemini 4, its next major model. For context, Google typically launches these flagship models at the very end of the year, so Gemini 4 is likely not a quick fix for the months when 3.5 Pro is delayed. Still, the pre-training start is meaningful. It is how Google tells the market it is not idling. The question for decision-makers is whether the company’s interim strategy, led by cheaper Flash variants, can hold traction until the next major leap.
For executives at AI-first companies, this is a procurement and roadmap story as much as a product story. If your teams have been watching token spend more closely since budgets tightened, Google is making the case for a hybrid model strategy: use efficient Flash versions for most tasks, and reserve heavier frontier models for moments when quality truly moves the needle. The strategic stakes are simple. If Gemini 3.5 Pro slips again, customers and competitors will keep optimizing for cost and operational fit now, and Google will have to earn mindshare twice: first by saving money in production, and later by proving the frontier jump is worth the wait.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business

Anthropic’s Levant Alpöge cracks the Jacobian conjecture after 87 years
A Harvard valedictorian used Claude to hit a 1939 breakthrough, but the missing “why” is the real problem.

Uber buys Delivery Hero for nearly $15B, vaulting to top food delivery outside China
The deal doubles Uber's dual-services footprint and pushes a ride-and-eats bundling play into 50 more markets.

Epic and Google drop settlement bid, forcing rival Android app stores by July 22
Google told the court it is ready to carry third-party app stores starting Wednesday, July 22.

