All posts

Google ships 3 new Gemini models, still no 3.5 Pro

Huma ShaziaJuly 21, 2026 at 11:01 PM5 min read

Key Takeaways

  • Gemini 3.6 Flash cuts token usage by 17% compared to 3.5 Flash, lowering costs for production workloads
  • Google's flagship Gemini 3.5 Pro remains delayed after internal performance issues reported by Bloomberg
  • A new cybersecurity model, 3.5 Flash Cyber, will only be available to governments and select partners

Google DeepMind released three new Gemini models on Tuesday: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The company pitched these as efficiency-focused releases for teams building AI agents at scale. What Google didn't ship is more interesting. The long-awaited Gemini 3.5 Pro, teased in May and originally promised for June, is still missing.

This matters if you're choosing an AI provider. OpenAI has released GPT-5.5 and started rolling out GPT-5.6. Anthropic has launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to its Fable 5 model. Google's flagship reasoning model hasn't been updated since February.

Advertisements

What did Google actually release?

Gemini 3.6 Flash is the headline model. Google calls it their "workhorse" and claims it improves coding, knowledge work, and multimodal performance while using up to 17% fewer tokens than 3.5 Flash. Fewer tokens means lower API bills. For startups running high-volume inference workloads, that's real money.

Gemini 3.5 Flash-Lite is the budget option. Google positions it as the most cost-effective model in the Flash family, targeting use cases where latency and price matter more than peak capability.

The third release, Gemini 3.5 Flash Cyber, is a specialized model fine-tuned for finding and fixing security vulnerabilities. It won't be available to most developers. Google is limiting access to governments and "trusted partners" through a pilot program. If you're building cybersecurity tooling, you'll need to apply.

Why is Gemini Pro delayed?

Google teased 3.5 Pro in May, saying it was "already being used internally" and would roll out "next month." That didn't happen. Last week, Bloomberg reported Google was struggling to meet internal performance goals for the model.

Gemini Pro models are Google's top-tier offerings for complex reasoning and coding tasks. Flash models trade some capability for speed and cost. The distinction matters because startups building agents that need to plan, reason through multi-step problems, or handle sophisticated code generation typically need Pro-class models.

Google DeepMind product lead Logan Kilpatrick said the company is testing 3.5 Pro with partners and hopes to "land soon." He also mentioned the team has started pre-training for Gemini 4, which he called their "most ambitious" run yet. That's forward-looking, but it doesn't solve the current gap.

Advertisements

How does this affect model selection for startups?

If you're building production AI features right now, you have a clearer picture of the tradeoffs. For high-volume, latency-sensitive workloads like chatbots, summarization, or real-time suggestions, Gemini 3.6 Flash looks competitive. The 17% token reduction helps unit economics.

For complex reasoning, code generation, or agentic workflows that require planning, the picture is murkier. OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.8 are available now. Gemini 3.5 Pro is not. If your product depends on frontier reasoning capability, you're choosing between OpenAI and Anthropic until Google delivers.

The cybersecurity model is a niche play. Government contracts and enterprise security are lucrative, but the restricted access means most startups won't touch it. If you're building security tooling, it's worth applying to the pilot. Everyone else can ignore it.

The competitive context

The AI model market has accelerated. Anthropic and OpenAI are shipping major releases every few months. Google's February-to-July gap on Pro is starting to look like a competitive liability. The company still has advantages: deep integration with Google Cloud, strong enterprise relationships, and pricing that undercuts OpenAI on many workloads.

But perception matters. Developers building today tend to default to whatever's newest and highest-capability. Every month without a Pro update is a month where new projects start on GPT-5.5 or Claude.

ℹ️

Logicity's Take

Google's Flash releases are solid incremental improvements, but the Pro delay is a strategic problem. Startups evaluating AI providers care about two things: capability and roadmap confidence. Google is delivering on cost efficiency while falling behind on frontier capability. If you're locked into Google Cloud already, 3.6 Flash is a reasonable choice for production workloads. If you're starting fresh and need the best reasoning model available today, OpenAI and Anthropic have the edge. Watch for 3.5 Pro in the next 4-6 weeks. If it slips again, that tells you something about Google's execution.

Frequently Asked Questions

What is the difference between Gemini Flash and Gemini Pro?

Flash models prioritize speed and cost efficiency for high-volume production workloads. Pro models offer higher capability for complex reasoning, coding, and multi-step tasks but cost more per token.

When will Gemini 3.5 Pro be released?

Google has not given a specific date. Product lead Logan Kilpatrick said it's being tested with partners and will "land soon." The model was originally expected in June 2026.

How much cheaper is Gemini 3.6 Flash than previous models?

Google claims 3.6 Flash uses up to 17% fewer tokens than 3.5 Flash for equivalent tasks, which directly reduces API costs.

Can startups access Gemini 3.5 Flash Cyber?

Not currently. Google is limiting access to governments and trusted partners through a pilot program. There's no timeline for broader availability.

Is Google falling behind OpenAI and Anthropic?

On flagship model releases, yes. OpenAI has shipped GPT-5.5 and is rolling out 5.6. Anthropic has released Claude Opus 4.8 and Sonnet 5. Google's Pro model hasn't been updated since February.

Also Read
Augustus hits $1B valuation with OCC banking approval

Another major enterprise AI development from this week

ℹ️

Need Help Implementing This?

Choosing between AI providers for your startup? Logicity's consulting team helps founders evaluate model performance, API costs, and integration complexity. Reach out at consulting@logicity.in.

Source: Enterprise News | TechCrunch / Rebecca Bellan

H

Huma Shazia

Senior AI & Tech Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.