Google’s Gemini 3.6 Flash cuts token usage by up to 17%, but the sharper signal in Google’s new Gemini models release is what still didn’t arrive: Gemini 3.5 Pro.

Missing Gemini 3.5 Pro Overshadows New Gemini Models
XOOMAR Intelligence
Analyst Take
Google DeepMind released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber on Tuesday, according to TechCrunch. The company is pitching the trio around efficiency, latency, and reliability for customers building AI agents at scale.
That is a useful launch. It is also a carefully scoped one. Google added cheaper and more specialized Flash models, while leaving its long-anticipated Pro update off the table again.
Google ships three new Gemini models, but Gemini 3.5 Pro is still absent
Gemini 3.6 Flash is Google’s “workhorse model,” aimed at coding, knowledge work, and multimodal tasks. The key commercial hook is lower token usage, up to 17% less than its predecessor, 3.5 Flash, which Google says makes it cheaper to run.
Gemini 3.5 Flash-Lite is positioned as the most cost-effective model in the class. That makes it the obvious option for high-volume applications where cost and response time matter more than maximum reasoning depth.
Gemini 3.5 Flash Cyber is the specialized release. Google says it was fine-tuned for finding and fixing cybersecurity vulnerabilities at a decent price point. It won’t be broadly available at launch.
Gemini 3.5 Flash Cyber will be available exclusively to governments and trusted partners as part of a limited access pilot program, according to Google.
The product split is clear: Flash models are built for speed, cost control, and production deployment. Pro models generally carry Google’s highest-capability push for complex reasoning and coding tasks.
That makes the missing Gemini 3.5 Pro hard to ignore. Google’s update widens the Gemini menu, but it does not answer the central question developers and AI buyers have been watching: when Google will ship its next flagship Pro model.
For readers tracking Google’s broader AI push, this launch sits alongside XOOMAR’s recent coverage of 6 to 10x Google Gemini Chip Jolts Alphabet AI Bets and Google Opens Android App Stores but Keeps the Cash, two separate fronts in how Google is trying to defend and extend its platform position.
New Gemini Flash models emphasize faster tools and tighter security access
Google says the focus of these releases is efficiency, latency, and reliability for customers building AI agents at scale. That language matters because agents are only useful in production if they can respond quickly, run repeatedly without runaway costs, and behave consistently enough for developers to trust them inside workflows.
The three models map to three different needs:
| Model | Google’s stated positioning | Likely buyer focus based on the release |
|---|---|---|
| Gemini 3.6 Flash | “Workhorse model” with better coding, knowledge work, and multimodal performance | General production apps needing faster, cheaper responses |
| Gemini 3.5 Flash-Lite | Most cost-effective model in the class | Lightweight or high-volume AI tasks |
| Gemini 3.5 Flash Cyber | Fine-tuned to find and fix cybersecurity vulnerabilities | Governments and trusted partners in the pilot |
The cyber model is the most restricted and the most specialized. Google did not describe it as a general release, and the limited-access pilot keeps its actual performance, pricing details, and deployment rules out of public view for now.
That restraint matters. A model tuned to find and fix vulnerabilities can be valuable for defensive work, but Google is not opening the door to everyone yet. The company is choosing a controlled rollout through governments and trusted partners.
XOOMAR analysis: the Flash lineup now looks more like a segmented product stack than a single model family update. Google is not only chasing capability. It is carving Gemini into models for cost-sensitive use, general production use, and restricted security work.
The unanswered part is whether this segmentation can offset the lack of a new Pro model. Faster, cheaper models help developers ship. They don’t replace the benchmark and credibility role that flagship models usually play.
Gemini 3.5 Pro delay keeps pressure on Google’s flagship AI roadmap
The most important absence remains Gemini 3.5 Pro. TechCrunch reports that Gemini Pro was last updated in February, and Google had teased the Pro version during the 3.5 Flash release in May.
Google said then that the Pro version was:
“already being used internally, and we look forward to rolling it out next month.”
That timing did not hold. Last week, Bloomberg reported that Google was facing internal delays launching 3.5 Pro as it struggled to meet internal performance goals, according to the source material.
The competitive backdrop is moving quickly. Since Google’s last Pro update, OpenAI has released GPT-5.5 and begun rolling out GPT-5.6, while Anthropic has launched Claude Opus 4.8 and Claude Sonnet 5 and expanded access to its frontier Fable 5 model.
Google DeepMind product lead Logan Kilpatrick said Tuesday that the company is currently testing Gemini 3.5 Pro with partners and hopes to “land soon.” He also said the team has started its most ambitious pre-training run yet for Gemini 4.
That gives Google two messages at once. One is immediate and practical: Flash models are getting cheaper and more specialized. The other is unfinished: the Pro roadmap is still pending, and Gemini 4 is already being discussed before Gemini 3.5 Pro has shipped publicly.
XOOMAR analysis: this is the tension Google cannot package away. The new Gemini models strengthen the middle of the lineup, but the missing flagship keeps the story centered on whether Pro is delayed, being held for quality, or being repositioned inside the broader Gemini roadmap.
Developers now get to test Google’s efficiency claims in real workloads
The next phase moves from launch claims to production tests. Developers will measure Gemini 3.6 Flash and 3.5 Flash-Lite on speed, cost, coding quality, multimodal output, and reliability under repeated use.
For Gemini 3.5 Flash Cyber, the public signal will be more limited because access is restricted. The key questions are availability, pilot scope, pricing, and what Google shares about the model’s performance in finding and fixing vulnerabilities.
Several details are still not clear from the launch material:
- Pricing: Google says 3.6 Flash is cheaper than 3.5 Flash because of lower token usage, but the supplied material does not provide full pricing tiers.
- Access: Flash Cyber is limited to governments and trusted partners, with no broader availability timeline supplied.
- Benchmarks: Google cites improved capabilities, but the source material does not include public benchmark tables.
- Pro timing: Kilpatrick said Gemini 3.5 Pro hopes to “land soon,” but no release date was given.
The new Gemini models give Google fresh product momentum where developers care about cost and latency. The sharper market signal will come when Google either ships Gemini 3.5 Pro or explains why its flagship roadmap is taking longer than expected.
The Bottom Line
- Google is emphasizing cheaper, faster AI models for developers building agents at scale.
- The absence of Gemini 3.5 Pro leaves unanswered questions about Google’s top-end AI roadmap.
- Limited access to Gemini 3.5 Flash Cyber shows cybersecurity AI remains a controlled deployment area.
New Gemini Models
| Model | Positioning | Key Detail | Availability |
|---|---|---|---|
| Gemini 3.6 Flash | Workhorse model for coding, knowledge work, and multimodal tasks | Uses up to 17% fewer tokens than Gemini 3.5 Flash | Released |
| Gemini 3.5 Flash-Lite | Cost-effective model for high-volume applications | Prioritizes cost and response time over maximum reasoning depth | Released |
| Gemini 3.5 Flash Cyber | Specialized cybersecurity model | Fine-tuned for finding and fixing vulnerabilities | Limited access pilot for governments and trusted partners |
| Gemini 3.5 Pro | Expected higher-capability Pro model | Still absent from the release | Not released |
Gemini 3.6 Flash Token Usage Reduction vs 3.5 Flash
Written by
XOOMAR Insights Team
Research and Editorial Desk
The XOOMAR Insights Team pairs automated research with human editorial judgment. We track hundreds of sources across technology, fintech, trading, SaaS, and cybersecurity, cross-check the facts, and explain what happened, why it matters, and what to watch next. We do not just rewrite headlines. Every article is fact-checked and scored for reliability before it goes live, and we link back to the original sources so you can verify anything yourself.
Explore More Topics
Related Articles
Technology6 to 10x Google Gemini Chip Jolts Alphabet AI Bets
Alphabet jumped 3% as Frozen v2 promised a 6 to 10x efficiency answer to Google's AI cost problem, but the chip may not arrive until 2028.
Technology$6,880 Vertu Alphafold Stumbles on AI Agent Promise
Vertu’s $6,880 Alphafold sells executive AI status, but early testing shows it still isn’t trustworthy enough for the job.
TechnologyPerfect AI Agent Evaluation Masks Enterprise Breakage
A clean agent transcript can hide broken user patterns. LangChain, Conviva and CoreWeave say cohort-level evaluation is the real test.
TechnologyApproval Fail Sinks ChatGPT Work, Claude Cowork Wins
ChatGPT Work organized 447 PDFs but ignored approval controls. Claude Cowork looks safer for file agents right now.
TechnologyMeta Warns AI Agents Infrastructure May Crack in 20 Months
Meta says companies may have just 20 months to rebuild infrastructure before AI agents become the dominant system users.
SaaS & ToolsJack Dorsey Buzz Hits Slack With Open AI Agent Chat
Buzz brings AI agents into team chat, giving Jack Dorsey a direct shot at Slack, GitHub, and the next workplace control layer.
Cybersecurity17,000 AI Agent Actions Crack Open Hugging Face Breach
Hugging Face says an autonomous AI agent breached production systems, stole some credentials, and triggered an AI-assisted defense.
CybersecurityBanks Brace for Gold Eagle AI Cybersecurity Pressure
Gold Eagle is voluntary, but banks may feel pressure to use its federal vulnerability intelligence before examiners start asking.
TechnologySubstack AI Detector Turns Every Writer Into a Suspect
Substack's AI detector gives readers a scan button for posts and comments, putting writer trust under a microscope.
Global TrendsStates Freeze Paramount Warner Merger After DOJ Approval
The Paramount Warner merger cleared DOJ review, then hit a state-led court freeze. CFOs can't treat federal approval as the finish line.
Don't miss the signal
Get our weekly roundup of the stories that matter across tech, fintech, and trading. No noise, just signal.
Free forever. No spam. Unsubscribe anytime.