- Google just launched three new AI models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. They're all built for speed and cheap operation in AI agent workflows.
- Gemini 3.6 Flash is now the main model in the Gemini app. It replaces 3.5 Flash, which just launched and is already deprecated.
- The pricing is tiered. Gemini 3.5 Flash-Lite is the cheapest at $0.30 per million input tokens, while Gemini 3.6 Flash costs $1.50 per million.
Google's AI strategy feels less like a roadmap and more like a game of whack-a-mole. The company just announced another three Gemini models, and if you're confused, you should be. This isn't about building a better chatbot for you to talk to. It's a naked attempt to win over developers building automated systems, or "AI agents," by competing on price and speed. But the real story here is the corpse left on the battlefield: the Gemini 3.5 Flash model they touted at I/O a few months ago is already dead. That tells you everything about the frantic, unstable pace of this race.
Meet the New Gemini Models
Forget one model to rule them all. Google is now pitching a toolbox. They're releasing three separate models, each with a specific, narrow job. The idea is simple: if you're building an automated customer service bot that needs to process a million tickets, you don't need the smartest model. You need the fastest and cheapest one. That's what this launch is about.
Gemini 3.6 Flash: The New Default
Gemini 3.6 Flash is your new middle-of-the-road option. Google says it's meant to "balance speed with intelligence for agentic and multimodal tasks." It's now generally available in their API. More importantly, it's immediately taking the place of Gemini 3.5 Flash as the main model inside the consumer Gemini app. Think about that. A model they launched with fanfare in May is being swapped out before summer even ends. It's a brutal upgrade cycle that makes it risky to build anything on this foundation.
Gemini 3.5 Flash-Lite: The Cheap Option
Then there's Gemini 3.5 Flash-Lite. This is the budget special. Google bills it as "the fastest, lowest-cost 3.5 model for high-throughput execution." If your AI agent just needs to sort data, tag content, or handle other simple, repetitive jobs at massive scale, this is the model they want you to use. Its entire reason for existing is to undercut competitors on price.
Gemini 3.5 Flash Cyber: The Security Specialist
The third model is Gemini 3.5 Flash Cyber. It's a niche product trained specifically for cybersecurity work, like summarizing threat reports or writing detection rules. It's Google's play for a slice of the lucrative AI security market, going up against tools like Microsoft's Security Copilot. But we haven't seen independent tests yet, so its actual usefulness is still a promise.
Performance, Pricing, and the Deprecated Predecessor
Let's get to the numbers, because that's where Google hopes to win. They're not claiming these models beat GPT-4o on reasoning tests. They're claiming they're cheaper to run. Here's the breakdown straight from the API docs.
| Model | API Name | Input Token Cost (per 1M) | Output Token Cost (per 1M) | Primary Use Case |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | gemini-3.5-flash-lite | $0.30 | $2.50 | High-throughput, low-cost tasks |
| Gemini 3.6 Flash | gemini-3.6-flash | $1.50 | $7.50 | Agentic & multimodal workflows |
Notice they don't list a price for the Cyber model yet. But the big news isn't the new prices, it's the old model. As Ars Technica reported, "Gemini 3.5 Flash, which was the star of the show at I/O, has already been deprecated." Google says they updated it based on feedback, but that's a wild pace. It means any developer who built something on 3.5 Flash just got a forced migration. That's a great way to erode trust.
The AI Agent Workflow Focus
You keep hearing "agentic workflows." What does that actually mean? Picture an AI that doesn't just answer a question, but completes a job. It's the difference between an AI that tells you your refund policy and an AI that actually finds your order, checks it's eligible, and starts the return process for you. That chain of actions is an agent.
Key Features for Agents
Google built specific features into Gemini 3.6 Flash for this kind of work:
- Code Generation: Agents often need to write little bits of code to talk to other software.
- Spatial/Multimodal Reasoning: They need to understand charts, images, or diagrams inside documents.
- Multi-step Reasoning: This is the core skill: breaking a big request like "analyze this report" into a checklist of smaller tasks.
The focus here is purely practical. It's not about a flashy demo. It's about being a reliable, affordable cog in a bigger machine.
The Elephant in the Room: Gemini 3.5 Pro and Product Confusion
Hold on. Where's Gemini 3.5 Pro? You know, the more powerful model that's supposed to handle complex reasoning? Good question. Ars Technica pointed out that "none of the new models is the delayed Gemini 3.5 Pro, which was supposed to launch in June." So Google's top-tier model is still missing.
This creates a weird hole in their lineup. And launching three models at once doesn't clear things up. It adds to the noise. A comment on Hacker News put it bluntly, arguing Google often just creates new products instead of fixing core ones. An Instagram post about the launch even suggested that dropping multiple models "isn't always a sign of leadership: it can be a sign that none of them are quite ready." Ouch.
What This Means for India
For developers and companies in India, this launch is about access and cost. These models should be available through the Gemini API in India, but Google has been known to change access rules without much warning.
Pricing and Developer Impact
The cost is the main draw. At $0.30 per million input tokens, Gemini 3.5 Flash-Lite is dirt cheap. For Indian startups watching every rupee, that low barrier could fuel a lot of experiments in automation for customer service or data processing.
Language Support and Local Alternatives
Here's the catch: Google hasn't said if these new models are any better at Indian languages like Hindi or Tamil. Their performance is an open question until someone tests it. Developers should absolutely run their own tests and remember there are local options like Krutrim, which are built from the ground up for Indian languages.
Another unanswered question is on-device processing. Can these models run on a phone, or do they need the cloud? For applications dealing with sensitive data in India, that's a major privacy question. Google isn't saying, while Apple is pushing its on-device "Apple Intelligence" as a selling point.
Frequently Asked Questions
Is Gemini 3.6 Flash available in India?
It should be, as it's listed as generally available in the API. But always double-check Google's official documentation for the latest regional info.
Which model is the cheapest for simple tasks?
Hands down, it's Gemini 3.5 Flash-Lite at $0.30 per million input tokens.
What happened to Gemini 3.5 Flash?
Google deprecated it. The model they launched in May 2026 has already been replaced by 3.6 Flash.
Is Gemini 3.5 Pro available now?
No. It's still delayed and didn't launch with this batch.
The Bottom Line
Look, this isn't a breakthrough. It's a price war. Google is throwing cheaper, faster models at developers to stop them from going to OpenAI or Anthropic. But killing a flagship model within months shows a shocking lack of product stability. If you're building, the tiered pricing might save you money today. Just don't be surprised if the model you pick is gone by tomorrow. The real signal isn't in these three new names, it's in the one that's still missing: Gemini 3.5 Pro. Its continued absence is the clearest sign that Google is scrambling.
Sources
- blog.google
- arstechnica.com
- ai.google.dev
- 9to5google.com
- news.ycombinator.com
- facebook.com
- instagram.com