Google launches Gemini 3.6 Flash and 3.5 Flash-Lite for scalable AI agents
Published
Google has launched Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two lower-cost models designed for production AI agents, coding, document processing a...

Google has launched Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two lower-cost models designed for production AI agents, coding, document processing and high-volume multimodal workloads.
Gemini 3.6 Flash targets more efficient agentic work
Gemini 3.6 Flash is positioned as Google’s new workhorse model. The company says it improves coding, knowledge work and multimodal performance while using fewer output tokens and fewer reasoning steps than Gemini 3.5 Flash.
Google reports that the model uses 17 percent fewer output tokens on the Artificial Analysis Index and can reduce token use by as much as 65 percent on selected software-engineering evaluations. These figures come from Google and benchmark providers and may not reflect every production workload.
Pricing and capabilities
Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens. Google says the model delivers stronger results in coding, computer use, document analysis and multi-step agent workflows while lowering the cost per completed task.
Computer use is now available as a built-in client-side tool through the Gemini API and Gemini Enterprise, allowing the model to interact with software interfaces as part of supervised workflows.
Flash-Lite is built for high-throughput workloads
Gemini 3.5 Flash-Lite is the fastest and least expensive model in the 3.5 family. Google lists pricing of $0.30 per million input tokens and $2.50 per million output tokens, with reported generation speed of 350 output tokens per second.
The model is intended for agentic search, document processing, translation, summarisation and other workloads where latency and volume are more important than maximum reasoning depth.
A specialised cyber model joins the lineup
Google also introduced Gemini 3.5 Flash Cyber, a specialised model integrated with the CodeMender security agent. It is designed to find, validate and patch software vulnerabilities, but will initially be available only to governments and trusted partners through a limited-access pilot.
Availability
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available through the Gemini API in Google AI Studio and Android Studio, as well as through Google’s enterprise platforms. Gemini 3.6 Flash is also available in Google Antigravity, while Flash-Lite is rolling out in Google Search.
What remains pending
Google says Gemini 3.5 Pro is still being tested with partners and will be released broadly when it is ready. The company also confirmed that pre-training has started on Gemini 4.
Practical significance
The release shows that model competition is moving beyond benchmark leadership toward the economics of running agents at scale. For developers, the main question is increasingly how much useful work a model can complete per token, per second and per dollar.
Google announced the models in an official product post .
Source: Google