Google DeepMind targets enterprise AI costs with three new Gemini models
On 21 July 2026, Google DeepMind released three new artificial intelligence models under its Gemini family: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Rather than pursuing headline-grabbing benchmark records, this release is a deliberate commercial pivot. Google is targeting the single biggest obstacle to widespread enterprise adoption of autonomous AI systems: the runaway cost of running persistent, multi-step AI agents at scale. For professional services firms, including those operating in environmental consulting, engineering, law, and planning, this development is more consequential than it may first appear.
The release also came with a notable absence. Gemini 3.5 Pro, which had been teased publicly in May 2026 for a June rollout, did not ship. According to Bloomberg, the model has been pushed back by several months due to internal difficulties meeting performance targets on complex coding benchmarks. Google AI Studio Product Lead Logan Kilpatrick addressed the delay publicly on X, stating that Gemini 3.5 Pro is “currently testing with partners” and will be made broadly available “as soon as it’s ready.” Separately, Kilpatrick confirmed that pre-training for the next-generation frontier model, Gemini 4, is now officially underway, describing it as Google’s “most ambitious pre-training run yet.”
For professionals and business leaders who have been watching agentic AI with interest but held back due to cost concerns, this release changes the economic equation in a material way. The combination of architectural efficiency improvements and direct price reductions has produced claimed real-world cost reductions of between 30 and 65 per cent for long-horizon enterprise workflows. That is not a marginal efficiency gain. It is the kind of shift that moves AI deployment from a pilot-stage experiment to a commercially defensible operational tool.
Key details of the Gemini 3.6 Flash release and token economics
Gemini 3.6 Flash is the centrepiece of the July 2026 release and is positioned as the primary workhorse for autonomous agent applications. It is priced at USD 1.50 per million input tokens and USD 7.50 per million output tokens. The output price represents a reduction from USD 9.00 per million tokens on the previous Gemini 3.5 Flash model. The model carries a 1-million-token context window and supports a maximum output of 64,000 tokens per call. These specifications make it well suited to tasks that require processing large bodies of text, such as lengthy regulatory documents, technical reports, or complex multi-document datasets, before generating structured outputs.
The performance gains on autonomous task benchmarks are substantial. On the DeepSWE benchmark, which evaluates a model’s capacity to autonomously complete software engineering tasks, Gemini 3.6 Flash scored 49 per cent, up from 37 per cent on its predecessor. On MLE-Bench, which tests performance on machine learning engineering tasks completed end-to-end without human intervention, the model scored 63.9 per cent, compared with 49.7 per cent on Gemini 3.5 Flash. These are not incremental improvements. They indicate a measurably stronger capacity to complete complex, multi-step workflows independently, which is precisely the capability profile needed for autonomous agent deployments in professional services environments.
The second model in the release, Gemini 3.5 Flash-Lite, is designed for applications requiring extremely high throughput at minimum cost. Priced at USD 0.30 per million input tokens and USD 2.50 per million output tokens, it processes text at approximately 350 tokens per second, which is roughly twice the speed of its predecessor. Its coding performance is reported to closely match the standard Gemini 3 Flash despite the significant price reduction. This model is aimed at high-volume, lower-complexity tasks where speed and cost per call matter more than deep reasoning capability.
The third model, Gemini 3.5 Flash Cyber, is a specialist release paired with Google’s CodeMender agent and fine-tuned specifically for identifying software vulnerabilities and generating patches. It scored 83.2 per cent on the Cybergym benchmark, which evaluates cybersecurity task performance. Because of the clear dual-use risk associated with a fast, low-cost tool capable of identifying software exploits, Google has restricted access to governments and vetted cybersecurity partners only. This access restriction reflects a deliberate governance decision on Google’s part and sets a precedent for how specialised, high-risk AI capabilities may be distributed in future releases. Beyond the price cuts, Google reports that the underlying model architecture was redesigned to reduce output verbosity by approximately 17 per cent on average, meaning the models reach correct outputs using fewer reasoning steps and tool calls. When this architectural efficiency is combined with the 17 per cent reduction in output token pricing, the compounding effect produces the 30 to 65 per cent cost reduction cited for long-horizon agentic workflows.

Australian context: implications for professional services and enterprise AI adoption
Australia’s professional services sector has been cautious in scaling agentic AI deployments, and for understandable reasons. Firms operating under regulatory obligations, whether in environmental consulting, town planning, legal practice, or financial services, carry professional liability that makes unchecked AI output a serious risk. The cost barrier has, in some ways, served as a natural brake on premature or poorly governed deployment. What the Gemini 3.6 Flash release does is shift that conversation. When the operational cost of running a persistent AI agent drops by up to 65 per cent, the business case for deployment strengthens considerably, and firms that have been waiting on the economics now face a more pressing question: not whether agentic AI is affordable, but whether their governance frameworks are ready to support it responsibly.
References and related sources
- Primary source: venturebeat.com
- codingsalt.com
- pureai.com
- neomanex.com
- kie.ai
How iEnvi can help
iEnvi integrates technology and data-driven approaches into environmental consulting. We monitor AI and technology developments that affect how environmental professionals deliver services to clients.
This is an iEnvi Machete news summary. Prepared by iEnvi to summarise the source article for environmental professionals tracking AI, data, and technology developments that affect consulting and project delivery.
Published: 26 Jul 2026
Need advice on this topic? Speak to an iEnvi expert at info@ienvi.com.au or 1300 043 684, or contact us online.
Need advice on this issue? iEnvi provides practical, senior-led environmental consulting across contaminated land, remediation, ecology and environmental risk.
Contaminated land advice Remediation services Discuss your site Talk to iEnvi