Google has released two new flash-based models for its Gemini AI platform, designed to cut latency and token costs for enterprise AI agents. The Gemini 3.6 Flash and 3.5 Flash-Lite are the latest additions to Google's Gemini lineup, which is specifically tailored for running autonomous software agents in production environments.
As a result of these new flash-based models, companies can expect to save significant amounts on their token costs. The economics of working with AI agents come down to a fixed equation that few vendors openly advertise directly. To perform complex tasks competently, an agent needs to reason through multiple steps, but the cost of this task is typically high due to the complexity and computational requirements.
The introduction of Gemini 3.6 Flash and 3.5 Flash-Lite addresses these challenges by reducing latency and token costs for enterprise AI agents. By optimizing the architecture of their Gemini platform, Google can provide more efficient execution of complex tasks without increasing token costs. This not only reduces operational expenses but also opens up new possibilities for businesses looking to leverage AI in their production environments.