Google has released Gemini 3.6 Flash and 3.5 Flash-Lite as new workhorses designed to cut latency and token costs for enterprise AI agents.
The economics of running autonomous software agents inside a production environment come down to a fixed equation few vendors advertise directly. A model needs to reason through a multi-step task competently, but the cost of tokens required to execute that model is often prohibitively high due to factors such as overhead and variability in system performance.
By introducing Gemini 3.6 Flash and 3.5 Flash-Lite, Google claims its enterprise AI agents can run at lower latency costs while still maintaining robust capabilities. These new flash-based architectures are designed to reduce the cost of token usage by optimizing data processing and reducing memory requirements. As a result, they offer significant benefits for organizations looking to deploy more complex AI models in their production environments without breaking the bank.