Gemini 3.5 Flash API Pricing & Cost Calculator
Gemini 3.5 Flash is one of the most cost-effective fast models on the market. It is ideal for high-volume multimodal analysis and low-latency completions.
Gemini 3.5 Flash is Google's ultra-fast multimodal model, designed for extreme speed and high scalability.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 30,000 input tokens (50% cached), 3,000 output tokens, and 1,000 requests.
How this is calculated
Gemini 3.5 Flash is priced at $1.50 per million input tokens and $9.00 per million output tokens, supporting a 75% caching discount ($0.38 per million) and a 50% batch discount ($0.75 per million inputs, $4.50 per million outputs).
Verdict
Base pricing sits at $1.50/M input and $9/M output, with a 75% cache discount landing cached input at $0.38/M and a 50% batch discount dropping both rows to $0.75/M input and $4.50/M output. The cache discount is shallower than Anthropic's 90% but the lower base output rate compensates, and Flash's native multimodal support is what closes the case: audio, image, and video inputs hurtle through a single endpoint at the same per-token price, so classification and transcription pipelines that would otherwise need multiple models collapse into one bill.
More API Standalones scenarios
Related guides
Frequently asked questions
Does Gemini 3.5 Flash support prompt caching?
How much does Gemini 3.5 Flash cost per million output tokens?
Is Gemini 3.5 Flash the cheapest multimodal model?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with syntax highlighting.
Use tool ➜