TechCompare LogoTechCompare

Gemini 3.5 Flash API Pricing & Cost Calculator

Gemini 3.5 Flash is one of the most cost-effective fast models on the market. It is ideal for high-volume multimodal analysis and low-latency completions.

Gemini 3.5 Flash is Google's ultra-fast multimodal model, designed for extreme speed and high scalability.

By TechCompare · Updated

Input tokens
30,000
per request
Output tokens
3,000
per request
Volume
1,000 / monthly
Standard API

Calculator

Cost Comparison

Based on 30,000 input tokens (50% cached), 3,000 output tokens, and 1,000 requests.

How this is calculated

Gemini 3.5 Flash is priced at $1.50 per million input tokens and $9.00 per million output tokens, supporting a 75% caching discount ($0.38 per million) and a 50% batch discount ($0.75 per million inputs, $4.50 per million outputs).

Verdict

Base pricing sits at $1.50/M input and $9/M output, with a 75% cache discount landing cached input at $0.38/M and a 50% batch discount dropping both rows to $0.75/M input and $4.50/M output. The cache discount is shallower than Anthropic's 90% but the lower base output rate compensates, and Flash's native multimodal support is what closes the case: audio, image, and video inputs hurtle through a single endpoint at the same per-token price, so classification and transcription pipelines that would otherwise need multiple models collapse into one bill.

More API Standalones scenarios

GPT-5.5 Pricing
Single-model gpt-5.5 cost estimate
View details ➜
GPT-5.4 Pricing
Single-model gpt-5.4 cost estimate
View details ➜
Claude Opus 4.8 Pricing
Single-model claude-opus-4.8 cost estimate
View details ➜

Frequently asked questions

Does Gemini 3.5 Flash support prompt caching?
Yes, Google offers a 75% discount on cached input tokens, reducing the input price to $0.38 per million for matches.
How much does Gemini 3.5 Flash cost per million output tokens?
Gemini 3.5 Flash charges $9.00 per million output tokens at standard pricing. With the 50% batch discount, that drops to $4.50 per million. Its 75% prompt caching discount brings cached input to $0.38 per million, making high-frequency multimodal workloads very cost-effective.
Is Gemini 3.5 Flash the cheapest multimodal model?
Among first-party frontier-tier multimodal models, yes at $0.50/M input and $3/M output (some sources quote $1.50/$9, which is the Gemini Flash standard rate). Native audio and video input are included, which is rare at this price tier. Compared to GPT-5.4 Mini ($0.75/$4.50), Gemini Flash wins specifically on multimodal input support.