Gemini 3.1 Flash-Lite
The cheapest active Gemini — high-volume, low-cost.
Modality
Text + vision
Context
Large context
Pricing (per 1M)
$0.25 in / $1.5 out
What Gemini 3.1 Flash-Lite is for
Flash Lite is the cheapest tier in Google's lineup, built for scale — the model you call in a loop over a million records without checking the bill nervously.
It is best understood as infrastructure rather than as an assistant. The question is not whether it is smart, but whether it is consistent enough to sit inside a pipeline unattended.
Where it does well
Cost per call, low enough to change architecture. Work that would be uneconomic at mid-tier prices — classifying every item in a large dataset, scoring every document, tagging every message — becomes routine. Latency is low enough for real-time paths.
Paired with a larger model behind a routing decision, it reduces total system cost while improving quality, because the expensive model only sees the cases that need it.
Where it falls short
Capability is genuinely limited. It needs unambiguous prompts, a narrow task definition, and validation on the output. Reasoning, nuance and long-context accuracy are all weak, and it will produce confident output regardless.
Do not use it for anything a customer reads unedited, and do not use it where a wrong answer propagates silently.
Should you use it?
Use Flash Lite for high-volume classification, extraction and filtering with validated output. Treat it as the first stage of a pipeline, never the last. The systems that get good results from cheap models are the ones that check the output rather than trusting it.
Capabilities
Pros
- Very cheap
- Low latency
- Scales well
Cons
- Lightweight on hard reasoning
Best for
Try a prompt for Gemini 3.1 Flash-Lite
Tag each message with a topic label. Return JSON only.
Other Gemini models
Pricing indicative, per 1M tokens, reviewed July 2026.
Visit official site