Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.
At a glance
- Input / 1M
- $0.25
- Output / 1M
- $1.50
- Context
- 1.05M
- Max output
- 66K
- Intelligence
- 15.6
- Released
- May 2026
Pricing
Gemini 3.1 Flash Lite API pricing
Per-token rates, plus what common tasks actually cost.
| Input tokensPer 1M tokens | $0.25 |
|---|---|
| Output tokensPer 1M tokens | $1.50 |
| Cached input (read)Per 1M tokens | $0.03 |
| Cache writePer 1M tokens | $0.08 |
| Web searchPer request | $0.014 |
Chat message
2K in, 500 out
$1.25
per 1,000 requests
Coding task
30K in, 4K out
$13.50
per 1,000 requests
Long document summary
150K in, 2K out
$40.50
per 1,000 requests
Benchmarks
Gemini 3.1 Flash Lite benchmark scores
Independent scores from Artificial Analysis. Higher is better.
Intelligence index
15.6
Coding index
34.7
Agentic index
1.6
Specs
Context window and capabilities
What it can read, what it can write, and which API features it supports.
- Context window
- 1.05M tokens
- Max output
- 66K tokens
- Input types
- Audio, File, Image, Text, Video
- Output types
- Text
- Reasoning effort
- High, Medium, Low, Minimal
- Release date
- May 7, 2026
- Tool calling
- Structured outputs
- JSON mode
- Reasoning
- Temperature
- Stop sequences
- Deterministic seed
- Verbosity control
Compare
Gemini 3.1 Flash Lite vs other models
See how Gemini 3.1 Flash Lite stacks up head to head on price, benchmarks, and features.
Gemini 3.1 Flash Lite vs Gemini 2.5 Flash
Google vs Google
See comparisonGemini 3.1 Flash Lite vs Gemma 4 31B
Google vs Google
See comparisonGemini 3.1 Flash Lite vs Gemma 4 26B A4B
Google vs Google
See comparisonGemini 3.1 Flash Lite vs Claude Haiku 4.5
Google vs Anthropic
See comparisonGemini 3.1 Flash Lite vs Qwen3.6 35B A3B
Google vs Qwen
See comparisonGemini 3.1 Flash Lite vs Gemini 2.5 Pro
Google vs Google
See comparisonMore Gemini 3.1 Flash Lite comparisons
- Gemini 3.1 Flash Lite vs DeepSeek V3.1
- Gemini 3.1 Flash Lite vs Muse Glimmer 30B
- Gemini 3.1 Flash Lite vs GLM 4.6
- Gemini 3.1 Flash Lite vs GLM 4.7 Flash
- Gemini 3.1 Flash Lite vs DeepSeek V3.2 Exp
- Gemini 3.1 Flash Lite vs R1 0528
- Gemini 3.1 Flash Lite vs LongCat 2.0
- Gemini 3.1 Flash Lite vs Qwen3.5-122B-A10B
- Gemini 3.1 Flash Lite vs Kimi K2 0905
- Gemini 3.1 Flash Lite vs Mistral Medium 3.5
- Gemini 3.1 Flash Lite vs Step 3.5 Flash
- Gemini 3.1 Flash Lite vs DeepSeek V3.1 Terminus
- Gemini 3.1 Flash Lite vs o4 Mini
- Gemini 3.1 Flash Lite vs Mercury 2
- Gemini 3.1 Flash Lite vs Nova 2 Lite
- Gemini 3.1 Flash Lite vs GLM 4.5
- Gemini 3.1 Flash Lite vs Kimi K2 0711
- Gemini 3.1 Flash Lite vs MiniMax M1
Alternatives
Gemini 3.1 Flash Lite alternatives
Similar models worth considering, and where each one has the edge.
Z.ai
GLM 4.6
Qwen
Qwen3.6 35B A3B
Meta
Muse Glimmer 30B
Want to learn about other models? Browse all models
FAQ
Frequently asked questions
How much does Gemini 3.1 Flash Lite cost?
Gemini 3.1 Flash Lite costs $0.25 per 1M input tokens and $1.50 per 1M output tokens through the API. A typical chat message costs about $1.25 per 1,000 messages.
What is the context window of Gemini 3.1 Flash Lite?
Gemini 3.1 Flash Lite accepts up to 1.05M tokens of input and can write up to 66K tokens in a single response.
What is the best alternative to Gemini 3.1 Flash Lite?
GLM 4.6 from Z.ai is a strong alternative. It offers: higher intelligence score.
Can I use Gemini 3.1 Flash Lite alongside other models?
Yes. Shortcut Chat sends one prompt to Gemini 3.1 Flash Lite and any other models you pick at the same time, so you can compare answers side by side.
Try Gemini 3.1 Flash Lite next to every other model.
Use Gemini 3.1 Flash Lite alongside all the other models. Use one at a time, or multiple together. Get the best answer from every AI.