Meta
Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
At a glance
- Input / 1M
- $0.22
- Output / 1M
- $0.50
- Context
- 131K
- Max output
- 16K
- Intelligence
- 7.7
- Released
- Dec 2024
Pricing
Llama 3.3 70B Instruct API pricing
Per-token rates, plus what common tasks actually cost.
| Input tokensPer 1M tokens | $0.22 |
|---|---|
| Output tokensPer 1M tokens | $0.50 |
| Cached input (read)Per 1M tokens | $0.11 |
| Cache writePer 1M tokens | Not available |
Chat message
2K in, 500 out
$0.69
per 1,000 requests
Coding task
30K in, 4K out
$8.60
per 1,000 requests
Long document summary
150K in, 2K out
$34
per 1,000 requests
Benchmarks
Llama 3.3 70B Instruct benchmark scores
Independent scores from Artificial Analysis. Higher is better.
Intelligence index
7.7
Coding index
11.9
Agentic index
Not available
Specs
Context window and capabilities
What it can read, what it can write, and which API features it supports.
- Context window
- 131K tokens
- Max output
- 16K tokens
- Input types
- Text
- Output types
- Text
- Reasoning effort
- Not available
- Release date
- December 6, 2024
- Knowledge cutoff
- December 31, 2023
- Tool calling
- Structured outputs
- JSON mode
- Reasoning
- Temperature
- Stop sequences
- Deterministic seed
- Verbosity control
Compare
Llama 3.3 70B Instruct vs other models
See how Llama 3.3 70B Instruct stacks up head to head on price, benchmarks, and features.
Llama 3.3 70B Instruct vs Gemini 2.5 Flash Lite
Meta vs Google
See comparisonLlama 3.3 70B Instruct vs GPT-4o-mini
Meta vs OpenAI
See comparisonLlama 3.3 70B Instruct vs Gemma 3 27B
Meta vs Google
See comparisonLlama 3.3 70B Instruct vs Llama 3.1 8B Instruct
Meta vs Meta
See comparisonLlama 3.3 70B Instruct vs GPT-4.1 Nano
Meta vs OpenAI
See comparisonLlama 3.3 70B Instruct vs Llama 4 Maverick
Meta vs Meta
See comparisonMore Llama 3.3 70B Instruct comparisons
- Llama 3.3 70B Instruct vs Llama 4 Scout
- Llama 3.3 70B Instruct vs Qwen3 VL 30B A3B Instruct
- Llama 3.3 70B Instruct vs Qwen3 VL 8B Instruct
- Llama 3.3 70B Instruct vs Qwen2.5 72B Instruct
- Llama 3.3 70B Instruct vs Sonar
- Llama 3.3 70B Instruct vs Qwen3 30B A3B
- Llama 3.3 70B Instruct vs Solar Pro 3
- Llama 3.3 70B Instruct vs Hermes 3 405B Instruct
- Llama 3.3 70B Instruct vs Hermes 3 70B Instruct
- Llama 3.3 70B Instruct vs Llama 3.1 70B Instruct
- Llama 3.3 70B Instruct vs Saba
- Llama 3.3 70B Instruct vs Llama 3.2 3B Instruct
- Llama 3.3 70B Instruct vs Llama 3.2 1B Instruct
Alternatives
Llama 3.3 70B Instruct alternatives
Similar models worth considering, and where each one has the edge.
Gemini 2.5 Flash Lite
Meta
Llama 4 Maverick
Meta
Llama 4 Scout
Want to learn about other models? Browse all models
FAQ
Frequently asked questions
How much does Llama 3.3 70B Instruct cost?
Llama 3.3 70B Instruct costs $0.22 per 1M input tokens and $0.50 per 1M output tokens through the API. A typical chat message costs about $0.69 per 1,000 messages.
What is the context window of Llama 3.3 70B Instruct?
Llama 3.3 70B Instruct accepts up to 131K tokens of input and can write up to 16K tokens in a single response.
What is the best alternative to Llama 3.3 70B Instruct?
Gemini 2.5 Flash Lite from Google is a strong alternative. It offers: 1.3x cheaper output, higher intelligence score, larger 1.05m context.
Can I use Llama 3.3 70B Instruct alongside other models?
Yes. Shortcut Chat sends one prompt to Llama 3.3 70B Instruct and any other models you pick at the same time, so you can compare answers side by side.
Try Llama 3.3 70B Instruct next to every other model.
Use Llama 3.3 70B Instruct alongside all the other models. Use one at a time, or multiple together. Get the best answer from every AI.