ColinBuilds.com · Open-weight profile · provider pricing separate
Llama 3.3 70B Instruct
- Published Meta context window
- 128,000 tokens
Provider pricing
USD / 1M tokens| Provider | Input / 1M | Output / 1M | Source |
|---|---|---|---|
| OpenRouter
Best deal
| $0.10 | $0.32 | Provider docs |
| DeepInfra
Best deal
| $0.10 | $0.32 | Provider docs |
| Groq | $0.59 | $0.79 | Provider docs |
Provider prices are verified from third-party docs — not official model-maker pricing unless labelled Official.
Compare models
Same per-1M input/output rates as above. Enter your monthly token volumes to compare bills.
Llama 3.3 70B Instruct
- Input: 1M × $0.10 = $0.10
- Output: 1M × $0.32 = $0.32
- Total
- $0.42
GPT-4o
- Input: 1M × $2.50 = $2.50
- Output: 1M × $10.00 = $10.00
- Total
- $12.50
Difference
$12.08
GPT-4o costs $12.08 more than Llama 3.3 70B Instruct.
Benchmark Matrix
Reported model metrics from published sources.
MMLU Score
86%
MMLU-Pro Score
68.9%
Active Parameters
70B
Reported active parameter count at inference time.
Popular comparisons
30 total featuring Llama 3.3 70B Instruct.
Llama 3.3 70B Instruct is Meta’s updated 70B open-weight chat model: a December 2024 refresh of the Llama 3.1 70B tier with stronger published benchmark scores while keeping the 128K context window.
What this model is
This is the 70B instruction-tuned member of the Llama 3.3 family. Meta describes it as a multilingual text-in/text-out model optimized for dialogue use cases, with support for multilingual text and code.
The main thing to understand is that Llama 3.3 70B Instruct is not one single priced API product from Meta. Meta provides the model weights under the Llama 3.3 Community License. Hosted inference prices depend on the provider, route name, quantization, and throughput tier. This page records model facts from Meta’s model card and provider-hosted prices from separate docs.
Pricing notes
OpenRouter lists Llama 3.3 70B Instruct at $0.10 per 1M input tokens and $0.32 per 1M output tokens, checked on 2026-07-08.
DeepInfra lists Llama 3.3 70B Instruct Turbo at $0.10 per 1M input tokens and $0.32 per 1M output tokens, checked on 2026-07-08.
Groq lists Llama 3.3 70B Versatile at $0.59 per 1M input tokens and $0.79 per 1M output tokens, checked on 2026-07-08.
The calculator on this page uses the lowest combined provider rate on this page (OpenRouter and DeepInfra tie at $0.10 input / $0.32 output). This is provider-hosted pricing, not an official Meta creator price.
Benchmarks and specs
Meta’s Llama 3.3 model card reports these values for the instruction-tuned 70B model:
- Context window: 128,000 tokens
- MMLU (CoT): 86.0
- MMLU-Pro (CoT): 68.9
- Parameters: 70B
- Model release date: 2024-12-06
- Knowledge cutoff: December 2023
- Supported languages listed by Meta: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai
Best fit
- Open-weight 70B comparisons where Llama 3.1 70B is too old but 405B is too heavy
- Shopping hosted provider prices with a clear best-deal table
- RAG, assistant, coding, and multilingual evaluations where teams want model-weight access
- Comparing the 70B open-weight tier against Qwen, DeepSeek, Mistral, GPT, Claude, and Gemini options
Sources
- https://raw.githubusercontent.com/meta-llama/llama-models/main/models/llama3_3/MODEL_CARD.md
- https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_3/
- OpenRouter Llama 3.3 70B pricing
- DeepInfra Llama 3.3 70B Instruct Turbo pricing
- Groq Llama 3.3 70B Versatile pricing
- https://huggingface.co/api/models/meta-llama/Llama-3.3-70B-Instruct