ColinBuilds.com

ColinBuilds.com · Open-weight profile · provider pricing separate

Llama 3.3 70B Instruct

Published Meta context window
128,000 tokens

Provider pricing

USD / 1M tokens
Provider Input / 1M Output / 1M Source
OpenRouter Best deal
$0.10 $0.32 Provider docs
DeepInfra Best deal
$0.10 $0.32 Provider docs
Groq
$0.59 $0.79 Provider docs

Provider prices are verified from third-party docs — not official model-maker pricing unless labelled Official.

Compare models

Same per-1M input/output rates as above. Enter your monthly token volumes to compare bills.

Llama 3.3 70B Instruct

Input: 1M × $0.10 = $0.10
Output: 1M × $0.32 = $0.32
Total
$0.42

GPT-4o

Input: 1M × $2.50 = $2.50
Output: 1M × $10.00 = $10.00
Total
$12.50

Difference

$12.08

GPT-4o costs $12.08 more than Llama 3.3 70B Instruct.

Benchmark Matrix

Reported model metrics from published sources.

MMLU Score

86%

MMLU-Pro Score

68.9%

Active Parameters

70B

Reported active parameter count at inference time.

Popular comparisons

30 total featuring Llama 3.3 70B Instruct.

Open compare hub →

Llama 3.3 70B Instruct is Meta’s updated 70B open-weight chat model: a December 2024 refresh of the Llama 3.1 70B tier with stronger published benchmark scores while keeping the 128K context window.

What this model is

This is the 70B instruction-tuned member of the Llama 3.3 family. Meta describes it as a multilingual text-in/text-out model optimized for dialogue use cases, with support for multilingual text and code.

The main thing to understand is that Llama 3.3 70B Instruct is not one single priced API product from Meta. Meta provides the model weights under the Llama 3.3 Community License. Hosted inference prices depend on the provider, route name, quantization, and throughput tier. This page records model facts from Meta’s model card and provider-hosted prices from separate docs.

Pricing notes

OpenRouter lists Llama 3.3 70B Instruct at $0.10 per 1M input tokens and $0.32 per 1M output tokens, checked on 2026-07-08.

DeepInfra lists Llama 3.3 70B Instruct Turbo at $0.10 per 1M input tokens and $0.32 per 1M output tokens, checked on 2026-07-08.

Groq lists Llama 3.3 70B Versatile at $0.59 per 1M input tokens and $0.79 per 1M output tokens, checked on 2026-07-08.

The calculator on this page uses the lowest combined provider rate on this page (OpenRouter and DeepInfra tie at $0.10 input / $0.32 output). This is provider-hosted pricing, not an official Meta creator price.

Benchmarks and specs

Meta’s Llama 3.3 model card reports these values for the instruction-tuned 70B model:

  • Context window: 128,000 tokens
  • MMLU (CoT): 86.0
  • MMLU-Pro (CoT): 68.9
  • Parameters: 70B
  • Model release date: 2024-12-06
  • Knowledge cutoff: December 2023
  • Supported languages listed by Meta: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai

Best fit

  • Open-weight 70B comparisons where Llama 3.1 70B is too old but 405B is too heavy
  • Shopping hosted provider prices with a clear best-deal table
  • RAG, assistant, coding, and multilingual evaluations where teams want model-weight access
  • Comparing the 70B open-weight tier against Qwen, DeepSeek, Mistral, GPT, Claude, and Gemini options

Sources