ColinBuilds.com · Alibaba Cloud Model Studio hosted pricing
Qwen3-235B-A22B
- Input / 1M
- $0.70
- Output / 1M
- $2.80
- Native 32K context (131K with YaRN scaling)
- 32,768 tokens
Current Alibaba Cloud Model Studio international pricing for qwen3-235b-a22b; non-thinking output $2.80/1M, thinking output $8.40/1M · as of June 30, 2026
Compare models
Same per-1M input/output rates as above. Enter your monthly token volumes to compare bills.
Qwen3-235B-A22B
- Input: 1M × $0.70 = $0.70
- Output: 1M × $2.80 = $2.80
- Total
- $3.50
GPT-4o
- Input: 1M × $2.50 = $2.50
- Output: 1M × $10.00 = $10.00
- Total
- $12.50
Difference
$9.00
GPT-4o costs $9.00 more than Qwen3-235B-A22B.
Benchmark Matrix
Reported model metrics from published sources.
Active Parameters
22B
Reported active parameter count at inference time.
Popular comparisons
30 total featuring Qwen3-235B-A22B.
Qwen3-235B-A22B is Alibaba/Qwen’s large open-weight Mixture-of-Experts model for builders who want a serious China-model alternative to closed frontier APIs, with Alibaba Cloud Model Studio pricing used for practical cost comparison.
What this model is
Qwen3-235B-A22B is one of the flagship Qwen3 MoE releases. The Hugging Face model card lists 235B total parameters and 22B activated parameters, which is why this page labels it as an MoE model rather than a dense 235B model.
Qwen3 is built around both thinking and non-thinking modes. In beginner terms, that means the same model family can be used for deeper reasoning-style answers or faster direct responses depending on how it is called.
Pricing notes
Alibaba Cloud Model Studio lists qwen3-235b-a22b for the International service with separate output pricing for non-thinking and thinking modes. The listed input price is $0.70 per 1M tokens. Non-thinking output is $2.80 per 1M tokens, while thinking-mode output is $8.40 per 1M tokens.
The calculator on this page uses the non-thinking output price by default because it is the safer baseline for quick comparison against GPT-4o, Claude Sonnet, DeepSeek, GLM, and Mistral. Thinking-mode output can be materially more expensive, so it should be treated as a separate cost case rather than silently blended into the default calculator rate.
This is Alibaba Cloud Model Studio hosted pricing, not OpenRouter or another reseller route.
Benchmarks and specs
The Hugging Face model card lists Qwen3-235B-A22B as 235B total parameters with 22B activated parameters. It also says Qwen3 natively supports context lengths up to 32,768 tokens and that Qwen validated performance up to 131,072 tokens using YaRN scaling.
The Qwen3 announcement says Qwen3 supports hybrid thinking modes and was released under the Apache 2.0 license. Qwen provides benchmark comparisons on its announcement and model card, but this page does not list specific benchmark scores until they are tied to a named source and test setup.
Best fit
Qwen3-235B-A22B is best for China/open-weight ecosystem comparisons, explaining MoE active parameters to beginners, and comparing Alibaba-hosted Qwen pricing against DeepSeek, GLM, Kimi, GPT-4o, and Claude Sonnet.
MoE sizing is a useful lesson here: total parameters (235B) and activated parameters (22B) are not the same number, and thinking-mode output pricing is priced separately from normal quick-response output.