Back to models

Qwen 3.6 35B-A3B

New

Qwen/Qwen3.6-35B-A3B-FP8

An FP8 mixture-of-experts model with 35B total parameters — high quality with efficient FP8 serving on Ada/Hopper GPUs.

Qwen 3.6 35B-A3B is an FP8 (E4M3) mixture-of-experts model with ~37.5 GB of weights. It requires an FP8-capable GPU — Ada Lovelace (L40S / L4) or Hopper (H100 / H200). Single-GPU tiers use fp8 KV-cache to fit within 48GB, delivering strong quality at low-to-medium concurrency; the 4× L40S large tier unlocks 128K context and high throughput.

ProviderAlibaba Qwen
CategoryReasoning
Parameters35B (A3B MoE)
PrecisionFP8
Context window128K
LicenseApache 2.0

Starting at

$1.86 / hour

g6e.xlarge · NVIDIA L40S

You'll need a GPU Router account to deploy this model.

Min GPUL40S (48GB)
Params35B (A3B MoE)
Context128K