Back to models
Qwen 3.6 35B-A3B
NewQwen/Qwen3.6-35B-A3B-FP8
An FP8 mixture-of-experts model with 35B total parameters — high quality with efficient FP8 serving on Ada/Hopper GPUs.
Qwen 3.6 35B-A3B is an FP8 (E4M3) mixture-of-experts model with ~37.5 GB of weights. It requires an FP8-capable GPU — Ada Lovelace (L40S / L4) or Hopper (H100 / H200). Single-GPU tiers use fp8 KV-cache to fit within 48GB, delivering strong quality at low-to-medium concurrency; the 4× L40S large tier unlocks 128K context and high throughput.
ProviderAlibaba Qwen
CategoryReasoning
Parameters35B (A3B MoE)
PrecisionFP8
Context window128K
LicenseApache 2.0
Starting at
$1.86 / hour
g6e.xlarge · NVIDIA L40S
You'll need a GPU Router account to deploy this model.
Min GPUL40S (48GB)
Params35B (A3B MoE)
Context128K