Back to models
GLM-5.2
Frontierzai-org/GLM-5.2-FP8
A frontier-scale FP8 MoE model with sparse attention and 1M native context. Runs on an 8× H200 node.
GLM-5.2 is a frontier-scale FP8 (E4M3) mixture-of-experts model with MLA + DSA sparse attention, ~756 GB of weights across 141 shards, 78 layers, 256 routed experts, and 1M native context. Its sparse MLA attention needs vLLM’s FlashMLA-Sparse backend, which only supports Hopper or Blackwell datacenter GPUs. The only supported single-instance topology is an 8× H200 node (p5e.48xlarge). Weights are streamed from a same-region public S3 bucket for fast, free in-region transfer. Always deploy with a TTL as a cost guard.
ProviderZhipu AI
CategoryReasoning
Parameters756 GB MoE
PrecisionFP8
Context window1M (capped 32K)
LicenseCustom (GLM)
Starting at
$35.00 / hour
p5e.48xlarge · 8× NVIDIA H200
You'll need a GPU Router account to deploy this model.
Min GPU8× H200
Params756 GB MoE
Context1M (capped 32K)