GLM-5.2
Frontierzai-org/GLM-5.2-FP8
A frontier-scale FP8 MoE model with sparse attention and 1M native context. Runs on an 8× H200 node.
GLM-5.2 is a frontier-scale FP8 (E4M3) mixture-of-experts model with MLA + DSA sparse attention, ~704 GiB of weights across 150 shards, 78 layers, 256 routed experts, and 1M native context. Its sparse MLA attention needs vLLM’s FlashMLA-Sparse backend, which only supports Hopper or Blackwell datacenter GPUs, so it runs on an 8× H200 node (p5en.48xlarge) or an 8× B200 node. It launches as a persistent spot request: capacity for these instances is scarce, so the request stays open and takes capacity the moment it appears rather than failing immediately. Weights stream from a same-region public S3 bucket.
Starting at
$20.01 / hour
p5en.48xlarge · 8× NVIDIA H200
You'll need a GPU Router account to deploy this model.