Model catalog
Four production-ready foundation models. Pick one, choose a GPU tier, and deploy in a few clicks.
Filters
Provider
Category
GPU Type
Showing 4 of 4 models
GLM-4 9B Chat
Zhipu AI
A capable 9B BF16 dense chat model that runs natively on any Ampere-or-newer GPU — no FP8 tensor cores required.
Min GPU
A10G (24GB)
Memory
24 GB
Params
9B
Precision
BF16
Chat128K
Qwen 3.5 4B
Alibaba Qwen
A compact vision-language model that pairs strong reasoning with image understanding — light enough for a single 24GB GPU.
Min GPU
A10G (24GB)
Memory
24 GB
Params
4B
Precision
BF16
Vision-Language128K
Qwen 3.6 35B-A3B
Alibaba Qwen
An FP8 mixture-of-experts model with 35B total parameters — high quality with efficient FP8 serving on Ada/Hopper GPUs.
Min GPU
L40S (48GB)
Memory
48 GB
Params
35B (A3B MoE)
Precision
FP8
Reasoning128K
GLM-5.2
Zhipu AI
A frontier-scale FP8 MoE model with sparse attention and 1M native context. Runs on an 8× H200 node.
Min GPU
8× H200
Memory
1128 GB
Params
756 GB MoE
Precision
FP8
Reasoning1M (capped 32K)