Back to models

GLM-4 9B Chat

Popular

zai-org/glm-4-9b-chat-hf

A capable 9B BF16 dense chat model that runs natively on any Ampere-or-newer GPU — no FP8 tensor cores required.

GLM-4 9B Chat is a BF16 dense 9B model (~18.8 GB of weights) built for general chat and assistant workloads. Because it runs in BF16, it works on the cheaper 24GB g5/g6 instances, and its grouped-query attention keeps the KV cache small so low-concurrency 8K serving fits on a single card. Scale up to L40S or a 4× L40S node for long context and higher throughput.

ProviderZhipu AI
CategoryChat
Parameters9B
PrecisionBF16
Context window128K
LicenseApache 2.0

Starting at

$1.01 / hour

g5.xlarge · NVIDIA A10G

You'll need a GPU Router account to deploy this model.

Min GPUA10G (24GB)
Params9B
Context128K