Model catalog

Deploy an open-weight model onto a GPU in your own AWS account. Connect once, then launch and delete from here.

Step 1 - connect your AWS account

Models deploy into your own AWS account, so you need an account here first. The GPU is billed by AWS directly - never by us.

Filters

Provider

Category

Showing 4 of 4 models. 2 available to deploy today; the rest are on the way.

Qwen 3.5 4B

Alibaba Qwen

Trending

A compact FP8 model with tool calling and a 256K context window, served on a single 24GB GPU for about a dollar an hour.

Min GPU
A10G (24GB)
Memory
24 GB
Params
4B
Precision
FP8
Chat256K$1.01/hr

GLM-5.2

Zhipu AI

Frontier

A frontier-scale FP8 MoE model with sparse attention and 1M native context. Runs on an 8× H200 node.

Min GPU
8× H200
Memory
1128 GB
Params
756 GB MoE
Precision
FP8
Reasoning1M (capped 32K)$20.01/hr

GLM-4 9B Chat

Zhipu AI

Coming soon

A capable 9B BF16 dense chat model that runs natively on any Ampere-or-newer GPU — no FP8 tensor cores required.

Min GPU
A10G (24GB)
Memory
24 GB
Params
9B
Precision
BF16
Chat128K

Qwen 3.6 35B-A3B

Alibaba Qwen

Coming soon

An FP8 mixture-of-experts model with 35B total parameters — high quality with efficient FP8 serving on Ada/Hopper GPUs.

Min GPU
L40S (48GB)
Memory
48 GB
Params
35B (A3B MoE)
Precision
FP8
Reasoning128K