Back to models

Qwen 3.5 4B

Trending

Qwen/Qwen3.5-4B

A compact FP8 model with tool calling and a 256K context window, served on a single 24GB GPU for about a dollar an hour.

Qwen 3.5 4B is a compact FP8 model served on vLLM with tool calling and reasoning parsing enabled. Its hybrid attention keeps only 8 of 32 layers at full attention, which is what lets a 262,144-token context fit in 24GB of VRAM alongside the weights - a window normally out of reach on a single A10G. It deploys to one g5.xlarge in your own AWS account and is ready to serve in 10-15 minutes.

ProviderAlibaba Qwen
CategoryChat
Parameters4B
PrecisionFP8
Context window256K
LicenseApache 2.0

Starting at

$1.01 / hour

g5.xlarge · NVIDIA A10G

You'll need a GPU Router account to deploy this model.

Min GPUA10G (24GB)
Params4B
Context256K