Deploy AI Models in Minutes. Not Hours.
GPU Router removes infrastructure complexity and lets developers deploy AI models to GPU infrastructure with just a few clicks.
gpurouter deploy glm-4-9b-chat --size smallReal-time AWS Spot capacity
AWS Spot Placement Score — 0 (unlikely) to 10 (very likely to get capacity right now).
Live score history is still building — check back soon.
Everything you need to serve a model
No EC2 wrangling, no security groups, no server flags. Just the parts that matter.
One-click deploy
Pick a model and a GPU tier — we boot the instance, load the model, and hand you a URL and API key.
GPU that fits
We match each model to the smallest GPU that can serve it, from a single A10G to an 8× H200 node.
OpenAI-compatible API
Every endpoint speaks the standard chat-completions API. Point any HTTP client at the URL.
TTL cost guards
Set an auto-terminate window so dev instances never quietly rack up a bill overnight.
Private by default
Endpoints are protected by an API key and can be locked to a CIDR range you control.
Built-in load testing
Measure TTFT and throughput from both client and server before you ship to production.
How it works
From zero to a running endpoint in four steps.
Pick a model
Choose from the catalog. Each entry shows the GPU tiers it runs on and what it costs.
Choose a GPU tier
Small, medium, or large. We size the instance, disk, and server flags for you.
Deploy
We boot the instance, load the model, and run a health check until it is ready.
Call your endpoint
Get a URL + API key. Point any OpenAI-compatible client at it and start shipping.
Supported models
Four production-ready foundation models, each pre-tuned for the GPUs it runs on.
GLM-4 9B Chat
Zhipu AI
A capable 9B BF16 dense chat model that runs natively on any Ampere-or-newer GPU — no FP8 tensor cores required.
Qwen 3.5 4B
Alibaba Qwen
A compact vision-language model that pairs strong reasoning with image understanding — light enough for a single 24GB GPU.
Qwen 3.6 35B-A3B
Alibaba Qwen
An FP8 mixture-of-experts model with 35B total parameters — high quality with efficient FP8 serving on Ada/Hopper GPUs.
GLM-5.2
Zhipu AI
A frontier-scale FP8 MoE model with sparse attention and 1M native context. Runs on an 8× H200 node.
Why GPU Router
Built for people who want a private GPU-backed endpoint without hand-rolling infrastructure.
- No EC2, security groups, drivers, or vLLM flags to hand-roll
- GPU auto-matched to each model — pay only for what fits
- State lives in AWS tags, so endpoint keys are never lost
- One command up, one command down — billing stops instantly
- Server-side metrics scraped for real throughput numbers
- Restrict access to a CIDR range and rotate keys anytime
“We went from a blank AWS account to a live inference endpoint before our standup finished. It is genuinely a few clicks.”
“The GPU-fit sizing alone saved us from over-provisioning. The load-test report is exactly the data I needed.”
“The CLI was already great. The dashboard makes it something I can hand to the rest of the team without a walkthrough.”
Frequently asked questions
Everything you might want to know before deploying.
No. GPU Router picks a GPU that fits your model, boots the instance, loads the model, and returns a ready-to-call URL and API key. When you are done, one action tears it down and stops billing.
Ship your first endpoint today
Deploy a foundation model to your own GPU in minutes. No infra code, no console clicking.