What do you want to run?

Pick a model, get a machine. Enter a model and quantization, check it against a real GPU (or your own VRAM), and see exactly what fits — no dead-end math, no generic advice.

Quantization
Context length
VRAM
Pick a model and a GPU (or enter VRAM) to see the fit verdict.

Browse by VRAM

What you can run at each common GPU capacity, from an 8 GB starter card to a 48 GB workstation part.

Model registry

Flagship models with published quant sizes, minimum VRAM, and estimated throughput.

Estimated tok/s shown on a reference RTX 5090 (32 GB).

Guides

Unified memory: Strix Halo & mini-PCs

AMD’s Strix Halo (Ryzen AI Max) mini-PCs and Apple Silicon Macs share one pool of memory between CPU and GPU instead of a fixed VRAM budget on a discrete card. A 64–128 GB unified-memory configuration can hold models that would otherwise need two or three discrete GPUs pooled together — at the cost of the raw bandwidth a top discrete GPU has. It’s the fastest-growing category with almost no independent coverage today, and it shows up as its own row in every model’s “Runs on” module.

Real hardware, real numbers

Every VRAM figure on this site comes from published quant file sizes plus standard KV-cache math, and every throughput figure is a bandwidth-derived estimate, clearly labeled. See the full benchmark leaderboard for real submitted hardware results.

How the fit math works

Weights, KV cache, and overhead — see exactly how VRAM need and tokens/sec are estimated.

Read the methodology