1. VRAM
Fit weights, runtime state, KV cache, batch and framework overhead with headroom.
The best GPU is not a universal model name. It is the configuration that fits your memory requirement and delivers the latency, throughput and cost your application needs.
Fit weights, runtime state, KV cache, batch and framework overhead with headroom.
Measure the latency and throughput your actual application cares about.
Compare cost per useful workload result, not hourly price in isolation.
For LLM serving, model parameters and numerical precision establish a starting memory requirement. Context length, KV cache, batching and concurrency add memory pressure. Quantization can reduce memory requirements, but the exact result depends on the model and runtime.
Compare Ampere and Hopper classes.
Compare 80 GB and 141 GB Hopper memory classes.
Compare Hopper and Blackwell classes.
See the methodology used for measured results.