Latency
Time-to-first-token and generation latency for model-serving tests.
A fast GPU on paper is not automatically the fastest GPU for your model. BurstDock benchmarks are designed around reproducible AI workloads, with the test conditions published next to the result.
Time-to-first-token and generation latency for model-serving tests.
Tokens per second and request throughput where the workload supports it.
VRAM usage, model precision and relevant runtime configuration.
We do not publish invented benchmark numbers or mix incompatible test conditions into a leaderboard. A BurstDock result should be reproducible enough for another engineer to understand what was actually measured.
Time-to-first-token, output throughput and concurrency under fixed model and context settings.
Memory pressure and latency as prompt length grows, with precision and runtime held constant.
Step time, memory use and throughput for explicitly documented training configurations.
Measured workload output relative to the BurstDock customer price during the test window.
BurstDock benchmark results use documented workloads and test conditions so developers can understand what was measured and compare results on a meaningful basis.
80 GB Hopper-class GPU compute.
141 GB for memory-intensive AI workloads.
180 GB Blackwell-class GPU compute.
High-memory Blackwell Ultra GPU compute.