BurstDock
BURSTDOCK LABS / GPU BENCHMARKS

GPU benchmarks.
Numbers with context.

A fast GPU on paper is not automatically the fastest GPU for your model. BurstDock benchmarks are designed around reproducible AI workloads, with the test conditions published next to the result.

Latency

Time-to-first-token and generation latency for model-serving tests.

Throughput

Tokens per second and request throughput where the workload supports it.

Memory

VRAM usage, model precision and relevant runtime configuration.

What every result
must disclose.

We do not publish invented benchmark numbers or mix incompatible test conditions into a leaderboard. A BurstDock result should be reproducible enough for another engineer to understand what was actually measured.

  • Exact GPU model and GPU count
  • Model, parameter size and numerical precision
  • Inference or training framework and version
  • Input/output length, batch and concurrency
  • Warm-up policy and number of measured runs
  • Median plus tail latency where applicable

LLM serving

Time-to-first-token, output throughput and concurrency under fixed model and context settings.

Long-context inference

Memory pressure and latency as prompt length grows, with precision and runtime held constant.

Fine-tuning

Step time, memory use and throughput for explicitly documented training configurations.

Cost efficiency

Measured workload output relative to the BurstDock customer price during the test window.

MEASUREMENT STANDARD

Results you can put in context.

BurstDock benchmark results use documented workloads and test conditions so developers can understand what was measured and compare results on a meaningful basis.