BurstDock
TUTORIAL / LLM INFERENCE

Deploy an LLM.
Measure what matters.

A framework-neutral workflow for taking an LLM inference workload from sizing to a measured GPU deployment.

1. Define the
serving target.

Record the exact model, precision, maximum context, expected prompt/output lengths and concurrency. Without those inputs, a GPU recommendation is guesswork.

  • Model and numerical precision
  • Context and output length
  • Concurrent requests
  • Latency and throughput target

2. Size and
deploy.

Estimate model memory, add serving headroom, then select a suitable GPU class in BurstDock. Deploy the runtime you intend to use in production rather than benchmarking an unrelated stack.

GPU sizing calculator

3. Validate
under load.

Warm the model, run representative prompts and capture time-to-first-token, output throughput and tail latency at realistic concurrency. Compare total workload cost, not just hourly GPU price.

Deploy GPU compute