vLLM has become the default serving engine for teams that want strong throughput without sacrificing operational flexibility: easier to iterate on than TensorRT-LLM, with most of the performance techniques that made TensorRT-LLM fast now available as configurable options. Whether you’re running it on H100 or B200, the defaults leave real performance on the table.
What makes vLLM fast
PagedAttention is vLLM’s core innovation, it manages KV cache memory the way an operating system manages virtual memory, splitting it into non-contiguous blocks. This eliminates the memory fragmentation that historically capped batch sizes and caused out-of-memory errors under variable-length request loads.
Continuous batching lets new requests join an in-progress batch as soon as GPU capacity frees up, rather than waiting for the entire batch to complete, a major throughput advantage over static batching for real-world traffic with mixed request lengths.
Tuning vLLM for H100
- FP8 quantization is natively supported and should be close to a default choice on H100, it delivers a solid throughput gain with minimal quality impact for most instruction-tuned models.
max_num_batched_tokensneeds to be tuned to your specific VRAM budget and traffic pattern, too conservative wastes capacity, too aggressive risks memory pressure under load spikes.- KV cache quantization (INT8 or INT4) can reclaim a meaningful chunk of VRAM, freeing it up for more concurrent requests, particularly valuable for long-context workloads.
Tuning vLLM for B200
- FP4 support via vLLM’s integration with Blackwell’s Transformer Engine unlocks further throughput beyond FP8, though, as with any FP4 deployment, it needs validation against your eval suite before production rollout.
- Larger batch sizes are viable given the B200’s 192GB of memory, but vLLM’s defaults may not automatically scale batch configuration to exploit that headroom, this often needs to be tuned manually.
- Speculative decoding (using a smaller draft model to propose tokens verified by the larger model) has shown meaningful cost-per-token reductions in production benchmarks, particularly for code-heavy or low-concurrency workloads where the decode phase leaves compute underused.
Common production pitfalls
- Treating vLLM as install-and-forget. Default configurations are reasonable starting points, not tuned endpoints.
- Not separating SLA-bound and batch traffic. Latency-sensitive requests and tolerant background jobs (embeddings, batch summarization) benefit from different scheduling and even different infrastructure, mixing them on the same pool wastes capacity.
- Skipping utilization monitoring. A GPU showing low utilization on
nvidia-smiwhile still incurring full hourly cost is one of the most common, and most fixable, sources of wasted GPU spend. - Ignoring engine choice per workload. vLLM’s flexibility is valuable during iteration; TensorRT-LLM’s raw throughput may be worth the reduced flexibility once a workload is stable and throughput-critical.
Where to focus first
If you’re running vLLM in production today without having explicitly tuned max_num_batched_tokens, quantization, and batching strategy for your specific hardware and traffic pattern, there’s very likely throughput being left unused, the default configuration is built to work broadly, not to be optimal for any one deployment.
Ensigncode’s LLM Inference Infrastructure and AI Inference Optimization services cover vLLM deployment and tuning across both H100 and Blackwell hardware. Book a free consultation to get a second set of eyes on your serving configuration.