vLLM's PagedAttention and continuous batching make it fast out of the box, but production throughput depends on tuning it further: FP8 (and max_num_batched_tokens) on H100, FP4 and larger batch sizes on B200, and separating SLA-bound traffic from tolerant background jobs. Default configuration is built to work broadly, not to be optimal for any one deployment.

vLLM has become the default serving engine for teams that want strong throughput without sacrificing operational flexibility: easier to iterate on than TensorRT-LLM, with most of the performance techniques that made TensorRT-LLM fast now available as configurable options. Whether you’re running it on H100 or B200, the defaults leave real performance on the table.

What makes vLLM fast

PagedAttention is vLLM’s core innovation, it manages KV cache memory the way an operating system manages virtual memory, splitting it into non-contiguous blocks. This eliminates the memory fragmentation that historically capped batch sizes and caused out-of-memory errors under variable-length request loads.

Continuous batching lets new requests join an in-progress batch as soon as GPU capacity frees up, rather than waiting for the entire batch to complete, a major throughput advantage over static batching for real-world traffic with mixed request lengths.

Tuning vLLM for H100

  • FP8 quantization is natively supported and should be close to a default choice on H100, it delivers a solid throughput gain with minimal quality impact for most instruction-tuned models.
  • max_num_batched_tokens needs to be tuned to your specific VRAM budget and traffic pattern, too conservative wastes capacity, too aggressive risks memory pressure under load spikes.
  • KV cache quantization (INT8 or INT4) can reclaim a meaningful chunk of VRAM, freeing it up for more concurrent requests, particularly valuable for long-context workloads.

Tuning vLLM for B200

  • FP4 support via vLLM’s integration with Blackwell’s Transformer Engine unlocks further throughput beyond FP8, though, as with any FP4 deployment, it needs validation against your eval suite before production rollout.
  • Larger batch sizes are viable given the B200’s 192GB of memory, but vLLM’s defaults may not automatically scale batch configuration to exploit that headroom, this often needs to be tuned manually.
  • Speculative decoding (using a smaller draft model to propose tokens verified by the larger model) has shown meaningful cost-per-token reductions in production benchmarks, particularly for code-heavy or low-concurrency workloads where the decode phase leaves compute underused.

Common production pitfalls

  1. Treating vLLM as install-and-forget. Default configurations are reasonable starting points, not tuned endpoints.
  2. Not separating SLA-bound and batch traffic. Latency-sensitive requests and tolerant background jobs (embeddings, batch summarization) benefit from different scheduling and even different infrastructure, mixing them on the same pool wastes capacity.
  3. Skipping utilization monitoring. A GPU showing low utilization on nvidia-smi while still incurring full hourly cost is one of the most common, and most fixable, sources of wasted GPU spend.
  4. Ignoring engine choice per workload. vLLM’s flexibility is valuable during iteration; TensorRT-LLM’s raw throughput may be worth the reduced flexibility once a workload is stable and throughput-critical.

Where to focus first

If you’re running vLLM in production today without having explicitly tuned max_num_batched_tokens, quantization, and batching strategy for your specific hardware and traffic pattern, there’s very likely throughput being left unused, the default configuration is built to work broadly, not to be optimal for any one deployment.

Ensigncode’s LLM Inference Infrastructure and AI Inference Optimization services cover vLLM deployment and tuning across both H100 and Blackwell hardware. Book a free consultation to get a second set of eyes on your serving configuration.

#vLLM#Inference#H100#B200

FAQ

Common questions

What makes vLLM fast by default?

PagedAttention, which manages KV cache memory like an OS manages virtual memory in non-contiguous blocks, eliminating fragmentation that used to cap batch sizes; and continuous batching, which lets new requests join an in-progress batch as soon as capacity frees up.

What's different about tuning vLLM for B200 vs H100?

On B200, FP4 support unlocks throughput beyond FP8 (with eval validation needed), larger batch sizes become viable given 192GB of memory, and speculative decoding shows meaningful cost-per-token reductions. On H100, FP8 quantization and tuning max_num_batched_tokens to your VRAM budget are the primary levers.

Let us build something great

Have a Project in Mind?

Tell us about your goals and our engineers will recommend the right approach across GPU, AI, and Odoo ERP. Reach out for a free consultation.