Every GPU optimization effort should start the same way: profiling. Yet it’s the step most often skipped in favor of applying known “best practices” and hoping they apply. NVIDIA’s Nsight tools exist specifically to replace that guesswork with data, here’s how to actually use them to find where your GPU time is going.
Nsight Systems: the system-wide view
Nsight Systems gives you a timeline view across CPU and GPU activity: kernel launches, memory transfers, CUDA API calls, and where gaps exist between them. This is the right starting point for any profiling effort, because it answers the first and most important question: is your bottleneck actually on the GPU, or is the GPU sitting idle waiting on something else (data loading, host-side preprocessing, PCIe transfers)?
Common findings at this level:
- Large gaps between kernel launches, indicating the GPU is starved for work
- Excessive host-to-device or device-to-host memory transfers
- Serialization where operations that could overlap are running sequentially
Nsight Compute: the per-kernel deep dive
Once Nsight Systems identifies which kernels matter most, Nsight Compute provides detailed per-kernel metrics:
- Occupancy, how many warps are active per streaming multiprocessor relative to the theoretical maximum. Low occupancy often means memory latency isn’t being hidden effectively.
- Memory throughput, actual achieved bandwidth versus the GPU’s theoretical peak, which reveals whether a kernel is memory-bound and, if so, whether memory access patterns are coalesced efficiently.
- Compute throughput, utilization of the GPU’s arithmetic units, revealing whether a kernel is genuinely compute-bound or whether it’s being held back by something else entirely.
- Warp stall reasons, a breakdown of exactly why warps aren’t executing at any given moment (memory dependency, execution dependency, synchronization, and so on), which points directly at the fix.
A practical profiling workflow
- Start with Nsight Systems to identify the highest-impact section of your pipeline, don’t optimize a kernel that only accounts for 2% of runtime.
- Drill into that section with Nsight Compute to understand whether it’s compute-bound, memory-bound, or latency-bound.
- Match the fix to the diagnosis. Memory-bound kernels need better access patterns or reduced data movement, not more compute optimization. Compute-bound kernels benefit from precision reduction or algorithmic changes, not memory tuning.
- Re-profile after each change. Optimizations interact, fixing a memory bottleneck can reveal a compute bottleneck that was previously hidden behind it.
- Repeat until the marginal gain from further tuning is smaller than the engineering time it costs. Profiling tells you not just where to optimize, but when to stop.
Why this matters more as hardware gets more expensive
On a $30-40K B200, or across a 72-GPU NVL72 rack, the cost of running an untuned, poorly profiled workload compounds fast. Profiling isn’t an optional nicety, it’s the difference between knowing you’re getting your money’s worth out of expensive hardware and just assuming you are.
Ensigncode’s CUDA Performance Profiling service uses Nsight Systems and Nsight Compute to find exactly where your GPU workloads are losing performance, then fixes it at the kernel level. If your GPUs feel underutilized but you’re not sure why, book a free consultation and we’ll show you.