CUDA profiling with Nsight starts with Nsight Systems for a system-wide timeline to confirm the bottleneck is actually on the GPU, then Nsight Compute for per-kernel metrics like occupancy, memory throughput, and warp stall reasons to diagnose whether a kernel is compute-bound, memory-bound, or latency-bound, before matching the fix to that diagnosis and re-profiling after each change.

Every GPU optimization effort should start the same way: profiling. Yet it’s the step most often skipped in favor of applying known “best practices” and hoping they apply. NVIDIA’s Nsight tools exist specifically to replace that guesswork with data, here’s how to actually use them to find where your GPU time is going.

Nsight Systems: the system-wide view

Nsight Systems gives you a timeline view across CPU and GPU activity: kernel launches, memory transfers, CUDA API calls, and where gaps exist between them. This is the right starting point for any profiling effort, because it answers the first and most important question: is your bottleneck actually on the GPU, or is the GPU sitting idle waiting on something else (data loading, host-side preprocessing, PCIe transfers)?

Common findings at this level:

  • Large gaps between kernel launches, indicating the GPU is starved for work
  • Excessive host-to-device or device-to-host memory transfers
  • Serialization where operations that could overlap are running sequentially

Nsight Compute: the per-kernel deep dive

Once Nsight Systems identifies which kernels matter most, Nsight Compute provides detailed per-kernel metrics:

  • Occupancy, how many warps are active per streaming multiprocessor relative to the theoretical maximum. Low occupancy often means memory latency isn’t being hidden effectively.
  • Memory throughput, actual achieved bandwidth versus the GPU’s theoretical peak, which reveals whether a kernel is memory-bound and, if so, whether memory access patterns are coalesced efficiently.
  • Compute throughput, utilization of the GPU’s arithmetic units, revealing whether a kernel is genuinely compute-bound or whether it’s being held back by something else entirely.
  • Warp stall reasons, a breakdown of exactly why warps aren’t executing at any given moment (memory dependency, execution dependency, synchronization, and so on), which points directly at the fix.

A practical profiling workflow

  1. Start with Nsight Systems to identify the highest-impact section of your pipeline, don’t optimize a kernel that only accounts for 2% of runtime.
  2. Drill into that section with Nsight Compute to understand whether it’s compute-bound, memory-bound, or latency-bound.
  3. Match the fix to the diagnosis. Memory-bound kernels need better access patterns or reduced data movement, not more compute optimization. Compute-bound kernels benefit from precision reduction or algorithmic changes, not memory tuning.
  4. Re-profile after each change. Optimizations interact, fixing a memory bottleneck can reveal a compute bottleneck that was previously hidden behind it.
  5. Repeat until the marginal gain from further tuning is smaller than the engineering time it costs. Profiling tells you not just where to optimize, but when to stop.

Why this matters more as hardware gets more expensive

On a $30-40K B200, or across a 72-GPU NVL72 rack, the cost of running an untuned, poorly profiled workload compounds fast. Profiling isn’t an optional nicety, it’s the difference between knowing you’re getting your money’s worth out of expensive hardware and just assuming you are.

Ensigncode’s CUDA Performance Profiling service uses Nsight Systems and Nsight Compute to find exactly where your GPU workloads are losing performance, then fixes it at the kernel level. If your GPUs feel underutilized but you’re not sure why, book a free consultation and we’ll show you.

#CUDA#Profiling#Nsight#Performance

FAQ

Common questions

What's the difference between Nsight Systems and Nsight Compute?

Nsight Systems gives a system-wide timeline view across CPU and GPU activity to find where gaps and bottlenecks are at a high level. Nsight Compute drills into individual kernels with detailed metrics like occupancy, memory throughput, and warp stall reasons.

Why does profiling matter more as GPU hardware gets more expensive?

On a $30-40K B200, or across a 72-GPU NVL72 rack, the cost of running an untuned, poorly profiled workload compounds fast. Profiling is the difference between knowing you're getting your money's worth and just assuming you are.

Let us build something great

Have a Project in Mind?

Tell us about your goals and our engineers will recommend the right approach across GPU, AI, and Odoo ERP. Reach out for a free consultation.