NVIDIA Nsight Developer Tools
NVIDIA Nsight is NVIDIA's family of developer tools for profiling, debugging and analyzing software on NVIDIA GPUs. Nsight Systems shows a timeline of CPU, GPU and system activity, Nsight Compute profiles individual CUDA kernels, and Nsight Graphics covers graphics, alongside debuggers and IDE plugins.1
Also known as Nsight
At a glance
- What is it?
- Nsight is a set of tools, libraries and SDKs for building, debugging and profiling software on NVIDIA hardware. The main members are Nsight Systems, a system-wide performance analyzer; Nsight Compute, an interactive profiler for CUDA and OptiX kernels; and Nsight Graphics, a debugger and profiler for Direct3D, Vulkan and OpenGL applications. The family also includes Nsight Visual Studio Edition and Visual Studio Code Edition, CUDA-GDB, Compute Sanitizer, Nsight Copilot (an AI coding assistant for CUDA), a JupyterLab extension, Nsight Python, and lower-level interfaces such as CUPTI and NVTX. Current versions are Nsight Systems 2026.5.1 and Nsight Compute 2026.3.1.123
- What does it do?
- Nsight Systems records what an application does across CPUs, GPUs, memory, network and the operating system and places it on one timeline, so developers can see idle GPUs, slow data transfers or serialized work. It traces CUDA, cuBLAS, cuDNN and TensorRT as well as graphics APIs, samples GPU metrics such as SM and Tensor Core activity, PCIe and NVLink traffic, and can analyze many nodes at once. Nsight Compute then examines a single kernel in depth: it collects hardware metrics, flags performance limiters through guided analysis, compares runs against a baseline and maps metrics back to source lines in CUDA C++, PTX, SASS, Fortran or Python. Compute Sanitizer checks for memory and synchronization errors, and CUDA-GDB debugs kernels on real hardware.124
- Who needs it?
- Developers who write or tune GPU code and need to know why it is slower than expected: CUDA developers, AI engineers looking at training or inference pipelines, HPC teams scaling across nodes, and graphics developers chasing frame stutter. NVIDIA recommends starting with Nsight Systems for a whole-application view and moving to Nsight Compute or Nsight Graphics for specific kernels or frames.1
- What does it need?35
- An NVIDIA GPU supported by the tool you use; Nsight Compute requires Turing or newer
- A supported host: Nsight Compute runs on Windows, Linux (x86_64 and aarch64 SBSA), WSL2, and macOS 13 or later as a viewer and remote host
- A supported NVIDIA driver; for Nsight Systems GPU Metrics, NVIDIA lists minimum driver versions per GPU architecture (for example r440 for Turing)
- The CUDA Toolkit, HPC SDK or JetPack, or the standalone tool downloads
- On Windows, administrator rights to run the Nsight Systems CLI
- What it is not
- Nsight is not one application but a family of separate tools with their own downloads and system requirements. It does not optimize code automatically; it shows where time and resources go and suggests causes. Nsight Compute does not support older GPU architectures such as Maxwell, Pascal and Volta. In our view, the tools do not replace application-level monitoring in production; they suit development and performance investigations.13
Availability and licensing. Nsight tools are available as standalone downloads from their product pages and are also packaged in the CUDA Toolkit, the HPC SDK and NVIDIA JetPack. Nsight Compute documentation refers to the NVIDIA Software License Agreement for its terms. Current versions are Nsight Systems 2026.5.1 and Nsight Compute 2026.3.1.46
The problem it solves
GPU applications often run far below the hardware's capability because work is serialized, data transfers stall the GPU, kernels use memory inefficiently or CPU threads cannot feed the device fast enough. These causes are hard to find from wall-clock timings or GPU utilization numbers alone.
Nsight tools collect detailed traces and hardware counters and present them on timelines and analysis pages, so developers can see which part of the system limits performance and which lines of code are responsible.2
How it works
The family splits the work by level of detail:
- Nsight Systems (nsys): attaches to an application with low overhead and records CUDA, NVTX, OS, library and graphics API events plus sampled GPU metrics into a report that is viewed on a timeline. The CLI command nsys profile collects data, and nsys stats and nsys analyze produce summary reports. Multi-node analysis and plugins extend it to clusters and custom data sources.
- Nsight Compute (ncu): replays or instruments individual kernels to collect hardware metrics organized in sections. The GUI offers guided analysis, memory charts and source correlation; the CLI writes reports (for example ncu -o profile) and collects the full metric set with --set full. Rules, custom sections and a Python report interface allow automated analysis.
- Correctness and debugging: Compute Sanitizer and CUDA-GDB, with IDE integration through Nsight Visual Studio Edition and Visual Studio Code Edition.
- Interfaces: NVTX lets applications annotate ranges that appear in the tools, and CUPTI lets others build profiling tools.
Many tools ship inside the CUDA Toolkit, the HPC SDK and JetPack, and a newer version is sometimes available as a standalone download when it was released after the toolkit shipped.1567
Diagram as a list
Applications & solutions
- Your application (CUDA, Python, AI framework, graphics)The workload being analyzed; may add NVTX rangesConnects to Nsight Systems (nsys), Nsight Compute (ncu), Compute Sanitizer and CUDA-GDB
- Reports, GUI and IDE plugins (VS, VS Code, JupyterLab)View, compare and script analysis results
Operations & orchestration
- Nsight Systems (nsys)System-wide timeline of CPU, GPU, OS and network activityConnects to CUPTI and NVTX interfaces, Reports, GUI and IDE plugins (VS, VS Code, JupyterLab)
- Nsight Compute (ncu)Per-kernel hardware metrics and guided analysisConnects to CUPTI and NVTX interfaces, Reports, GUI and IDE plugins (VS, VS Code, JupyterLab)
- Compute Sanitizer and CUDA-GDBCorrectness checks and kernel debuggingConnects to NVIDIA GPU and driver
Accelerated computing
- CUPTI and NVTX interfacesCollect trace and metric data from the CUDA stackConnects to NVIDIA GPU and driver
- NVIDIA GPU and driverHardware whose counters and activity are measured
Capabilities
System-wide timeline (Nsight Systems)2
Shows CPU threads, GPU work, CUDA library calls, OS events and network activity on one timeline.
Why it matters: Reveals idle GPUs, transfer stalls and CPU bottlenecks across the whole application.
Limits: Shows where time goes but not detailed kernel internals; use Nsight Compute for that.
GPU metrics sampling25
Samples PCIe, NVLink and DRAM activity, SM utilization, Tensor Core activity and warp occupancy.
Why it matters: Connects hardware usage to the code that caused it.
Limits: GPU Metrics needs a Turing or newer GPU, Linux (x86-64 or aarch64) or Windows targets, and elevated permissions.
Multi-node and Python analysis2
Nsight Systems analyzes many nodes at once and samples Python backtraces; a JupyterLab extension profiles notebook cells.
Why it matters: Helps AI and HPC teams tune distributed training and Python pipelines.
Limits: Large multi-node traces produce large reports that need analysis time.
Kernel profiling with guided analysis (Nsight Compute)347
Collects detailed hardware metrics per kernel, flags performance limiters and recommends actions, with baseline comparison.
Why it matters: Developers do not need to be GPU architecture experts to find kernel bottlenecks.
Limits: Supports Turing and newer GPUs only; collecting many metrics can require replaying each kernel several times.
Source-level correlation4
Maps metrics and warp stall samples to SASS, PTX and source lines in CUDA C/C++, Fortran, OpenACC or Python.
Why it matters: Points to the exact lines that cause memory or latency problems.
Limits: How much source detail appears depends on the build and language; check the profiling guide for your setup.
Correctness and debugging tools1
Compute Sanitizer checks memory access errors, shared memory hazards, uninitialized reads and synchronization misuse; CUDA-GDB debugs kernels on hardware.
Why it matters: Finds bugs that only appear on the GPU.
Limits: We recommend running it as a test step rather than in production runs.
Graphics and game tools12
Nsight Graphics debugs and profiles Direct3D, Vulkan, OpenGL and ray tracing apps; Nsight Systems detects frame stutter; Aftermath creates GPU crash dumps.
Why it matters: Covers rendering performance and GPU crash analysis for games and visualization.
Limits: Separate tools with their own platform and API coverage.
Practical use cases
A deep learning training job shows low GPU utilization and the team does not know why.
- Approach
- Profile a few training steps with Nsight Systems, tracing CUDA and NVTX, and look for gaps between GPU kernels.
- Role of NVIDIA Nsight Developer Tools
- Nsight Systems shows whether data loading, CPU work or communication starves the GPU.
- Data, infrastructure and skills
- Access to the training host and permission to run a profiler.
- Type of benefit
- Better GPU utilization
- Caveats
- Profile a short, representative window; long traces are large.
- First step
- Run nsys profile --trace=cuda,nvtx on a short training run and open the report.
Sources 2
A custom CUDA kernel is slower than expected on new GPUs.
- Approach
- Profile it with Nsight Compute, follow the guided analysis and use source correlation to find the slow lines.
- Role of NVIDIA Nsight Developer Tools
- Nsight Compute identifies memory, occupancy or instruction limits.
- Data, infrastructure and skills
- A Turing or newer GPU and the source code of the kernel.
- Type of benefit
- Faster kernels
- Caveats
- Metric collection may replay each kernel several times, so measure end-to-end time in normal runs.
- First step
- Run ncu -o profile on the application and review the details page.
Sources 4
An HPC application scales poorly across many GPU nodes.
- Approach
- Use Nsight Systems multi-node analysis with network metrics to find communication or load-balance limits.
- Role of NVIDIA Nsight Developer Tools
- Nsight Systems correlates activity across nodes, GPUs and interconnects.
- Data, infrastructure and skills
- A cluster where profiling is allowed and storage for reports.
- Type of benefit
- Better scaling
- Caveats
- Collection on many nodes needs coordination with the job scheduler.
- First step
- Profile a small multi-node run and compare per-node timelines.
Sources 2
Works with
Optional integration
- NVIDIA CUDA ToolkitNsight Systems, Nsight Compute and Compute Sanitizer ship with the CUDA Toolkit.
- NVIDIA JetsonNsight tools are packaged in JetPack and Nsight Systems supports Jetson.
Complementary tools
- NVIDIA TensorRT LLMNVIDIA documents profiling TensorRT LLM execution with Nsight Systems.
- NVIDIA DGXNsight Systems profiles workloads on DGX systems, including multi-node runs.
Optional integration for
- NVIDIA RTX RemixThe dxvk-remix Runtime can self-inject Nsight Graphics for programmatic frame captures.
- NVIDIA CUDA ToolkitNsight Compute, Nsight Systems and related tools ship with the CUDA Toolkit.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Get the tools1
Use the Nsight Systems and Nsight Compute versions included in your CUDA Toolkit, HPC SDK or JetPack, or check the product pages, which sometimes offer a newer standalone version.
Check: The nsys and ncu command line tools run on the target machine.
Annotate your code (optional)1
Add NVTX ranges around key phases such as data loading, forward pass or solver iterations.
Check: The named ranges appear on the Nsight Systems timeline.
Capture a system timeline5
Run nsys profile --trace=cuda,nvtx on a short, representative run, then use nsys stats for summary tables or open the report in the GUI.
Check: The timeline shows CPU threads, CUDA API calls and GPU kernels.
Profile the slowest kernels7
Run ncu -o profile on the application, adding --set full for the complete metric set, and open the report in Nsight Compute.
Check: The details page lists performance limiters and recommendations for each profiled kernel.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA Nsight Developer Tools as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA developer tools overview (opens in a new tab)
- NVIDIA Nsight Systems product page (opens in a new tab)
- Nsight Compute release notes (opens in a new tab)
- NVIDIA Nsight Compute product page (opens in a new tab)
- Nsight Systems user guide (opens in a new tab)
- Nsight Compute documentation hub (opens in a new tab)
- Nsight Compute CLI documentation (opens in a new tab)
- CUDA Toolkit 13.4 Update 1 release notes (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.