Developer Tools & Frameworks
The base layer for programming NVIDIA GPUs: the CUDA Toolkit, CUDA-X libraries, Nsight profilers and debuggers, the NGC software catalog and AI Workbench for managing project environments.
Technology profiles for this category are in research.
Overview
NVIDIA's GPU libraries are built on CUDA. Developers who write kernels, port applications or tune performance work with this layer directly; others meet it when a driver and framework version do not match.
The CUDA Toolkit provides a C/C++ compiler, a runtime, GPU libraries and debugging tools, and version 13.4 adds support for the Rubin architecture. CUDA-X is the wider set of libraries built on CUDA, from cuDNN and CUTLASS to NCCL and data processing libraries, so many teams can call a library instead of writing kernels. Nsight Systems shows where time goes across CPU and GPU, and Nsight Compute inspects single kernels. NGC is a catalog of GPU-optimized containers, models and Helm charts, not a cloud. AI Workbench is a free tool for managing project environments across laptops, servers and cloud.
If you only use a high-level framework or a managed service, you will seldom touch these tools beyond installing drivers. Profiling becomes worth learning once performance or GPU cost starts to matter.12345
Problems it addresses
Finding the real bottleneck3
Slow jobs are often held up by data loading or CPU work, not the GPU. Nsight Systems gives a system-wide timeline across CPUs and GPUs.
Kernel-level tuning3
Custom kernels need detailed metrics. Nsight Compute is an interactive profiler for CUDA kernels.
Memory bugs in GPU code3
Compute Sanitizer checks for memory access errors, shared memory hazards, uninitialized reads and synchronization misuse.
Version mismatches4
Matching drivers, CUDA and frameworks by hand is error-prone. NGC containers package GPU-optimized software for workstations, clusters and cloud instances.
Rewriting common routines2
CUDA-X libraries provide tuned routines for math, deep learning, communication and data processing.
A typical workflow
Set up the environment1
Install the CUDA Toolkit, or pull an NGC container that already includes it.
Use libraries first2
Call CUDA-X libraries such as cuDNN, NCCL or CUB before writing custom kernels.
Profile the whole application3
Capture a timeline with Nsight Systems to see CPU, GPU and transfer gaps.
Tune and check hot kernels3
Open the slowest kernels in Nsight Compute and check correctness with Compute Sanitizer.
Package and share5
Ship the result as a container, or manage the project across machines with AI Workbench.
NVIDIA Solution Architect
Describe your project and get an explainable architecture.
Next steps
Join the free NVIDIA Developer Program for SDK access, forums and training.
Profile one real workload with Nsight Systems and note the largest idle gap before optimizing anything.
Pin driver, CUDA and framework versions in a container so results can be reproduced.
Use the NVIDIA Solution Architect to see which higher-level NVIDIA software sits on top of this layer for your use case.
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.