Skip to content

Developer Tools & Frameworks

The base layer for programming NVIDIA GPUs: the CUDA Toolkit, CUDA-X libraries, Nsight profilers and debuggers, the NGC software catalog and AI Workbench for managing project environments.

Technology profiles for this category are in research.

Overview

NVIDIA's GPU libraries are built on CUDA. Developers who write kernels, port applications or tune performance work with this layer directly; others meet it when a driver and framework version do not match.

The CUDA Toolkit provides a C/C++ compiler, a runtime, GPU libraries and debugging tools, and version 13.4 adds support for the Rubin architecture. CUDA-X is the wider set of libraries built on CUDA, from cuDNN and CUTLASS to NCCL and data processing libraries, so many teams can call a library instead of writing kernels. Nsight Systems shows where time goes across CPU and GPU, and Nsight Compute inspects single kernels. NGC is a catalog of GPU-optimized containers, models and Helm charts, not a cloud. AI Workbench is a free tool for managing project environments across laptops, servers and cloud.

If you only use a high-level framework or a managed service, you will seldom touch these tools beyond installing drivers. Profiling becomes worth learning once performance or GPU cost starts to matter.12345

Problems it addresses

  • Finding the real bottleneck3

    Slow jobs are often held up by data loading or CPU work, not the GPU. Nsight Systems gives a system-wide timeline across CPUs and GPUs.

  • Kernel-level tuning3

    Custom kernels need detailed metrics. Nsight Compute is an interactive profiler for CUDA kernels.

  • Memory bugs in GPU code3

    Compute Sanitizer checks for memory access errors, shared memory hazards, uninitialized reads and synchronization misuse.

  • Version mismatches4

    Matching drivers, CUDA and frameworks by hand is error-prone. NGC containers package GPU-optimized software for workstations, clusters and cloud instances.

  • Rewriting common routines2

    CUDA-X libraries provide tuned routines for math, deep learning, communication and data processing.

A typical workflow

  1. Set up the environment1

    Install the CUDA Toolkit, or pull an NGC container that already includes it.

  2. Use libraries first2

    Call CUDA-X libraries such as cuDNN, NCCL or CUB before writing custom kernels.

  3. Profile the whole application3

    Capture a timeline with Nsight Systems to see CPU, GPU and transfer gaps.

  4. Tune and check hot kernels3

    Open the slowest kernels in Nsight Compute and check correctness with Compute Sanitizer.

  5. Package and share5

    Ship the result as a container, or manage the project across machines with AI Workbench.

NVIDIA Solution Architect

Describe your project and get an explainable architecture.

Open the lab

Next steps

  1. Join the free NVIDIA Developer Program for SDK access, forums and training.

  2. Profile one real workload with Nsight Systems and note the largest idle gap before optimizing anything.

  3. Pin driver, CUDA and framework versions in a container so results can be reproduced.

  4. Use the NVIDIA Solution Architect to see which higher-level NVIDIA software sits on top of this layer for your use case.

Sources

  1. NVIDIA CUDA Toolkit (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA CUDA-X Libraries (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA developer tools overview (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA NGC (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA AI Workbench (product page) (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)