Skip to content

NVIDIA CUDA Toolkit

The NVIDIA CUDA Toolkit is the development kit for programming NVIDIA GPUs: a compiler, runtime and driver APIs, math and parallel libraries, debugging and profiling tools, and documentation. The current release is CUDA 13.4 Update 1, and the GPU driver is now installed separately.12

Also known as CUDA, CUDA Toolkit, CTK

At a glance

What is it?
The CUDA Toolkit is NVIDIA's software development kit for building GPU-accelerated applications. It contains the NVCC compiler and runtime compilation tools, the CUDA Runtime and Driver APIs, core libraries such as cuBLAS, cuFFT, cuRAND, cuSOLVER, cuSPARSE, Thrust, CUB and libcu++, and developer tools such as Nsight Compute, Nsight Systems, Compute Sanitizer and cuda-gdb. Applications built with it run on systems from embedded devices and workstations to data centers, clouds and supercomputers. The current release is 13.4 Update 1; recent versions add CUDA Tile, a tile-based programming model available in C++ and in Python as cuTile Python.13
What does it do?
Developers write GPU kernels in CUDA C++ (or tile kernels in C++ and Python), compile them with NVCC into code for specific GPU architectures, and call the runtime to move data and launch work on the GPU. Most applications also call the bundled libraries for linear algebra, Fourier transforms, sparse math, random numbers and parallel algorithms instead of writing kernels by hand. The toolkit's profilers and debuggers then show where time goes and catch memory errors. CUDA 13.4 adds support for Windows on Arm on RTX Spark devices, preview support for the NVIDIA Rubin architecture and a third version of the Multi-Process Service (MPS) for sharing GPUs.23
Who needs it?
Anyone writing or building software that runs code directly on NVIDIA GPUs: developers of simulations, scientific and HPC codes, deep learning frameworks, data processing libraries, graphics and media tools, and embedded applications. In our view, teams that only use ready-made frameworks need the full toolkit mainly when they compile their own GPU code.1
What does it need?24
  • A CUDA-capable NVIDIA GPU listed on NVIDIA's CUDA GPUs page
  • A supported operating system, for example Ubuntu 22.04, 24.04 or 26.04 LTS, RHEL, SUSE, Debian, Fedora, Amazon Linux or Azure Linux, or Windows
  • An NVIDIA driver installed separately: R615 or later for CUDA 13.4 features, 580 or later to run existing CUDA 13.x applications
  • A supported host compiler such as gcc on Linux for building code (not needed just to run applications)
What it is not
The CUDA Toolkit is not the GPU driver: since CUDA 13.4 on Linux (and 13.1 on Windows) the driver is no longer bundled and must be installed on its own. It is not CUDA-X, the wider set of domain libraries such as those for data science, which build on top of CUDA. It needs a CUDA-capable NVIDIA GPU; NVIDIA keeps a list of CUDA-capable GPU products. Running a finished CUDA application needs a CUDA-capable GPU and a compatible driver; the compiler toolchain is needed only for development.245

Availability and licensing. The CUDA Toolkit is distributed by NVIDIA under the License Agreement for NVIDIA Software Development Kits with a CUDA-specific supplement; NVIDIA's developer page presents it among its free tools. CUDA containers are also available from the NGC catalog. The current release is CUDA 13.4 Update 1.16

The problem it solves

CPUs alone cannot keep up with the arithmetic needed by large simulations, AI training, signal processing or data analytics. GPUs offer thousands of parallel threads, but using them requires a compiler that targets GPU architectures, a runtime to manage memory and launches, tuned math libraries and tools to find performance and correctness problems.

The CUDA Toolkit bundles these pieces in one versioned release, with documented driver compatibility, so software teams can write GPU code once and build it for the NVIDIA architectures they target.

How it works

The toolkit is organized in layers:

  • Compilers: NVCC compiles CUDA C++ source, NVRTC compiles at runtime, and nvJitLink and nvFatbin handle linking and packaging. PTX is the virtual instruction set that GPU code is compiled through; CUDA 13.4 supports PTX ISA 9.4.
  • APIs: the CUDA Runtime API and the lower-level Driver API manage devices, memory, streams and kernel launches, together with the CUDA Math API.
  • Libraries: cuBLAS, cuFFT, cuRAND, cuSOLVER, cuSPARSE, NPP and nvJPEG, plus the CUDA C++ Core Libraries (Thrust, CUB, libcu++).
  • Tile programming: CUDA Tile IR with CUDA Tile C++ and cuTile Python for writing kernels at the level of data tiles.
  • Tools: Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb, CUPTI and NVML.

Each toolkit release pairs with a driver branch (R615 for CUDA 13.4). Under CUDA minor version compatibility, CUDA 13.x applications run on drivers from 580 onward, while new 13.4 features need R615 or later.23

NVIDIA CUDA Toolkit architecture: components by layer and how they connectApplications &solutionsOperations &orchestrationAcceleratedcomputingYour application (C++ or Python): Calls CUDA libraries and launches kernelsYour application (C++ orPython)Tools (Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb): Profile and debug the applicationTools (Nsight Systems,Nsight Compute, Compute…CUDA libraries (cuBLAS, cuFFT, cuSPARSE, Thrust, CUB): Tuned math and parallel algorithmsCUDA libraries(cuBLAS, cuFFT,…Compilers (NVCC, NVRTC, nvJitLink, Tile IR): Compile host and device code through PTXCompilers (NVCC,NVRTC,…CUDA Runtime and Driver APIs: Manage memory, streams and kernel launchesCUDA Runtime andDriver APIsNVIDIA GPU driver (installed separately): Executes work on the GPU; R615 branch for CUDA 13.4NVIDIA GPUdriver…NVIDIA GPU: Runs the compiled kernelsNVIDIA GPU
Diagram as a list
  1. Applications & solutions

    • Your application (C++ or Python)Calls CUDA libraries and launches kernelsConnects to CUDA libraries (cuBLAS, cuFFT, cuSPARSE, Thrust, CUB), Compilers (NVCC, NVRTC, nvJitLink, Tile IR), CUDA Runtime and Driver APIs
  2. Operations & orchestration

    • Tools (Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb)Profile and debug the applicationConnects to Your application (C++ or Python)
  3. Accelerated computing

    • CUDA libraries (cuBLAS, cuFFT, cuSPARSE, Thrust, CUB)Tuned math and parallel algorithmsConnects to CUDA Runtime and Driver APIs
    • Compilers (NVCC, NVRTC, nvJitLink, Tile IR)Compile host and device code through PTXConnects to CUDA Runtime and Driver APIs
    • CUDA Runtime and Driver APIsManage memory, streams and kernel launchesConnects to NVIDIA GPU driver (installed separately)
    • NVIDIA GPU driver (installed separately)Executes work on the GPU; R615 branch for CUDA 13.4Connects to NVIDIA GPU
    • NVIDIA GPURuns the compiled kernels
Components and connections as documented by NVIDIA.2

Capabilities

  • GPU compilers34

    NVCC compiles CUDA C++, NVRTC compiles at runtime, and nvJitLink links code for target GPUs through PTX.

    Why it matters: Turns your source code into programs that run on specific NVIDIA GPU generations.

    Limits: Needs a supported host compiler version; very new distribution compilers may require a compatibility package.

  • Math and parallel libraries23

    cuBLAS, cuFFT, cuRAND, cuSOLVER, cuSPARSE, Thrust, CUB and libcu++ ship with the toolkit.

    Why it matters: Most applications call tuned libraries instead of writing their own kernels.

    Limits: Library releases contain known issues on specific GPUs; read the release notes for your architecture.

  • Tile programming model1

    CUDA Tile IR, CUDA Tile C++ and cuTile Python let developers write kernels in terms of data tiles.

    Why it matters: Opens kernel writing to Python developers and simplifies some C++ kernels.

    Limits: A newer programming model; check which GPUs and features it supports in your release.

  • Profiling and debugging tools78

    Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb and CUPTI are part of the toolkit.

    Why it matters: Finds bottlenecks and memory errors without separate products.

    Limits: Standalone tool downloads can be newer than the versions bundled with a toolkit release.

  • Driver compatibility rules2

    Each release maps to a driver branch, and minor version compatibility lets CUDA 13.x applications run on drivers 580 and later.

    Why it matters: Lets teams upgrade the toolkit without always upgrading cluster drivers.

    Limits: New features and platforms in 13.4 need an R615 or later driver.

  • Broad platform coverage123

    Targets embedded systems, workstations, data centers, clouds and supercomputers, and from 13.4 Windows on Arm for RTX Spark.

    Why it matters: One programming model across the NVIDIA hardware range.

    Limits: Rubin support in 13.4 is a preview; CUDA 14.0 will raise the minimum Arm architecture for ARM64-SBSA.

Practical use cases

A simulation or scientific code spends most of its time in loops that could run in parallel.
Approach
Move the hot loops into CUDA C++ kernels or library calls such as cuBLAS and cuFFT, then profile with Nsight.
Role of NVIDIA CUDA Toolkit
The toolkit provides the compiler, libraries and profilers for the port.
Data, infrastructure and skills
A CUDA-capable GPU, a supported compiler and developers comfortable with parallel programming.
Type of benefit
Faster computation
Caveats
Gains depend on how parallel the code is and on data movement between CPU and GPU.
First step
Profile the CPU version to find the hot loops, then port one of them.

Sources 1

Python developers need a custom GPU operation that existing libraries do not offer.
Approach
Write the operation as a tile kernel with cuTile Python.
Role of NVIDIA CUDA Toolkit
The toolkit supplies the CUDA Tile model and its compilation path.
Data, infrastructure and skills
A recent CUDA 13 toolkit and a GPU supported by CUDA Tile.
Type of benefit
Developer productivity
Caveats
A new programming model; check support and examples for your GPU.
First step
Work through the cuTile Python documentation examples.

Sources 1

A team must build and ship GPU software consistently across developer machines and clusters.
Approach
Build inside CUDA containers from the NGC catalog and keep the driver on hosts within the compatibility range.
Role of NVIDIA CUDA Toolkit
Containers fix the CUDA version used for builds, and NVIDIA's release notes state which driver branches each CUDA version needs.
Data, infrastructure and skills
A container runtime with GPU support and hosts with compatible drivers.
Type of benefit
Operational consistency
Caveats
The NVIDIA driver is no longer bundled with the toolkit, so each host needs a compatible driver installed separately (580 or later for CUDA 13.x applications).
First step
Pull a CUDA container from NGC that matches your target CUDA version.

Sources 12

Works with

Optional integration

Complementary tools

  • NVIDIA CUDA-X Data ScienceCUDA-X data science libraries are built on CUDA and require a matching CUDA version.
  • NVIDIA cuOptcuOpt is built on CUDA and requires CUDA 12 or 13.
  • NVIDIA JetsonNVIDIA documents CUDA upgrades for Jetson devices; the toolkit targets embedded systems.

Required by

  • NVIDIA CUDA-X Data ScienceThe libraries need CUDA 12 or 13 and a matching driver; pip wheels must match the installed CUDA major version.
  • NVIDIA cuOptcuOpt needs CUDA 12.0+ or 13.0+ and a compatible driver.

Optional integration for

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Check the GPU and system4

    Run lspci | grep -i nvidia, confirm the GPU on NVIDIA's CUDA GPUs list, and check that the Linux distribution and gcc version are supported.

    Check: The GPU appears in lspci output and gcc --version reports a supported compiler.

  2. Install the driver2

    Install an NVIDIA driver separately, since CUDA 13.4 no longer bundles it on Linux. Use R615 or later for 13.4 features.

    Check: The installed driver version meets the branch listed in the release notes.

  3. Install the toolkit4

    Add NVIDIA's package repository for your distribution, install the cuda-toolkit package and add /usr/local/cuda-13.4/bin to PATH.

    Check: nvcc is found on the PATH from the CUDA 13.4 install directory.

  4. Verify with the samples4

    Build deviceQuery and bandwidthTest from the CUDA samples repository on GitHub and run them.

    Check: deviceQuery lists your GPU and bandwidthTest completes without errors.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA CUDA Toolkit as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA CUDA Toolkit product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. CUDA Toolkit 13.4 Update 1 release notes (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. CUDA Toolkit documentation hub (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. CUDA installation guide for Linux (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  5. NVIDIA CUDA-X for Data Science developer page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. CUDA Toolkit end user license agreement (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  7. NVIDIA Nsight Compute product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  8. NVIDIA developer tools overview (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)