NVIDIA CUDA Toolkit
The NVIDIA CUDA Toolkit is the development kit for programming NVIDIA GPUs: a compiler, runtime and driver APIs, math and parallel libraries, debugging and profiling tools, and documentation. The current release is CUDA 13.4 Update 1, and the GPU driver is now installed separately.12
Also known as CUDA, CUDA Toolkit, CTK
At a glance
- What is it?
- The CUDA Toolkit is NVIDIA's software development kit for building GPU-accelerated applications. It contains the NVCC compiler and runtime compilation tools, the CUDA Runtime and Driver APIs, core libraries such as cuBLAS, cuFFT, cuRAND, cuSOLVER, cuSPARSE, Thrust, CUB and libcu++, and developer tools such as Nsight Compute, Nsight Systems, Compute Sanitizer and cuda-gdb. Applications built with it run on systems from embedded devices and workstations to data centers, clouds and supercomputers. The current release is 13.4 Update 1; recent versions add CUDA Tile, a tile-based programming model available in C++ and in Python as cuTile Python.13
- What does it do?
- Developers write GPU kernels in CUDA C++ (or tile kernels in C++ and Python), compile them with NVCC into code for specific GPU architectures, and call the runtime to move data and launch work on the GPU. Most applications also call the bundled libraries for linear algebra, Fourier transforms, sparse math, random numbers and parallel algorithms instead of writing kernels by hand. The toolkit's profilers and debuggers then show where time goes and catch memory errors. CUDA 13.4 adds support for Windows on Arm on RTX Spark devices, preview support for the NVIDIA Rubin architecture and a third version of the Multi-Process Service (MPS) for sharing GPUs.23
- Who needs it?
- Anyone writing or building software that runs code directly on NVIDIA GPUs: developers of simulations, scientific and HPC codes, deep learning frameworks, data processing libraries, graphics and media tools, and embedded applications. In our view, teams that only use ready-made frameworks need the full toolkit mainly when they compile their own GPU code.1
- What does it need?24
- A CUDA-capable NVIDIA GPU listed on NVIDIA's CUDA GPUs page
- A supported operating system, for example Ubuntu 22.04, 24.04 or 26.04 LTS, RHEL, SUSE, Debian, Fedora, Amazon Linux or Azure Linux, or Windows
- An NVIDIA driver installed separately: R615 or later for CUDA 13.4 features, 580 or later to run existing CUDA 13.x applications
- A supported host compiler such as gcc on Linux for building code (not needed just to run applications)
- What it is not
- The CUDA Toolkit is not the GPU driver: since CUDA 13.4 on Linux (and 13.1 on Windows) the driver is no longer bundled and must be installed on its own. It is not CUDA-X, the wider set of domain libraries such as those for data science, which build on top of CUDA. It needs a CUDA-capable NVIDIA GPU; NVIDIA keeps a list of CUDA-capable GPU products. Running a finished CUDA application needs a CUDA-capable GPU and a compatible driver; the compiler toolchain is needed only for development.245
Availability and licensing. The CUDA Toolkit is distributed by NVIDIA under the License Agreement for NVIDIA Software Development Kits with a CUDA-specific supplement; NVIDIA's developer page presents it among its free tools. CUDA containers are also available from the NGC catalog. The current release is CUDA 13.4 Update 1.16
The problem it solves
CPUs alone cannot keep up with the arithmetic needed by large simulations, AI training, signal processing or data analytics. GPUs offer thousands of parallel threads, but using them requires a compiler that targets GPU architectures, a runtime to manage memory and launches, tuned math libraries and tools to find performance and correctness problems.
The CUDA Toolkit bundles these pieces in one versioned release, with documented driver compatibility, so software teams can write GPU code once and build it for the NVIDIA architectures they target.
How it works
The toolkit is organized in layers:
- Compilers: NVCC compiles CUDA C++ source, NVRTC compiles at runtime, and nvJitLink and nvFatbin handle linking and packaging. PTX is the virtual instruction set that GPU code is compiled through; CUDA 13.4 supports PTX ISA 9.4.
- APIs: the CUDA Runtime API and the lower-level Driver API manage devices, memory, streams and kernel launches, together with the CUDA Math API.
- Libraries: cuBLAS, cuFFT, cuRAND, cuSOLVER, cuSPARSE, NPP and nvJPEG, plus the CUDA C++ Core Libraries (Thrust, CUB, libcu++).
- Tile programming: CUDA Tile IR with CUDA Tile C++ and cuTile Python for writing kernels at the level of data tiles.
- Tools: Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb, CUPTI and NVML.
Each toolkit release pairs with a driver branch (R615 for CUDA 13.4). Under CUDA minor version compatibility, CUDA 13.x applications run on drivers from 580 onward, while new 13.4 features need R615 or later.23
Diagram as a list
Applications & solutions
- Your application (C++ or Python)Calls CUDA libraries and launches kernelsConnects to CUDA libraries (cuBLAS, cuFFT, cuSPARSE, Thrust, CUB), Compilers (NVCC, NVRTC, nvJitLink, Tile IR), CUDA Runtime and Driver APIs
Operations & orchestration
- Tools (Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb)Profile and debug the applicationConnects to Your application (C++ or Python)
Accelerated computing
- CUDA libraries (cuBLAS, cuFFT, cuSPARSE, Thrust, CUB)Tuned math and parallel algorithmsConnects to CUDA Runtime and Driver APIs
- Compilers (NVCC, NVRTC, nvJitLink, Tile IR)Compile host and device code through PTXConnects to CUDA Runtime and Driver APIs
- CUDA Runtime and Driver APIsManage memory, streams and kernel launchesConnects to NVIDIA GPU driver (installed separately)
- NVIDIA GPU driver (installed separately)Executes work on the GPU; R615 branch for CUDA 13.4Connects to NVIDIA GPU
- NVIDIA GPURuns the compiled kernels
Capabilities
GPU compilers34
NVCC compiles CUDA C++, NVRTC compiles at runtime, and nvJitLink links code for target GPUs through PTX.
Why it matters: Turns your source code into programs that run on specific NVIDIA GPU generations.
Limits: Needs a supported host compiler version; very new distribution compilers may require a compatibility package.
Math and parallel libraries23
cuBLAS, cuFFT, cuRAND, cuSOLVER, cuSPARSE, Thrust, CUB and libcu++ ship with the toolkit.
Why it matters: Most applications call tuned libraries instead of writing their own kernels.
Limits: Library releases contain known issues on specific GPUs; read the release notes for your architecture.
Tile programming model1
CUDA Tile IR, CUDA Tile C++ and cuTile Python let developers write kernels in terms of data tiles.
Why it matters: Opens kernel writing to Python developers and simplifies some C++ kernels.
Limits: A newer programming model; check which GPUs and features it supports in your release.
Profiling and debugging tools78
Nsight Systems, Nsight Compute, Compute Sanitizer, cuda-gdb and CUPTI are part of the toolkit.
Why it matters: Finds bottlenecks and memory errors without separate products.
Limits: Standalone tool downloads can be newer than the versions bundled with a toolkit release.
Driver compatibility rules2
Each release maps to a driver branch, and minor version compatibility lets CUDA 13.x applications run on drivers 580 and later.
Why it matters: Lets teams upgrade the toolkit without always upgrading cluster drivers.
Limits: New features and platforms in 13.4 need an R615 or later driver.
Broad platform coverage123
Targets embedded systems, workstations, data centers, clouds and supercomputers, and from 13.4 Windows on Arm for RTX Spark.
Why it matters: One programming model across the NVIDIA hardware range.
Limits: Rubin support in 13.4 is a preview; CUDA 14.0 will raise the minimum Arm architecture for ARM64-SBSA.
Practical use cases
A simulation or scientific code spends most of its time in loops that could run in parallel.
- Approach
- Move the hot loops into CUDA C++ kernels or library calls such as cuBLAS and cuFFT, then profile with Nsight.
- Role of NVIDIA CUDA Toolkit
- The toolkit provides the compiler, libraries and profilers for the port.
- Data, infrastructure and skills
- A CUDA-capable GPU, a supported compiler and developers comfortable with parallel programming.
- Type of benefit
- Faster computation
- Caveats
- Gains depend on how parallel the code is and on data movement between CPU and GPU.
- First step
- Profile the CPU version to find the hot loops, then port one of them.
Sources 1
Python developers need a custom GPU operation that existing libraries do not offer.
- Approach
- Write the operation as a tile kernel with cuTile Python.
- Role of NVIDIA CUDA Toolkit
- The toolkit supplies the CUDA Tile model and its compilation path.
- Data, infrastructure and skills
- A recent CUDA 13 toolkit and a GPU supported by CUDA Tile.
- Type of benefit
- Developer productivity
- Caveats
- A new programming model; check support and examples for your GPU.
- First step
- Work through the cuTile Python documentation examples.
Sources 1
A team must build and ship GPU software consistently across developer machines and clusters.
- Approach
- Build inside CUDA containers from the NGC catalog and keep the driver on hosts within the compatibility range.
- Role of NVIDIA CUDA Toolkit
- Containers fix the CUDA version used for builds, and NVIDIA's release notes state which driver branches each CUDA version needs.
- Data, infrastructure and skills
- A container runtime with GPU support and hosts with compatible drivers.
- Type of benefit
- Operational consistency
- Caveats
- The NVIDIA driver is no longer bundled with the toolkit, so each host needs a compatible driver installed separately (580 or later for CUDA 13.x applications).
- First step
- Pull a CUDA container from NGC that matches your target CUDA version.
Works with
Optional integration
- NVIDIA Nsight Developer ToolsNsight Compute, Nsight Systems and related tools ship with the CUDA Toolkit.
Complementary tools
- NVIDIA CUDA-X Data ScienceCUDA-X data science libraries are built on CUDA and require a matching CUDA version.
- NVIDIA cuOptcuOpt is built on CUDA and requires CUDA 12 or 13.
- NVIDIA JetsonNVIDIA documents CUDA upgrades for Jetson devices; the toolkit targets embedded systems.
Required by
- NVIDIA CUDA-X Data ScienceThe libraries need CUDA 12 or 13 and a matching driver; pip wheels must match the installed CUDA major version.
- NVIDIA cuOptcuOpt needs CUDA 12.0+ or 13.0+ and a compatible driver.
Optional integration for
- NVIDIA Nsight Developer ToolsNsight Systems, Nsight Compute and Compute Sanitizer ship with the CUDA Toolkit.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Check the GPU and system4
Run lspci | grep -i nvidia, confirm the GPU on NVIDIA's CUDA GPUs list, and check that the Linux distribution and gcc version are supported.
Check: The GPU appears in lspci output and gcc --version reports a supported compiler.
Install the driver2
Install an NVIDIA driver separately, since CUDA 13.4 no longer bundles it on Linux. Use R615 or later for 13.4 features.
Check: The installed driver version meets the branch listed in the release notes.
Install the toolkit4
Add NVIDIA's package repository for your distribution, install the cuda-toolkit package and add /usr/local/cuda-13.4/bin to PATH.
Check: nvcc is found on the PATH from the CUDA 13.4 install directory.
Verify with the samples4
Build deviceQuery and bandwidthTest from the CUDA samples repository on GitHub and run them.
Check: deviceQuery lists your GPU and bandwidthTest completes without errors.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA CUDA Toolkit as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA CUDA Toolkit product page (opens in a new tab)
- CUDA Toolkit 13.4 Update 1 release notes (opens in a new tab)
- CUDA Toolkit documentation hub (opens in a new tab)
- CUDA installation guide for Linux (opens in a new tab)
- NVIDIA CUDA-X for Data Science developer page (opens in a new tab)
- CUDA Toolkit end user license agreement (opens in a new tab)
- NVIDIA Nsight Compute product page (opens in a new tab)
- NVIDIA developer tools overview (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.