NVIDIA CUDA-X Data Science
NVIDIA CUDA-X Data Science, known until August 2026 as RAPIDS, is a collection of open source GPU libraries for data science: cuDF for dataframes, cuML for machine learning and cuGraph for graph analytics. Several of them can speed up existing pandas, scikit-learn or NetworkX code without code changes.12
Also known as RAPIDS, NVIDIA RAPIDS, CUDA-X for Data Science, CUDA-X libraries for data science
At a glance
- What is it?
- CUDA-X Data Science is NVIDIA's set of open source Python and C++ libraries (the core libraries cuDF, cuML and cuGraph are Apache 2.0 licensed) that run common data science work on NVIDIA GPUs. The main members are cuDF (dataframes, with accelerators for pandas, Polars, Dask and Apache Spark), cuML (machine learning with a scikit-learn style API) and cuGraph (graph analytics), plus cuxfilter for dashboards and Dask-based scale-out. NVIDIA moved the RAPIDS brand to CUDA-X beginning 11 August 2026 and says library functionality did not change. Package names, the conda metapackage (rapids) and parts of the documentation still use RAPIDS, and the Spark accelerator keeps the name RAPIDS Accelerator for Apache Spark.145
- What does it do?
- cuDF loads, filters, joins and aggregates tabular data on the GPU through a pandas-like API, and its cudf.pandas mode runs unmodified pandas code on the GPU, falling back to pandas on the CPU for operations it does not support. cudf-polars provides a GPU engine for Polars, dask-cudf scales dataframes across GPUs, and a plugin accelerates Apache Spark. cuML trains and runs clustering, regression, classification, dimensionality reduction and nearest-neighbor models, and cuml.accel runs existing scikit-learn, UMAP and HDBSCAN code on the GPU with CPU fallback. cuGraph runs graph algorithms such as PageRank on GPU dataframes and offers a zero-code-change backend for NetworkX.1678
- Who needs it?
- Data scientists and engineers whose pandas, Polars, scikit-learn, NetworkX, Dask or Spark jobs are slow on CPUs and who have access to NVIDIA GPUs, locally, in a cluster or in the cloud. It suits teams that want to keep their existing Python code and tools while moving the heavy work to GPUs.1
- What does it need?3
- An NVIDIA GPU of the Volta generation or newer (compute capability 7.0 or higher)
- Linux with glibc 2.28 or newer (for example Ubuntu 20.04+, Rocky, Alma or RHEL 8+, Debian 10+) or Windows 11 with WSL2
- CUDA 12 with driver 525.60.13 or newer, or CUDA 13 with driver 580.65.06 or newer
- A conda (for example miniforge), pip or Docker environment with a supported Python version
- What it is not
- CUDA-X Data Science is not a new product line with new features: it is the RAPIDS libraries under a new brand, with the same functionality. It is not a data platform, database or notebook service, but libraries you install in your own environment. It requires an NVIDIA GPU (Volta or newer), and the zero-code-change modes do not move every operation to the GPU: unsupported operations fall back to the CPU, which can limit the gain.23
Availability and licensing. The core libraries (cuDF, cuML and cuGraph) are open source under the Apache 2.0 license, and the libraries are free to install with conda, pip or Docker. The current release line is 26.08. NVIDIA moved the RAPIDS brand to CUDA-X beginning 11 August 2026; packages and some documentation still carry the RAPIDS name.78
The problem it solves
Dataframe processing, feature engineering, model training and graph analysis on large datasets can take minutes or hours on CPUs, which slows down iteration and makes some analyses impractical. Rewriting pipelines for a new engine is costly and risky.
CUDA-X Data Science runs the same kinds of operations on NVIDIA GPUs through APIs that mirror pandas, scikit-learn and NetworkX, and through accelerator modes that need no code changes, so teams can test GPU speed on their existing code first.1
How it works
The collection is built in layers on CUDA:
- cuDF: libcudf is a CUDA C++ library with Apache Arrow compatible data structures; pylibcudf exposes it to Python; the cudf package offers a pandas-like API; cudf.pandas, cudf-polars, dask-cudf and the Spark plugin plug the engine into existing tools. A Velox extension and Sirius, a GPU-native SQL engine with DuckDB extensions, are also listed.
- cuML: GPU estimators with fit, predict and transform methods; cuml.accel intercepts scikit-learn, UMAP and HDBSCAN calls; cuml.dask distributes selected algorithms across GPUs and nodes.
- cuGraph: graph algorithms operating on cuDF dataframes, with a NetworkX accelerator.
Data stays on the GPU between steps, so a cuDF table can feed cuML or cuGraph without copying back to the CPU. The current release line is 26.08, installed with conda, pip (wheels matched to the CUDA major version, such as -cu13) or Docker images.3478
Diagram as a list
Applications & solutions
- Existing Python code (pandas, Polars, scikit-learn, NetworkX)Data preparation, analytics and model trainingConnects to Zero-code-change modes (cudf.pandas, cuml.accel, NetworkX backend), cuDF (libcudf, pylibcudf, cudf-polars), cuML, cuGraph
- Zero-code-change modes (cudf.pandas, cuml.accel, NetworkX backend)Route supported calls to the GPU, fall back to CPU otherwiseConnects to cuDF (libcudf, pylibcudf, cudf-polars), cuML, cuGraph
Operations & orchestration
- Dask and Apache Spark integrationsDistribute work across GPUs and nodesConnects to cuDF (libcudf, pylibcudf, cudf-polars), cuML
Accelerated computing
- cuDF (libcudf, pylibcudf, cudf-polars)GPU dataframes and SQL-style operationsConnects to CUDA 12 or 13 and NVIDIA driver
- cuMLGPU machine learning algorithmsConnects to cuDF (libcudf, pylibcudf, cudf-polars), CUDA 12 or 13 and NVIDIA driver
- cuGraphGPU graph algorithms on cuDF dataConnects to cuDF (libcudf, pylibcudf, cudf-polars), CUDA 12 or 13 and NVIDIA driver
- CUDA 12 or 13 and NVIDIA driverRuntime the libraries are built onConnects to NVIDIA GPU (Volta or newer)
- NVIDIA GPU (Volta or newer)Executes the work
Capabilities
GPU dataframes (cuDF)14
A pandas-like dataframe library on the GPU, built on libcudf with Arrow-compatible data structures.
Why it matters: Speeds up loading, joining, grouping and cleaning large tables.
Limits: Scale-out goes through dask-cudf, a GPU backend for Dask DataFrames.
Zero-code-change pandas acceleration6
cudf.pandas runs existing pandas code on the GPU and falls back to pandas for unsupported operations.
Why it matters: Teams can test GPU speed on current notebooks and scripts without rewriting them.
Limits: Operations that fall back to the CPU run at CPU speed and may involve data transfers.
GPU machine learning (cuML)7
scikit-learn style estimators for clustering, regression, classification, dimensionality reduction and nearest neighbors, plus cuml.accel for unmodified scikit-learn, UMAP and HDBSCAN code. NVIDIA states speedups of up to 50x over scikit-learn on representative benchmarks.
Why it matters: Shortens model training and tuning loops.
Limits: Not every estimator or parameter is accelerated; speedups depend on algorithm, data and hardware.
Graph analytics (cuGraph)1
GPU graph algorithms such as PageRank that work on cuDF dataframes, with an accelerator for NetworkX.
Why it matters: Makes analysis of large graphs practical without specialized graph software.
Limits: Check which NetworkX algorithms the accelerator covers before relying on it.
Scale-out with Dask and Spark1
dask-cudf and cuml.dask spread work across GPUs and nodes; a plugin accelerates Apache Spark jobs.
Why it matters: Extends GPU acceleration to multi-GPU and cluster workloads.
Limits: Cluster setup and data partitioning add operational work.
Open source and broad packaging14
Libraries are Apache 2.0 licensed and installable with conda, pip or Docker, with guides for Kubernetes, Databricks, Colab and major clouds.
Why it matters: Fits existing Python environments and cloud platforms.
Limits: pip wheels must match the installed CUDA major version (cu12 or cu13).
Practical use cases
Analysts run pandas notebooks on large tables and wait minutes for each group-by or join.
- Approach
- Load cudf.pandas in the notebook so supported pandas operations run on the GPU.
- Role of NVIDIA CUDA-X Data Science
- cuDF executes the dataframe work; pandas handles anything unsupported.
- Data, infrastructure and skills
- A Volta or newer NVIDIA GPU and a cuDF installation.
- Type of benefit
- Faster analysis
- Caveats
- Gains shrink if many operations fall back to the CPU.
- First step
- Add %load_ext cudf.pandas at the top of one slow notebook and compare run times.
Sources 6
Model training and hyperparameter searches in scikit-learn take too long to iterate.
- Approach
- Run the existing scripts through cuml.accel, or port key steps to cuML estimators.
- Role of NVIDIA CUDA-X Data Science
- cuML runs supported estimators on the GPU.
- Data, infrastructure and skills
- GPU memory large enough for the training data.
- Type of benefit
- Shorter training cycles
- Caveats
- Check the cuml.accel compatibility list; unsupported estimators run on the CPU.
- First step
- Run one training script through the cuml.accel module and compare accuracy and time.
Sources 7
A data engineering team runs large Apache Spark or Dask jobs on CPU clusters.
- Approach
- Add the GPU accelerator for Spark or switch Dask dataframes to dask-cudf on GPU nodes.
- Role of NVIDIA CUDA-X Data Science
- CUDA-X libraries execute the heavy dataframe operations on GPUs across the cluster.
- Data, infrastructure and skills
- GPU nodes in the cluster and the matching plugin or packages.
- Type of benefit
- Higher throughput
- Caveats
- Cluster tuning and data partitioning still matter; not all operations are accelerated.
- First step
- Test one representative job on a single GPU node before changing the whole cluster.
Sources 4
Works with
Requires
- NVIDIA CUDA ToolkitThe libraries need CUDA 12 or 13 and a matching driver; pip wheels must match the installed CUDA major version.
Complementary tools
- NVIDIA Dynamo-TritonDynamo-Triton lists RAPIDS FIL among the frameworks it can serve.
- NVIDIA CUDA ToolkitCUDA-X data science libraries are built on CUDA and require a matching CUDA version.
Same family
- NVIDIA cuOptcuOpt is a CUDA-X library and follows the RAPIDS release schedule.
Optional integration for
- NVIDIA Dynamo-TritonThe RAPIDS FIL backend serves tree-based models from the CUDA-X Data Science (RAPIDS) stack.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Check the system3
Confirm a Volta or newer GPU, a Linux distribution with glibc 2.28 or newer (or Windows 11 with WSL2) and a driver of 580.65.06 or newer for CUDA 13.
Check: Your setup matches the system requirements in the installation guide.
Install the libraries1
Use the install selector: for example create a conda environment with the rapids=26.08 metapackage, or pip install the -cu13 wheels such as cudf-cu13 and cuml-cu13 from pypi.nvidia.com.
Check: import cudf and import cuml succeed in Python.
Accelerate existing pandas code6
Run a pandas script with python -m cudf.pandas, or add %load_ext cudf.pandas in Jupyter before importing pandas.
Check: The script produces the same results and supported operations run on the GPU.
Accelerate scikit-learn code7
Run an existing scikit-learn script through the cuml.accel module.
Check: The model trains successfully and you can compare its run time with the CPU run.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA CUDA-X Data Science as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA CUDA-X for Data Science developer page (opens in a new tab)
- NVIDIA Technical Blog: Accelerating Large-Scale Data Analytics with GPU-Native Velox and NVIDIA cuDF (opens in a new tab)
- CUDA-X Data Science installation guide (opens in a new tab)
- NVIDIA/cudf GitHub repository (README) (opens in a new tab)
- NVIDIA CUDA-X libraries for data science documentation hub (opens in a new tab)
- cuDF documentation: cudf.pandas (opens in a new tab)
- NVIDIA/cuml GitHub repository (README) (opens in a new tab)
- cuGraph GitHub repository (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.