Data Science & Analytics
GPU acceleration for dataframes, machine learning, graph analytics and optimization: CUDA-X Data Science (formerly RAPIDS) with cuDF, cuML and cuGraph, the RAPIDS Accelerator for Apache Spark, and the cuOpt decision optimization engine.
Technology profiles for this category are in research.
Overview
Data teams often spend more time waiting on joins, feature engineering and model training than on analysis. This category covers NVIDIA libraries that run familiar Python and Spark work on GPUs.
CUDA-X Data Science is the current name for the libraries long known as RAPIDS, and some parts, such as the RAPIDS Accelerator for Apache Spark, keep the old name. cuDF speeds up dataframe work and offers zero-code-change acceleration for pandas, Polars and Spark. cuML accelerates scikit-learn style algorithms, UMAP and HDBSCAN. cuGraph accelerates NetworkX graph analytics. cuOpt is an open-source engine for vehicle routing, linear programming and mixed-integer programming, with MILP support in beta.
The audience is data scientists and data engineers whose jobs take minutes or hours on CPUs.
For small tables that already finish in seconds, a GPU adds setup cost for little gain. Check that your operations are supported before migrating a pipeline.12
Problems it addresses
Slow dataframe pipelines1
Joins and group-bys on large tables take long on CPUs. cuDF runs pandas and Polars code on GPUs without rewrites.
Long training on tabular data1
cuML accelerates scikit-learn style algorithms, UMAP and HDBSCAN.
Costly Spark clusters1
The RAPIDS Accelerator for Apache Spark moves Spark data processing onto GPUs.
Graph analysis at scale1
cuGraph accelerates NetworkX workloads.
Routing and planning decisions2
cuOpt targets routing and linear programs with millions of variables and constraints.
A typical workflow
Profile the CPU pipeline
Time each stage and find the slowest operations.
Turn on GPU acceleration1
Load the cuDF pandas accelerator or Polars GPU engine and rerun the unchanged code.
Train models1
Move scikit-learn style training to cuML.
Scale out1
Use Dask or the RAPIDS Accelerator for Apache Spark for multi-node jobs.
Optimize decisions2
Feed results into cuOpt for routing or planning problems.
NVIDIA Solution Architect
Describe your project and get an explainable architecture.
Next steps
Run your slowest notebook with the cuDF pandas accelerator on one GPU and compare wall-clock time.
Check the supported-operations list for your libraries before planning a migration.
For routing or scheduling, install cuOpt from PIP, Conda or a container and test it on a sample instance.
Use the NVIDIA Solution Architect if analytics is one part of a wider AI project.
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.