Accelerated analytics
Speed up dataframe processing, machine learning and graph analytics on GPUs, often without code changes to pandas, Polars, scikit-learn, NetworkX or Spark, so data teams can iterate faster on large datasets.
The problem
Data teams wait minutes or hours for joins, aggregations, feature engineering and model training on large tables, which limits how many ideas they can test in a day. Scaling out CPU clusters adds cost and operations work, and rewriting pipelines for a new framework is a project of its own.
The practical question is whether existing Python and Spark code can run faster on GPUs without a rewrite, and where the remaining bottlenecks are, such as storage reads, data transfer or single-threaded code.
The approach
Profile first: find which steps dominate runtime and how big the data is. Then try GPU libraries on those steps. NVIDIA's CUDA-X Data Science libraries, whose packages still carry the RAPIDS name, include cuDF for pandas and Polars, cuML for scikit-learn, UMAP and HDBSCAN, and cuGraph for NetworkX, each with a zero-code-change mode. Dask and the RAPIDS Accelerator for Apache Spark spread work across nodes. Installation is through conda, pip or Docker, with guides for Kubernetes, managed notebook platforms and the main clouds.
On shared clusters, Run:ai can give data scientists pooled or fractional GPUs instead of fixed machines; for local work, a desktop system such as DGX Spark is one option.
Small tables that already finish in seconds gain little from GPUs, and SQL engines on columnar storage may be enough for query-heavy work. GPUs help most with large in-memory transformations, iterative machine learning and graph algorithms.123
Conceptual architecture
Diagram as a list
Applications & solutions
- Data lake or warehouse (Parquet and tables)Source data for pipelinesConnects to GPU dataframes (cuDF), Multi-node engines (Dask, Spark with RAPIDS Accelerator)
- Notebooks and pipelinesExisting pandas, Polars, scikit-learn, NetworkX or Spark codeConnects to GPU dataframes (cuDF), GPU machine learning (cuML), GPU graph analytics (cuGraph)
Models & frameworks
- GPU machine learning (cuML)Trains and applies classical machine learning modelsConnects to NVIDIA GPUs (desktop, server or cloud)
- GPU graph analytics (cuGraph)Runs graph algorithms on large networksConnects to NVIDIA GPUs (desktop, server or cloud)
Operations & orchestration
- Multi-node engines (Dask, Spark with RAPIDS Accelerator)Spread data and work across nodesConnects to NVIDIA GPUs (desktop, server or cloud)
- GPU scheduling (Run:ai)Shares GPUs among data scientists and jobsConnects to NVIDIA GPUs (desktop, server or cloud)
Accelerated computing
- GPU dataframes (cuDF)Runs joins, aggregations and feature engineering on the GPUConnects to NVIDIA GPUs (desktop, server or cloud)
- NVIDIA GPUs (desktop, server or cloud)Execute the accelerated libraries
Technologies and their roles
cuda-x-data-science1
GPU data libraries
cuDF, cuML, cuGraph, Dask integration and the Spark accelerator speed up familiar Python and Spark APIs.
run-ai4
Shared GPU access
Pools GPUs across teams with quotas and fractional allocation so analysts do not need dedicated machines.
dgx3
Local or cluster hardware
NVIDIA positions DGX Spark for data scientists working locally, and DGX systems for larger clusters.
What you need first1
- A profiled pipeline that shows where time is spent
- Linux or WSL2 with a supported NVIDIA GPU; current packages target CUDA 13
- Data that fits GPU memory per partition, or Dask or Spark to spread it
- Python data skills, plus Spark skills for the Spark accelerator
- A test that compares CPU and GPU results on the same data
Risks and how to reduce them
- Unsupported operations fall back to the CPU and erase the gain
- Check profiler output for fallbacks and rewrite only those steps.
- Small numerical differences between CPU and GPU results
- Compare outputs on a reference dataset and agree tolerances.
- Published speed-ups do not match your data
- Benchmark with your own data sizes and queries before committing.
- Dedicated GPUs sit idle between analyses
- Share GPUs through a scheduler or use cloud instances for bursty work.
Related
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.