Skip to content

Accelerated analytics

Speed up dataframe processing, machine learning and graph analytics on GPUs, often without code changes to pandas, Polars, scikit-learn, NetworkX or Spark, so data teams can iterate faster on large datasets.

The problem

Data teams wait minutes or hours for joins, aggregations, feature engineering and model training on large tables, which limits how many ideas they can test in a day. Scaling out CPU clusters adds cost and operations work, and rewriting pipelines for a new framework is a project of its own.

The practical question is whether existing Python and Spark code can run faster on GPUs without a rewrite, and where the remaining bottlenecks are, such as storage reads, data transfer or single-threaded code.

The approach

Profile first: find which steps dominate runtime and how big the data is. Then try GPU libraries on those steps. NVIDIA's CUDA-X Data Science libraries, whose packages still carry the RAPIDS name, include cuDF for pandas and Polars, cuML for scikit-learn, UMAP and HDBSCAN, and cuGraph for NetworkX, each with a zero-code-change mode. Dask and the RAPIDS Accelerator for Apache Spark spread work across nodes. Installation is through conda, pip or Docker, with guides for Kubernetes, managed notebook platforms and the main clouds.

On shared clusters, Run:ai can give data scientists pooled or fractional GPUs instead of fixed machines; for local work, a desktop system such as DGX Spark is one option.

Small tables that already finish in seconds gain little from GPUs, and SQL engines on columnar storage may be enough for query-heavy work. GPUs help most with large in-memory transformations, iterative machine learning and graph algorithms.123

Conceptual architecture

Accelerated analytics: conceptual architectureApplications &solutionsModels & frameworksOperations &orchestrationAcceleratedcomputingData lake or warehouse (Parquet and tables): Source data for pipelinesData lake or warehouse(Parquet and tables)Notebooks and pipelines: Existing pandas, Polars, scikit-learn, NetworkX or Spark codeNotebooks and pipelinesGPU machine learning (cuML): Trains and applies classical machine learning modelsGPU machine learning (cuML)GPU graph analytics (cuGraph): Runs graph algorithms on large networksGPU graph analytics(cuGraph)Multi-node engines (Dask, Spark with RAPIDS Accelerator): Spread data and work across nodesMulti-node engines (Dask,Spark with RAPIDS…GPU scheduling (Run:ai): Shares GPUs among data scientists and jobsGPU scheduling (Run:ai)GPU dataframes (cuDF): Runs joins, aggregations and feature engineering on the GPUGPU dataframes (cuDF)NVIDIA GPUs (desktop, server or cloud): Execute the accelerated librariesNVIDIA GPUs (desktop, serveror cloud)
Diagram as a list
  1. Applications & solutions

    • Data lake or warehouse (Parquet and tables)Source data for pipelinesConnects to GPU dataframes (cuDF), Multi-node engines (Dask, Spark with RAPIDS Accelerator)
    • Notebooks and pipelinesExisting pandas, Polars, scikit-learn, NetworkX or Spark codeConnects to GPU dataframes (cuDF), GPU machine learning (cuML), GPU graph analytics (cuGraph)
  2. Models & frameworks

    • GPU machine learning (cuML)Trains and applies classical machine learning modelsConnects to NVIDIA GPUs (desktop, server or cloud)
    • GPU graph analytics (cuGraph)Runs graph algorithms on large networksConnects to NVIDIA GPUs (desktop, server or cloud)
  3. Operations & orchestration

    • Multi-node engines (Dask, Spark with RAPIDS Accelerator)Spread data and work across nodesConnects to NVIDIA GPUs (desktop, server or cloud)
    • GPU scheduling (Run:ai)Shares GPUs among data scientists and jobsConnects to NVIDIA GPUs (desktop, server or cloud)
  4. Accelerated computing

    • GPU dataframes (cuDF)Runs joins, aggregations and feature engineering on the GPUConnects to NVIDIA GPUs (desktop, server or cloud)
    • NVIDIA GPUs (desktop, server or cloud)Execute the accelerated libraries
Conceptual: one common way to arrange the parts, not a required design.1

Technologies and their roles

  • cuda-x-data-science1

    GPU data libraries

    cuDF, cuML, cuGraph, Dask integration and the Spark accelerator speed up familiar Python and Spark APIs.

  • run-ai4

    Shared GPU access

    Pools GPUs across teams with quotas and fractional allocation so analysts do not need dedicated machines.

  • dgx3

    Local or cluster hardware

    NVIDIA positions DGX Spark for data scientists working locally, and DGX systems for larger clusters.

What you need first1

  • A profiled pipeline that shows where time is spent
  • Linux or WSL2 with a supported NVIDIA GPU; current packages target CUDA 13
  • Data that fits GPU memory per partition, or Dask or Spark to spread it
  • Python data skills, plus Spark skills for the Spark accelerator
  • A test that compares CPU and GPU results on the same data

Risks and how to reduce them

Unsupported operations fall back to the CPU and erase the gain
Check profiler output for fallbacks and rewrite only those steps.
Small numerical differences between CPU and GPU results
Compare outputs on a reference dataset and agree tolerances.
Published speed-ups do not match your data
Benchmark with your own data sizes and queries before committing.
Dedicated GPUs sit idle between analyses
Share GPUs through a scheduler or use cloud instances for bursty work.

Related

Sources

  1. NVIDIA CUDA-X for Data Science (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA Run:ai (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA DGX Platform (opens in a new tab)NVIDIA · Vendor-reported
  4. The NVIDIA Run:ai Scheduler: Concepts and Principles (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)