NVIDIA DeepStream SDK
DeepStream is NVIDIA's toolkit for building real-time video and multi-sensor analytics pipelines on GPUs and Jetson devices. It is based on GStreamer and part of Metropolis. Its source code has been on GitHub under Apache-2.0 since version 9.0, while the prebuilt runtime libraries stay under an NVIDIA license.123
Also known as DeepStream, DeepStream SDK
At a glance
- What is it?
- DeepStream is a software development kit for streaming analytics. A DeepStream application is a GStreamer pipeline built from NVIDIA plugins that decode camera streams, batch frames, run AI models, track objects across frames and send results to other systems. NVIDIA calls it an open source toolkit inside the Metropolis vision AI platform, although it notes that a few libraries remain closed. Version 9.1 is the current release in the documentation. The source moved into one GitHub repository with version 9.0; from 9.1 the install packages and Dockerfiles are published there too, while container images are still released through NGC. Developers write applications in C or C++, or in Python through the Service Maker API.123
- What does it do?
- DeepStream takes many live video streams (and, through extensions, other sensors such as lidar) and turns them into structured events in real time. It decodes and batches streams on the GPU, runs detection or classification models with TensorRT or through Dynamo-Triton, follows objects across frames and across cameras, applies rules such as line crossing or region counting, draws overlays, and publishes metadata to message brokers. It also ships reference applications, camera calibration tools and agent skills that let coding agents generate and profile pipelines from a plain-language request.1
- Who needs it?
- Teams that must analyze many camera streams continuously with low latency: traffic and smart city projects, retail and warehouse analytics, factory inspection and safety monitoring. It fits developers comfortable with Linux, GStreamer concepts and C++ or Python, who deploy on NVIDIA GPUs in servers, workstations or Jetson modules.
- What does it need?23
- An NVIDIA GPU (Turing or newer) or a Jetson Orin or Jetson Thor module
- For x86: Ubuntu 24.04, CUDA 13.2, TensorRT 10.16.x and driver 595 or newer, per the repository README
- For Jetson: JetPack 7.2 GA
- Git LFS for the sample media, and Git with submodule support to clone the DeepStream repository
- On DGX Spark and other Arm server (SBSA) systems: the DeepStream Docker container, since bare-metal install is not supported there
- A trained model in a format TensorRT or Dynamo-Triton can load, or a model trained with NVIDIA TAO
- What it is not
- DeepStream is not a finished video management system or a dashboard; it is the processing engine you build one with. It is not the same as Metropolis: Metropolis is the wider platform and partner program, and blueprints such as Video Search and Summarization sit on top of DeepStream. It is also not fully open source, because the prebuilt runtime libraries ship under an NVIDIA software license. Speech features are no longer part of it; NVIDIA points speech work to Riva.13
Availability and licensing. DeepStream 9.1 is the current release in the documentation. The GitHub repository's source code is Apache-2.0 and its documentation CC-BY-4.0, while the prebuilt runtime libraries and packages are released under NVIDIA's Software License Agreement for SDKs; NVIDIA notes that a few libraries remain closed. The repository is maintained but does not accept code contributions. DeepStream is also available as part of NVIDIA AI Enterprise, which adds validation, support and API stability. No price is stated for the SDK itself.12
The problem it solves
Running AI on one video file is easy. Running it on dozens of live camera feeds, around the clock, with low delay, is not: decoding, resizing and batching frames on the CPU quickly becomes the bottleneck, and every model, tracker and message format needs glue code.
Teams also need the same pipeline to work on a data center GPU and on a small edge device near the cameras.
DeepStream addresses this by keeping frames in GPU memory from decode to inference, providing ready plugins for each pipeline stage, and running the same pipeline design on x86 servers, DGX Spark and Jetson.
How it works
- Ingest. Source plugins such as Gst-nvurisrcbin and Gst-nvmultiurisrcbin read files, RTSP cameras or other inputs and decode them on the GPU.
- Batch. Gst-nvstreammux combines frames from many streams into one batch.
- Prepare. Gst-nvvideoconvert and Gst-nvdspreprocess prepare frames for the model, and the DeepStream Libraries (CV-CUDA, nvImageCodec, PyNvVideoCodec) offer lower-level GPU operations for the pre- and post-processing stages.
- Infer. Gst-nvinfer runs models with TensorRT; Gst-nvinferserver sends them to Dynamo-Triton, locally or over gRPC.
- Track and analyze. Gst-nvtracker follows objects; the multiview 3D tracking extension follows them across cameras; Gst-nvdsanalytics applies rules such as line crossing.
- Output. Gst-nvdsosd draws overlays, and Gst-nvmsgconv with Gst-nvmsgbroker sends metadata to brokers such as Kafka or MQTT.
Service Maker wraps these plugins in a C++ and Python API, and REST APIs change parameters while the pipeline runs.14
Diagram as a list
Applications & solutions
- Dashboards, alerts and VSS blueprintWhere events are used or searched
Inference & runtime software
- Decode and batch (nvurisrcbin, nvstreammux)GPU decoding and batching of many streamsConnects to Inference (nvinfer with TensorRT, nvinferserver)
- Inference (nvinfer with TensorRT, nvinferserver)Runs detection and classification modelsConnects to Tracking and analytics (nvtracker, MV3DT, nvdsanalytics)
- Dynamo-Triton serverOptional shared model server, local or remote over gRPCConnects to Inference (nvinfer with TensorRT, nvinferserver)
- Tracking and analytics (nvtracker, MV3DT, nvdsanalytics)Follows objects and applies rulesConnects to Overlays and messaging (nvdsosd, nvmsgbroker)
Operations & orchestration
- Overlays and messaging (nvdsosd, nvmsgbroker)Draws results and publishes metadataConnects to Dashboards, alerts and VSS blueprint
Accelerated computing
- NVIDIA GPUs, DGX Spark and JetsonHardware that runs the whole pipelineConnects to Decode and batch (nvurisrcbin, nvstreammux), Inference (nvinfer with TensorRT, nvinferserver)
Networking, power & facilities
- Cameras, RTSP streams and filesVideo and sensor inputsConnects to Decode and batch (nvurisrcbin, nvstreammux)
Capabilities
GPU plugins for every pipeline stage13
GStreamer plugins for decode, batching, preprocessing, TensorRT inference, Dynamo-Triton inference, tracking, analytics, overlays and messaging.
Why it matters: Frames stay on the GPU from decode to result, which is what makes many streams per device possible.
Limits: You still design the pipeline; dynamic resolution change is alpha quality and very high stream counts can hit known stability issues.
Service Maker API for C++ and Python34
An object-oriented layer over GStreamer for building DeepStream applications without writing raw GStreamer code.
Why it matters: Lowers the GStreamer learning curve and is NVIDIA's recommended route for Python.
Limits: The older Python bindings (pyds) are deprecated, so existing Python apps may need migration.
Model serving through TensorRT or Dynamo-Triton13
Gst-nvinfer runs models with TensorRT; Gst-nvinferserver uses Dynamo-Triton, including remote instances over gRPC. TAO models are supported.
Why it matters: Lets teams pick fast native inference or a shared model server.
Limits: Dynamo-Triton inside DeepStream supports a single GPU; TensorFlow, UFF and Caffe model support was removed in DeepStream 8.0.
Multi-camera tracking and calibration1
Multiview 3D tracking (MV3DT) extends the tracker across cameras, and AutoMagicCalib aligns several cameras to a floor plan.
Why it matters: Needed for counting and following people or vehicles across a whole site rather than per camera.
Limits: Calibration quality depends on camera placement and floor plan accuracy.
Agent skills and Inference Builder12
The repository ships skills for coding agents such as Claude Code and Cursor to generate, profile and evaluate pipelines, and its tools folder includes Inference Builder, which supplies templates agents use to produce consistent code.
Why it matters: Speeds up first prototypes and model onboarding.
Limits: Generated pipelines still need review and performance testing on the target hardware.
Same pipeline on edge and data center13
Packages target x86 GPUs, Arm servers including DGX Spark (via container), and Jetson Orin and Thor; deployment uses containers, Kubernetes and Helm.
Why it matters: One design can move from a lab GPU to devices near the cameras.
Limits: Jetson builds track a specific JetPack release; DLA is not supported on Jetson Thor.
Practical use cases
A city wants vehicle and pedestrian counts at busy intersections in real time.
- Approach
- Feed intersection cameras into a DeepStream pipeline with a detector, tracker and line-crossing rules, and publish counts to a message broker.
- Role of NVIDIA DeepStream SDK
- Runs decoding, detection, tracking and counting for all streams on the GPU.
- Data, infrastructure and skills
- Camera access, a detection model suited to local traffic, a GPU server or Jetson devices near the cameras.
- Type of benefit
- Faster operational insight
- Caveats
- Accuracy drops at night or in bad weather unless the model is trained for those conditions; privacy rules apply to video of public spaces.
- First step
- Run the deepstream-app sample on one recorded intersection video, then swap in your own model.
A warehouse needs to follow people and forklifts across several cameras to spot near misses.
- Approach
- Calibrate cameras to the floor plan with AutoMagicCalib, run multiview 3D tracking, and raise events when tracks come too close.
- Role of NVIDIA DeepStream SDK
- Provides cross-camera tracking and the rule engine.
- Data, infrastructure and skills
- Overlapping camera coverage, an accurate floor plan and a safety team to define rules.
- Type of benefit
- Improved safety monitoring
- Caveats
- Alerts need tuning to avoid false alarms; tracking does not replace physical safety measures.
- First step
- Run the MV3DT skill on recorded footage from two overlapping cameras.
A production line needs visual checks on every item without slowing the line.
- Approach
- Run a defect classification model in a DeepStream pipeline on a Jetson or GPU server at the line, sending results to the line control system.
- Role of NVIDIA DeepStream SDK
- Handles low-latency capture, inference and result messaging.
- Data, infrastructure and skills
- Labeled images of good and defective parts, a camera with suitable lighting, an integration point in the line control system.
- Type of benefit
- Consistent inspection coverage
- Caveats
- Rare defects need enough examples; validate on real production data before relying on it.
- First step
- Benchmark your model in a single-stream pipeline on the target device.
Works with
Optional integration
- NVIDIA Dynamo-TritonGst-nvinferserver serves models through Dynamo-Triton; 9.1 supports Triton 26.03 on x86 and 26.04 on Jetson.
- NVIDIA JetsonDeepStream 9.1 packages target Jetson Orin and Jetson Thor on JetPack 7.2.
- NVIDIA AI EnterpriseDeepStream is available as part of NVIDIA AI Enterprise with support and API stability.
- NVIDIA DGXThe documentation includes a DGX Spark setup; on DGX Spark DeepStream runs from its container.
Same family
- NVIDIA MetropolisDeepStream is the streaming analytics toolkit inside Metropolis and powers RT-CV in VSS.
Alternative approaches
- NVIDIA Holoscan SDKFor many-camera video analytics DeepStream is the closer fit; Holoscan targets low-latency sensor pipelines.
Optional integration for
- NVIDIA JetsonDeepStream is a supported SDK on JetPack.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Check the platform
Match your system to the 9.1 requirements: Ubuntu 24.04 with CUDA 13.2, TensorRT 10.16.x and driver 595 or newer, or JetPack 7.2 GA on Jetson.
Check: nvidia-smi or the JetPack version shows a supported setup.
Clone and build
Install Git LFS, clone github.com/NVIDIA/DeepStream with submodules, and run build/build.sh, which downloads the runtime packages and installs to /opt/nvidia/deepstream/deepstream-9.1/. On DGX Spark use the container instead.
Check: The install folder exists and the sample apps are built.
Run a reference pipeline
Start deepstream-app with one of the sample configuration files, for example source30_1080p_dec_infer-resnet_tiled_display.txt from the app's config folder.
Check: Annotated video appears with detected objects.
Bring your own model
Point the Gst-nvinfer configuration at your own model, or use the deepstream-import-vision-model skill with a coding agent.
Check: Your model's detections appear in the output and the pipeline holds its frame rate.
Connect outputs
Add Gst-nvmsgconv and Gst-nvmsgbroker to send events to your broker, or use the REST API to change settings at runtime.
Check: Events arrive in the downstream system in the expected format.
Decide on support
Use community support, or obtain DeepStream through NVIDIA AI Enterprise for validation, support and API stability.
Check: Support route is agreed before production.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA DeepStream SDK as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.