Skip to content

Vision AI & Smart Environments

Video analytics for cities, factories, stores and warehouses: the Metropolis platform, DeepStream pipelines, the video search and summarization blueprint, TAO model customization, and Jetson or RTX PRO hardware for deployment.

Technology profiles for this category are in research.

Overview

Many organizations have cameras but no one to watch them all. Vision AI turns video into structured events such as counts, safety alerts or answers to plain-language questions. This category covers NVIDIA tools for building those systems.

Metropolis is NVIDIA's vision AI application platform and partner ecosystem, from edge to cloud. The AI Blueprint for video search and summarization (VSS) builds agents that answer questions about live or recorded video. DeepStream is an open-source toolkit based on GStreamer for real-time pipelines in C/C++ or Python. TAO customizes vision foundation models, which can then be deployed as NIM. Deployment targets include Jetson Thor, DGX Spark and RTX PRO 6000 Blackwell GPUs.

Typical users are system integrators, city operators, manufacturers and retailers.

Video analytics raises privacy questions, so check local law on recording and retention first. A single-camera counting task may be handled by an off-the-shelf smart camera without custom AI.12

Problems it addresses

  • Too much video, too few people1

    The VSS blueprint lets operators query live or archived video in natural language.

  • Real-time multi-camera processing2

    DeepStream provides more than 40 hardware-accelerated GStreamer plug-ins for multi-camera pipelines.

  • Models that miss local conditions1

    TAO customizes vision foundation models with site data.

  • Traffic and transit monitoring1

    Metropolis is used for multi-camera tracking at transit corridors and intersections.

  • Manual inspection on production lines1

    Automated visual inspection is a listed Metropolis use in factories.

A typical workflow

  1. Define events and privacy rules

    List the events or questions that matter and the retention rules that apply.

  2. Build the stream pipeline2

    Ingest camera streams and run inference with DeepStream.

  3. Adapt the models1

    Customize vision models with TAO and serve them as NIM.

  4. Add search and summaries3

    Deploy the VSS blueprint for natural-language queries over video.

  5. Deploy at edge or center1

    Run on Jetson Thor near the cameras or on RTX PRO servers centrally.

Digital Twin Opportunity Lab

Find the first digital twin worth building for your site.

Open the lab

Next steps

  1. Run a privacy impact assessment for the cameras you plan to use.

  2. Start with recorded video from two or three cameras and one event type.

  3. Try the video-search-and-summarization code from the NVIDIA-AI-Blueprints GitHub organization on sample footage.

  4. Use the Digital Twin Opportunity Lab if you also want to simulate the site layout.

Sources

  1. NVIDIA Metropolis (product page) (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA DeepStream SDK (developer page) (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA AI Blueprints GitHub organization (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)