Vision AI & Smart Environments
Video analytics for cities, factories, stores and warehouses: the Metropolis platform, DeepStream pipelines, the video search and summarization blueprint, TAO model customization, and Jetson or RTX PRO hardware for deployment.
Technology profiles for this category are in research.
Overview
Many organizations have cameras but no one to watch them all. Vision AI turns video into structured events such as counts, safety alerts or answers to plain-language questions. This category covers NVIDIA tools for building those systems.
Metropolis is NVIDIA's vision AI application platform and partner ecosystem, from edge to cloud. The AI Blueprint for video search and summarization (VSS) builds agents that answer questions about live or recorded video. DeepStream is an open-source toolkit based on GStreamer for real-time pipelines in C/C++ or Python. TAO customizes vision foundation models, which can then be deployed as NIM. Deployment targets include Jetson Thor, DGX Spark and RTX PRO 6000 Blackwell GPUs.
Typical users are system integrators, city operators, manufacturers and retailers.
Video analytics raises privacy questions, so check local law on recording and retention first. A single-camera counting task may be handled by an off-the-shelf smart camera without custom AI.12
Problems it addresses
Too much video, too few people1
The VSS blueprint lets operators query live or archived video in natural language.
Real-time multi-camera processing2
DeepStream provides more than 40 hardware-accelerated GStreamer plug-ins for multi-camera pipelines.
Models that miss local conditions1
TAO customizes vision foundation models with site data.
Traffic and transit monitoring1
Metropolis is used for multi-camera tracking at transit corridors and intersections.
Manual inspection on production lines1
Automated visual inspection is a listed Metropolis use in factories.
A typical workflow
Define events and privacy rules
List the events or questions that matter and the retention rules that apply.
Build the stream pipeline2
Ingest camera streams and run inference with DeepStream.
Adapt the models1
Customize vision models with TAO and serve them as NIM.
Add search and summaries3
Deploy the VSS blueprint for natural-language queries over video.
Deploy at edge or center1
Run on Jetson Thor near the cameras or on RTX PRO servers centrally.
Digital Twin Opportunity Lab
Find the first digital twin worth building for your site.
Next steps
Run a privacy impact assessment for the cameras you plan to use.
Start with recorded video from two or three cameras and one event type.
Try the video-search-and-summarization code from the NVIDIA-AI-Blueprints GitHub organization on sample footage.
Use the Digital Twin Opportunity Lab if you also want to simulate the site layout.
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.