Edge AI
Run AI models on devices close to where data is created, such as cameras, machines, robots and medical equipment, for low latency, limited connectivity or data that should not leave the site.
The problem
Sending every video frame or sensor reading to the cloud is often too slow, too costly in bandwidth or not allowed. Robots need decisions within milliseconds, remote sites may have weak connections, and some data must stay on the premises.
Edge devices bring their own constraints: limited power and memory, harsh environments, long service lives, and fleets that must be updated and secured remotely.
The approach
Choose the device class from the constraints: power, size, temperature, latency and certification needs. NVIDIA's Jetson modules range from Orin Nano to Jetson Thor and run the JetPack stack with CUDA, TensorRT and frameworks such as PyTorch, vLLM and SGLang; for industrial-grade systems with functional safety, NVIDIA points to its IGX platform. Above that, DeepStream handles video pipelines, Holoscan handles real-time sensor streams and Isaac ROS supports robots. Compact open models such as Cosmos3-Edge and Nemotron run on Jetson.
Models are usually trained in the data center, then reduced in size and precision to fit the device. Fleet management, secure boot, encryption and over-the-air updates need planning from the start, and Jetson Linux supports them.
If latency, bandwidth and privacy allow, cloud inference is easier to update and scale. Microcontroller-class or other low-power accelerators can suit very small tasks such as keyword spotting or simple classification.123
Conceptual architecture
Diagram as a list
Applications & solutions
- Cameras and sensorsProduce video and signal data on siteConnects to Edge module (Jetson or IGX)
- Pipeline (DeepStream, Holoscan or Isaac ROS)Handles video, sensor or robot data flowConnects to Optimized model (TensorRT, reduced precision), Local actions and alerts
- Local actions and alertsActs on results without a round trip to the cloud
Inference & runtime software
- Optimized model (TensorRT, reduced precision)Fits accuracy and latency targets on the device
Operations & orchestration
- Fleet management and OTA updatesDeploys models and patches securely across devicesConnects to Edge module (Jetson or IGX)
Accelerated computing
- Edge module (Jetson or IGX)Runs inference within the device power budgetConnects to Pipeline (DeepStream, Holoscan or Isaac ROS)
- Data center training and retrainingProduces and updates the models shipped to devicesConnects to Optimized model (TensorRT, reduced precision)
Technologies and their roles
jetson1
Edge compute
Modules and developer kits from Orin Nano to Thor run the JetPack stack for on-device AI.
deepstream4
Video pipelines
Streaming analytics toolkit that runs on Jetson as well as servers.
holoscan5
Sensor streaming
Real-time processing of high-bandwidth sensor data, including through Holoscan Sensor Bridge.
isaac6
Robot software
Isaac ROS packages accelerate ROS 2 pipelines on Jetson.
cosmos3
Compact world model
Cosmos3-Edge is sized for Jetson AGX Orin and Thor.
nemotron1
On-device language models
NVIDIA lists Nemotron among the open models that run on Jetson.
What you need first
- Power, size, environmental and latency limits for the device
- A trained model and a plan to quantize and test it on the target hardware
- Embedded Linux and CUDA or TensorRT skills
- A fleet update and monitoring plan
- The certification requirements of the target industry
Risks and how to reduce them
- The model is too large or slow for the device
- Benchmark on the target module early and adjust the model or the hardware choice.
- Device tampering or theft
- Use secure boot, disk encryption and signed over-the-air updates.
- Hardware availability and lifecycle1
- Check module availability and support dates; some Jetson products show a notify-me option rather than ordering.
- Safety-critical use1
- NVIDIA points to IGX for functional safety needs; certification of the final product remains your responsibility.
- Privacy of on-device video
- Keep raw data on the device, send only events and set retention limits.
Related
Sources
- NVIDIA Jetson embedded systems (product page) (opens in a new tab)
- NVIDIA JetPack (developer page) (opens in a new tab)
- NVIDIA Cosmos repository README (opens in a new tab)
- NVIDIA DeepStream SDK (developer page) (opens in a new tab)
- NVIDIA Holoscan SDK (developer page) (opens in a new tab)
- NVIDIA Isaac ROS (developer page) (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.