Skip to content

Edge AI

Run AI models on devices close to where data is created, such as cameras, machines, robots and medical equipment, for low latency, limited connectivity or data that should not leave the site.

The problem

Sending every video frame or sensor reading to the cloud is often too slow, too costly in bandwidth or not allowed. Robots need decisions within milliseconds, remote sites may have weak connections, and some data must stay on the premises.

Edge devices bring their own constraints: limited power and memory, harsh environments, long service lives, and fleets that must be updated and secured remotely.

The approach

Choose the device class from the constraints: power, size, temperature, latency and certification needs. NVIDIA's Jetson modules range from Orin Nano to Jetson Thor and run the JetPack stack with CUDA, TensorRT and frameworks such as PyTorch, vLLM and SGLang; for industrial-grade systems with functional safety, NVIDIA points to its IGX platform. Above that, DeepStream handles video pipelines, Holoscan handles real-time sensor streams and Isaac ROS supports robots. Compact open models such as Cosmos3-Edge and Nemotron run on Jetson.

Models are usually trained in the data center, then reduced in size and precision to fit the device. Fleet management, secure boot, encryption and over-the-air updates need planning from the start, and Jetson Linux supports them.

If latency, bandwidth and privacy allow, cloud inference is easier to update and scale. Microcontroller-class or other low-power accelerators can suit very small tasks such as keyword spotting or simple classification.123

Conceptual architecture

Edge AI: conceptual architectureApplications &solutionsInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingCameras and sensors: Produce video and signal data on siteCameras and sensorsPipeline (DeepStream, Holoscan or Isaac ROS): Handles video, sensor or robot data flowPipeline (DeepStream,Holoscan or Isaac ROS)Local actions and alerts: Acts on results without a round trip to the cloudLocal actions and alertsOptimized model (TensorRT, reduced precision): Fits accuracy and latency targets on the deviceOptimized model (TensorRT,reduced precision)Fleet management and OTA updates: Deploys models and patches securely across devicesFleet management and OTAupdatesEdge module (Jetson or IGX): Runs inference within the device power budgetEdge module (Jetson or IGX)Data center training and retraining: Produces and updates the models shipped to devicesData center training andretraining
Diagram as a list
  1. Applications & solutions

    • Cameras and sensorsProduce video and signal data on siteConnects to Edge module (Jetson or IGX)
    • Pipeline (DeepStream, Holoscan or Isaac ROS)Handles video, sensor or robot data flowConnects to Optimized model (TensorRT, reduced precision), Local actions and alerts
    • Local actions and alertsActs on results without a round trip to the cloud
  2. Inference & runtime software

    • Optimized model (TensorRT, reduced precision)Fits accuracy and latency targets on the device
  3. Operations & orchestration

    • Fleet management and OTA updatesDeploys models and patches securely across devicesConnects to Edge module (Jetson or IGX)
  4. Accelerated computing

    • Edge module (Jetson or IGX)Runs inference within the device power budgetConnects to Pipeline (DeepStream, Holoscan or Isaac ROS)
    • Data center training and retrainingProduces and updates the models shipped to devicesConnects to Optimized model (TensorRT, reduced precision)
Conceptual: one common way to arrange the parts, not a required design.4

Technologies and their roles

  • jetson1

    Edge compute

    Modules and developer kits from Orin Nano to Thor run the JetPack stack for on-device AI.

  • deepstream4

    Video pipelines

    Streaming analytics toolkit that runs on Jetson as well as servers.

  • holoscan5

    Sensor streaming

    Real-time processing of high-bandwidth sensor data, including through Holoscan Sensor Bridge.

  • isaac6

    Robot software

    Isaac ROS packages accelerate ROS 2 pipelines on Jetson.

  • cosmos3

    Compact world model

    Cosmos3-Edge is sized for Jetson AGX Orin and Thor.

  • nemotron1

    On-device language models

    NVIDIA lists Nemotron among the open models that run on Jetson.

What you need first

  • Power, size, environmental and latency limits for the device
  • A trained model and a plan to quantize and test it on the target hardware
  • Embedded Linux and CUDA or TensorRT skills
  • A fleet update and monitoring plan
  • The certification requirements of the target industry

Risks and how to reduce them

The model is too large or slow for the device
Benchmark on the target module early and adjust the model or the hardware choice.
Device tampering or theft
Use secure boot, disk encryption and signed over-the-air updates.
Hardware availability and lifecycle1
Check module availability and support dates; some Jetson products show a notify-me option rather than ordering.
Safety-critical use1
NVIDIA points to IGX for functional safety needs; certification of the final product remains your responsibility.
Privacy of on-device video
Keep raw data on the device, send only events and set retention limits.

Related

Sources

  1. NVIDIA Jetson embedded systems (product page) (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA JetPack (developer page) (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA Cosmos repository README (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA DeepStream SDK (developer page) (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA Holoscan SDK (developer page) (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA Isaac ROS (developer page) (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)