Skip to content

Kaohsiung City Government, with integrator Linker Vision · City government and public infrastructure · Taiwan (Kaohsiung City)

Kaohsiung City Government: city video agents and a digital twin built by Linker Vision

Linker Vision built a vision AI platform for Kaohsiung City Government on NVIDIA Metropolis, Omniverse and Cosmos, with video agents that flag events such as roadway flooding and alert the right agencies. No measured results with a stated baseline have been published.

Status
Scaling
Status as of
1 Jun 2026

The challenge

City bureaus in Kaohsiung ran separate systems from different integrators and vendors. NVIDIA's account gives a concrete example: when the Water Resources Bureau detected flooding, that information did not reach the Transportation Bureau automatically, even though floods disrupt traffic and public safety.

Conventional vision models detect cars, people and buildings, but struggle to describe an unusual situation such as an accident, a flooded road or a fallen tree, which is what responders need to know.1

What was implemented

Linker Vision, the integrator, organizes the work in three stages. For simulation, it converts aerial and satellite pictures into OpenUSD scenes, then assembles a city digital twin in NVIDIA Omniverse on OVX servers, and uses Cosmos to generate synthetic video of rare events such as flooding or damaged infrastructure. For training, it curates and labels real footage with NeMo Curator and nv-grounding-dino, then fine-tunes vision language models on real and synthetic data. For runtime, it relies on the VSS Blueprint from Metropolis, which handles video search and summarization, running Cosmos vision language models on DGX servers.

Agents built this way watch thousands of live camera feeds; in one example given by NVIDIA, they spot a flooded main road and alert the responsible agencies and affected residents with location, timing and suggested actions. Outputs feed a command center linked to a live Omniverse twin. NVIDIA says Linker Vision is integrating 30,000 camera streams in Kaohsiung and covers more than 300 scenarios across over ten domains.

In June 2026 Linker Vision said a 2026 scale-up of the Kaohsiung lighthouse infrastructure now supports traffic management, public infrastructure, environmental monitoring and flooding response, and that it is extending video reasoning agents to Taipei City and Taiwan's Ministry of Transportation and Communications on its AI-GRID platform, which uses the NemoClaw blueprint, the VSS Blueprint and Cosmos.12

What others can learn, and the limits

What other cities can take from it: the value here comes less from detection accuracy than from routing one bureau's alerts to the others, so the first task is agreeing which events each department needs to receive. Generating synthetic footage of rare events is a practical answer to having too few real flood or collapse videos for training. Limits: both published results come from NVIDIA's story about its partner, with no baseline, method or time period, and the June 2026 scale-up is described only by Linker Vision. Kaohsiung's camera density and city-wide twin represent a large investment; a city with a few hundred cameras and one traffic center can start with off-the-shelf video analytics at intersections before building a twin.

Explore a similar project

Use the Digital Twin Opportunity Lab with your own assumptions. Results are independent of this case.

Open the lab

Sources

  1. Linker Vision Taps Into Vision AI to Optimize City Operations (opens in a new tab)NVIDIA · Vendor-reported
  2. Linker Vision Unveils Application-Driven AI Grid for Agentic Video Reasoning at Scale (opens in a new tab)Linker Vision · Customer-reported

Fill out the form below to request your copy.

Name(Required)