← All posts
· 3 min read

Run inference on the device, not in the cloud

Moving real-time YOLO detection from cloud servers to NVIDIA Jetson devices for a smart-city digital twin.

Computer visionEdge AIMLOps

At PROD'AIR I architected and led the AI side of a digital-twin platform for urban services and incident prevention. It needed to detect and classify urban objects from city cameras in real time.

The cloud was the bottleneck

Inference ran on DigitalOcean cloud servers. Every frame had to travel off-site before anything was detected, which made the system too slow, too costly and not accurate enough for real-time use.

Moving the model to the street

I moved YOLO inference onto NVIDIA Jetson devices on site. The trade-off is real: you now have hardware to manage in the field. But the results were clear: detection accuracy up 60% versus the cloud setup, cloud costs down 30%, and deployment time cut in half.

The rest of the data had to flow too

The same platform streamed area-level energy consumption from more than 50 IoT sensors at 95% data fidelity, so leaks and anomalies surfaced automatically. And a Llama-2 RAG pipeline with OCR made internal archives searchable, so employees could answer their own questions.

The takeaway

  • When latency matters, move the model to the data, not the data to the model.
  • Edge AI trades cloud bills for field operations; plan for both.
  • A digital twin is only as good as its pipelines: vision, sensors and documents.