Run inference on the device, not in the cloud
Moving real-time YOLO detection from cloud servers to NVIDIA Jetson devices for a smart-city digital twin.
At PROD'AIR I architected and led the AI side of a digital-twin platform for urban services and incident prevention. It needed to detect and classify urban objects from city cameras in real time.
The cloud was the bottleneck
Inference ran on DigitalOcean cloud servers. Every frame had to travel off-site before anything was detected, which made the system too slow, too costly and not accurate enough for real-time use.
Moving the model to the street
I moved YOLO inference onto NVIDIA Jetson devices on site. The trade-off is real: you now have hardware to manage in the field. But the results were clear: detection accuracy up 60% versus the cloud setup, cloud costs down 30%, and deployment time cut in half.
The rest of the data had to flow too
The same platform streamed area-level energy consumption from more than 50 IoT sensors at 95% data fidelity, so leaks and anomalies surfaced automatically. And a Llama-2 RAG pipeline with OCR made internal archives searchable, so employees could answer their own questions.
The takeaway
- When latency matters, move the model to the data, not the data to the model.
- Edge AI trades cloud bills for field operations; plan for both.
- A digital twin is only as good as its pipelines: vision, sensors and documents.