Latency matters, and for some workloads it is decisive. But most edge-inference decisions are made on cost of data movement, privacy exposure, and what happens when the link drops. A camera that streams everything to the cloud has an operating cost that grows with success. A camera that decides locally and sends conclusions does not.
The bill the buyer is actually reading
Consider a deployment of ten thousand cameras. At even modest resolution, streaming continuously to the cloud means petabytes a month across the backhaul, a cloud ingest and storage bill that scales with every unit installed, and a GPU fleet somewhere that has to be sized for peak. The pitch deck says "real-time video analytics." The finance team says "our cost per site goes up every time we sell one." That is not a business; it is a subsidy program for a cloud vendor.
Move the inference to the device and the picture inverts. The unit cost is fixed at manufacture. Bandwidth drops to metadata — events, counts, bounding boxes, the occasional clip that a human needs to see. The cloud does aggregation and fleet management, which scale gently. Every additional unit improves the margin instead of eroding it. Buyers understand this arithmetic quickly, and they understand it better than they understand milliseconds.
Three questions that decide the architecture
When we help a team decide where inference should live, the conversation rarely starts with latency. It starts with three questions. Who pays for the bandwidth, and does that cost scale with volume? Who owns the data, and is there a regulatory or contractual reason it cannot leave the site — a hospital, a factory floor with trade secrets, a retail chain that has promised customers their faces are never uploaded? And what must the system do when the link is down: keep working, degrade gracefully, or stop?
Answer those three and the architecture largely designs itself. The latency requirement is usually met as a side effect of the decision to compute locally. Sell the economics; the milliseconds are a proof point.
Where the engineering gets hard
None of this makes edge inference easy. It moves the difficulty from the cloud bill to the product. Models must fit in a fixed power and thermal envelope, which means quantization, pruning, and compiler work that most ML teams have never had to do. The silicon choice becomes a product decision with a five-year horizon rather than a monthly cloud invoice that can be renegotiated. Updating a model on ten thousand devices in the field is a release-engineering problem, not a deployment script. And the observability that comes free in the cloud — logs, metrics, traces — has to be designed in, with a bandwidth budget of its own.
Edge inference trades a variable cloud cost for a fixed engineering cost. The companies that win are the ones that treat the second as a product investment rather than a tax.
What good looks like
The edge products we have seen succeed share a pattern. The cost model was written down before the architecture was chosen, with the customer's numbers in it. Inference runs where the data is generated, and the cloud is used for what it is good at: fleet management, aggregation across sites, and retraining. The model, the runtime, and the hardware were chosen together, by one team, against a real power budget. And the system was designed to keep doing its job with the network cable pulled out — because at some point, at some site, it will be.
Get the economics right and the latency story takes care of itself. Get it wrong and no benchmark will save the business.
Red Tomatoes Enterprises · Published · Updated