For years the default picture of AI was a request travelling to a data centre, a big model running on someone else's hardware, and an answer travelling back. That picture is increasingly wrong. A growing share of inference now happens on the device in your hand — the phone unlocking with your face, live captions generated locally, a keyboard predicting your next word without a round trip. On-device AI is not merely the cloud made smaller; it is a different set of constraints and pay-offs, and it is worth understanding why the industry keeps pushing computation to the edge. This piece is the "why and how" companion to the case for small, efficient models.
Why move inference to the device
Four forces pull computation to the edge, and they reinforce each other.
- 1
Latency
A local prediction has no network round trip, so it can respond in the time it takes to draw the next frame. For anything interactive — camera effects, live transcription — that immediacy is the feature.
- 2
Privacy
Data that never leaves the device cannot be intercepted, logged, or breached in transit. For photos, health signals, and typing, keeping raw data local is a strong privacy posture.
- 3
Offline capability
On-device models keep working with no connection — on a plane, underground, or in a dead zone. The capability is not contingent on a network.
- 4
Cost
Inference on the user's own hardware is inference you do not pay a server bill for. At scale, moving work to the edge can be the difference between a viable feature and an unaffordable one.
The cloud does not disappear
On-device inference does not replace the data centre; it complements it. Training still overwhelmingly happens in the cloud, and the hardest requests can still be escalated to a large server-side model. The pattern that is winning is hybrid: handle the common case locally, fall back to the cloud for the rest.
The constraint: a device is not a data centre
The catch is obvious the moment you state it: a phone has a fraction of the memory, compute, and — critically — the energy budget of a server rack, and it has to spend that budget on everything else the device does. A model that runs comfortably in the cloud may be far too large to load into a phone's memory, too slow to run at an acceptable frame rate, or too power-hungry to use without draining the battery. On-device AI is fundamentally an exercise in fitting a capable model into a small envelope.
Making a model fit
Three techniques do most of the shrinking, and they compose:
- Quantization stores and computes the model's numbers in lower precision — for instance, small integers instead of full-precision floating point. This is the highest-leverage trick: it shrinks memory and speeds up arithmetic, often with little quality loss, and much on-device hardware is built to run low-precision math especially fast.
- Pruning removes weights or whole structures that contribute little, producing a smaller, sparser model that does nearly the same job.
- Distillation trains a compact "student" model to imitate a larger "teacher," capturing much of the behaviour in a package small enough to deploy.
These are covered as a general efficiency strategy in the small-models piece; on the edge they stop being optional optimisations and become the price of admission.
The runtime and the silicon
Shrinking the model is half the story; running it efficiently is the other half. Two things make local inference fast enough to feel instant.
The first is a dedicated inference runtime — a library built to execute models efficiently on constrained hardware, handling quantized math and mapping operations onto whatever accelerators the device has. Google's on-device runtime (recently renamed from TensorFlow Lite to LiteRT), Apple's Core ML, and the cross-platform ONNX Runtime are prominent examples.
The second is dedicated hardware. Modern phones and laptops increasingly ship a neural processing unit (NPU) — silicon specialised for the matrix-heavy math of neural networks — alongside the CPU and GPU. Offloading inference to an NPU is both faster and far more power-efficient than doing the same work on a general-purpose CPU, which is what makes always-on features like face unlock or live translation practical on battery.
Pros
- Instant, interactive responses with no network dependency.
- Strong privacy: raw data can stay on the device.
- Works offline and reduces server inference cost.
Cons
- Tight memory, compute, and battery limits cap model size and capability.
- You must ship and update models across many diverse devices.
- Fragmented hardware: accelerators and their capabilities vary widely across the device landscape.
Practical takeaway
On-device AI is worth reaching for whenever latency, privacy, offline use, or per-inference cost matters — which, increasingly, is most consumer features. The workflow is: pick or design a model small enough to have a chance, shrink it with quantization first (then pruning and distillation as needed), and run it through a dedicated runtime that can exploit the device's NPU. Keep the cloud in reserve for the hard tail of requests. The mental shift is from "how capable is the biggest model" to "how much capability can I fit in this envelope" — and that constraint, handled well, is where a lot of the most-used AI now lives.
Sources & Further Reading
- 01LiteRT (formerly TensorFlow Lite) — GoogleGoogle's on-device inference runtime for mobile and edge.
- 02TensorFlow Lite is now LiteRT — Google Developers BlogThe announcement of the runtime's rename.
- 03Core ML — Apple Developer DocumentationApple's framework for on-device model inference.
- 04ONNX Runtime — ONNX Runtime projectA cross-platform runtime for running models on many targets, including the edge.
Editorial note — A conceptual explainer of on-device inference and the techniques and runtimes behind it. Runtimes and hardware are described qualitatively; no benchmark speeds, memory figures, or battery numbers are quoted.


