How AI models get small enough to run on your phone
Beyond distillation, a handful of engineering tricks let AI that once needed a data centre run quietly inside a handset or laptop.
The problem: intelligence is heavy
A large language model is, at its core, a huge table of numbers, the weights, that get multiplied together billions of times to produce an answer. The biggest models run on racks of specialised chips in data centres, drawing serious amounts of power and needing fast memory to hold all those numbers at once. A phone has none of that headroom. It has limited memory, a battery to protect, and a processor built for efficiency rather than raw throughput.
Yet increasingly, phones, laptops and even some cars run AI features locally: transcription, photo editing, predictive text, voice assistants, translation. That only works because engineers have got very good at making models smaller without making them much less useful. Distillation, where a smaller model learns to mimic a larger one, is one well-known route to this. But it is only one tool among several, and understanding the others explains why the shrinking has kept going.
Quantisation: rounding the numbers
Most models are trained using high-precision numbers, essentially decimals carried out to many places, because that precision helps during learning. Once a model is trained, though, much of that precision is wasted. Quantisation reduces the number of bits used to store each weight, similar to rounding pound-and-pence prices to the nearest pound. The model takes up less memory and calculations run faster, because processors can move and multiply smaller numbers more efficiently.
The skill lies in deciding where precision can be sacrificed without the model’s answers noticeably degrading. Some parts of a network are more sensitive than others, so modern quantisation methods vary the precision across the model rather than applying one blunt cut everywhere. Done well, a model can be reduced to a quarter of its original size with barely measurable loss in quality.
Pruning: removing what isn’t earning its place
Training a large model tends to produce a lot of redundancy. Some connections between neurons barely influence the output at all. Pruning identifies and removes these low-impact connections, or sometimes whole chunks of the network, shrinking it directly.
This is closer to editing an over-long report: cutting sentences that add words but not meaning. Aggressive pruning can damage a model if done carelessly, so it is usually followed by a short further training step, letting the slimmed-down network adjust to compensate for what was removed.
Distillation, briefly, as part of the toolkit
Distillation trains a compact “student” model to reproduce the behaviour of a larger “teacher” model, learning not just correct answers but the teacher’s patterns of confidence across possible answers, which carries more information than a simple right-or-wrong label. It is powerful precisely because it can be combined with quantisation and pruning rather than replacing them. A model might first be distilled into a smaller architecture, then pruned, then quantised, each step removing a different kind of excess.
Better architectures, not just smaller ones
Alongside compressing existing models, researchers keep designing network architectures that are inherently more efficient, achieving similar results with fewer calculations per task. Techniques that let a model activate only the relevant portion of its parameters for a given input, rather than the whole network every time, mean a model can have a large total size on disk while only doing a fraction of that work for any single request. This is one reason model capability has kept climbing even as running costs, per query, have often fallen.
Hardware built for the job
Software tricks only go so far without hardware that can exploit them. Phone chips increasingly include dedicated neural processing units, silicon designed specifically for the repetitive multiply-and-add operations that AI models rely on, separate from the general-purpose processor. These units are efficient at exactly the low-precision, highly parallel arithmetic that quantised models use, which is part of why a heavily compressed model can run smoothly on a handset while an uncompressed version would not.
Why this matters beyond convenience
Running AI on the device itself, rather than sending data to a remote server, has real practical consequences. It can work without an internet connection, it reduces the amount of personal data leaving the device, and it removes the ongoing cost of server processing for every single request. For a company shipping millions of devices, that cost difference is significant, which is why so much engineering effort goes into shrinking models rather than simply relying on ever-bigger data centres.
What to watch for
Smaller is not automatically worse, but it is not automatically fine either. A heavily compressed model can behave subtly differently from its larger counterpart, particularly on unusual or edge-case inputs, so credible providers publish some form of evaluation showing how a compressed model performs against benchmarks, not just its size. If you are choosing a product or service that advertises “on-device AI”, it is reasonable to ask what evidence exists that shrinking the model has not quietly shrunk its reliability too.
For deeper technical background and current standards work on efficient AI, the UK’s AI Safety Institute and the Alan Turing Institute both publish accessible explainers, and the Ada Lovelace Institute covers the practical and policy implications of on-device and edge AI deployment.