Why Your Phone Crashes Trying to Run AI — And the Routing Trick That Fixes It
Most talk about "AI on your phone" skips an awkward hardware fact: a flagship phone running a small language model to generate text will crash after a handful of queries. Not throttle — crash. The software that runs the model on the graphics chip freezes or dies, and it happens even when the phone is cold to the touch.
That fact changes how you should think about where AI actually runs. Here is the mechanism, then the fix, then the portable idea buried in the fix — which is the most valuable part.
The dream: run it locally
A small language model is a compressed version of a big chatbot — shrunk enough to fit on a phone and run without the internet. Local execution has three real advantages, not marketing ones:
- Privacy — your data never leaves the device.
- Offline — it works on a plane or in a tunnel.
- Free — no per-query fee to a cloud provider.
So the obvious design is: run on the phone by default, and escalate to something bigger only when you must.
Why the dream breaks
Phones run AI math on the GPU — the graphics processor, which also does the matrix multiplication that neural networks are made of. To use it, the model's math must be compiled into GPU instructions through a toolchain called OpenCL (a standard for writing code that runs on graphics hardware).
Here is why sustained generation fails where a single query survives. Text generation is autoregressive: the model produces one token (a word-fragment), feeds it back in, produces the next, and repeats — often hundreds of times per answer. Each step launches fresh GPU work. On a phone, that means hundreds of consecutive OpenCL command dispatches, each allocating and freeing GPU memory buffers. The mobile GPU driver — the low-level software translating those commands into hardware operations — is tuned for the short, bursty workloads of rendering graphics, not for this relentless, long-running compute loop. Memory fragments, buffers leak, the driver's internal state corrupts, and the runtime wedges.
Two things make this important:
- It is not heat throttling. Everyone expects a hot phone to slow down. This is a full failure of the inference runtime, not a slowdown.
- It happens cold. The crash recurs even at low temperature, which proves it isn't purely thermal — it's rooted in the software stack and gets worse the longer the generation runs.
The lesson: the ceiling on local AI is reliability, not just speed. A phone doesn't merely run models slowly; past a certain sustained load it falls over, unpredictably. Any system that treats the phone as a dependable workhorse is building on sand.
The fix: a router that reads the phone's physical state
The proposed system, HybridInfer, is a router — a decision-maker that inspects each incoming query and picks where to run it, across three tiers of rising power:
- On-device: a 3-billion-parameter model on the phone.
- Edge: an 8-billion model on a nearby server, able to look things up (retrieval).
- Cloud: GPT-4o, the heavyweight.
What makes it new is what it looks at before deciding. Prior routers are thermal-blind — they ignore the phone's physical condition, and most were tested only in simulation or on non-phone hardware. HybridInfer reads two live signals: how much thermal headroom the phone has left, and how hard the query looks. Then it picks a tier.
The routing policy is learned by reinforcement learning — a method where a system tries actions, receives a numerical reward for good outcomes, and gradually settles on the policy that maximises total reward. The reward here trades off four quantities: answer quality (reward), latency and cost (penalty), and a thermal penalty for pushing the phone toward the failure state above.
The trap in the reward — the part worth keeping forever
Here is the result to carry out of this paper. With only those four factors in the reward, the trained policy did something stupid: it sent every query to the cloud.
The logic was airtight from the optimiser's point of view. The cloud never crashes your phone, never overheats, and gives good answers. The phone's real advantages — privacy, offline, free — appear nowhere in a reward that counts only quality, speed, and heat. So the arithmetic says: never use the phone. The learner didn't decide privacy was worthless. It never saw privacy on the scoreboard, so it drove it to zero.
The fix was a locality bonus — an explicit reward just for running on-device. Only with that term does thermal-aware routing behave sensibly; without it, the "smart" policy discards the entire reason for having a local model.
State the principle plainly: an optimiser annihilates any good you leave out of its scoring function — not from malice, but because the good is invisible to it. This is the same failure as a company measured only on quarterly revenue quietly gutting its research lab, or a school ranked only on test scores dropping everything untested. The values you don't price, you lose. If you want an optimiser — an algorithm, an employee, a market — to protect something, you must pay it for that thing explicitly and by name.
What the results actually mean
On a real Android benchmark of 210 prompts, the learned router beat two hand-tuned rule-based routers on quality, at the lowest cost of any adaptive setup. But the honest finding is about what kind of win it is.
Running everything on-device produced the same per-query quality — on the queries the phone could survive. It was just three to six times slower, and it failed outright on long queries. So the router does not win by giving smarter answers. It wins on latency, reliability, and coverage: faster, doesn't crash, and handles the hard queries the phone alone cannot.
That reframing is the payoff. The right question is not "which model is smartest?" It is "given a device that will physically fail under sustained load, how do you route work so the whole system stays fast and alive?"
Two things you now know that the headlines don't
- Local AI has a reliability ceiling, not just a speed ceiling. Phones run models by crashing, not merely crawling, because mobile GPU drivers weren't built for long autoregressive loops. Any "runs entirely on your device" claim deserves one question: for how many queries in a row?
- An optimiser destroys anything you forget to reward. If a value matters — privacy, safety, long-term research, staff morale — and it isn't written into the scoring function, the smartest possible system will optimise it out of existence and call the result efficiency.
Distilled from arXiv Machine Learning
Liked this one?
The week's best pieces, one email, every Sunday. Nothing else.