Many teams are currently working on striking the right balance of fine tuning and model size. Most aren't considering phones yet, but PCs off network.
It is entirely possible to have an LLM run "closed loop", but obviously Google and Apple want in that loop