Today, Large Language Models are almost always consumed via cloud services, with associated costs, latency, and privacy implications. But is it really necessary to always stay connected?
In this talk, we will see how to run LLMs directly on smartphones, completely offline, using Flutter and the Cactus framework. We will analyze the benefits and trade-offs of on-device inference and demonstrate a live Flutter app integrating a local language model.


