How it works

Running a language model on an Android phone

Storage for the download, memory to hold it, a processor to run it and a runtime to drive it, with Brello 1.0’s three models as a worked example.

Brello Research12 min readVersion 1.0

Summary

Running a language model on an Android phone takes four things: storage for the model file, enough RAM for its weights and a key–value cache that grows with the conversation, a processor to run it, usually the GPU with the CPU as fallback, and a runtime such as Google’s LiteRT-LM or llama.cpp to load the file and generate text. Quantization, storing weights in 4 or 8 bits instead of 16, is what lets models with billions of parameters fit. Brello 1.0 checks storage and memory before it downloads, then runs one of three models, with downloads from 977 MB to 3.66 GB, entirely on the phone.

  • A local model needs storage for its file, RAM for its weights and key–value cache, a processor to run it and a runtime to load it; Brello 1.0’s model files range from 977 MB to 3.66 GB.
  • Weight memory is roughly parameters × bits per weight ÷ 8: a 1.7-billion-parameter model needs about 3.4 GB at 16 bits and about 0.85 GB at 4 bits.
  • At 16-bit precision, a 4,096-token key–value cache for a model shaped like Qwen3 1.7B takes about 470 MB, and its full 32,768-token context would take about 3.76 GB.
  • Brello 1.0 checks for room for the model, its speed cache and 300 MB of headroom before downloading, and treats a model as a fit if the phone has at least 90% of its recommended RAM.
  • Brello 1.0 runs its models with LiteRT-LM on the GPU through OpenCL, falls back to the CPU, and does not use the NPU.
Contents10 sections

01

What does it take to run a language model on an Android phone?

Running a language model on an Android phone takes four things: storage for the model file, enough memory (RAM) to hold its weights while it works, a processor to do the arithmetic, usually the graphics processor (GPU), and a runtime that loads the file and generates text, one token at a time.

Each requirement maps to a number you can check before installing anything. Running locally means the model works offline once downloaded, and the prompt is processed on the phone instead of on a server, the idea at the heart of on-device AI. The cost is size: a model that fits on a phone is far smaller than the models behind cloud assistants, as our explainer on small language models sets out.

Figure 1 follows one model, Brello Core, through those steps on a phone with 4 GB of RAM. The sections after it take each requirement in turn, then work through Brello 1.0’s three models as an example.

From download to answer: running Brello Core on a phone with 4 GB of RAM Inside a dashed boundary labelled “On this phone” are three panels: storage, memory and processor. First, Brello checks that the phone has about 2.25 GB of free space, for the 977 MB model file, a speed cache of about 974 MB and 300 MB of headroom, and that the phone, which reports about 3.7 GB of RAM, has at least 3.6 GB, 0.9 times the recommended 4 GB. Second, the 977 MB file downloads from Hugging Face. Third, the first launch writes the speed cache, bringing storage in use to about 1.95 GB. Fourth, the weights load into memory and run on the GPU through OpenCL, with the CPU as fallback and the NPU unused. Fifth, the model answers on the phone, with a 4,096-token context and up to 1,200 tokens per answer. Hugging Face Model host litert-community On this phone Brello Core · 4 GB of RAM Storage Free space check Model file977 MB Speed cache≈974 MB Headroom300 MB Needs ≈2.25 GB free 977 MB in use ≈1.95 GB in use Memory RAM check Reports ≈3.7 GB Needs ≥3.6 GB 0.9 × 4 GB Fits Weights loaded Processor Backend GPU OpenCL CPU Fallback NPU Not used Answer Convert 350 °F to Celsius. 350 °F is about 177 °C. 4,096-token context · up to 1,200 tokens per answer
  1. Brello checks the phone before it downloads. It requires free space for the model, its speed cache and 300 MB of headroom, and warns if the phone reports less than 90% of the 4 GB of RAM recommended for Brello Core.
  2. The model file downloads once. Brello Core’s 977 MB file comes straight from Hugging Face, as a foreground download that keeps going if you leave the app.
  3. The first launch writes a speed cache. The runtime tunes the weights to this phone’s chip and saves them, about 974 MB more, so later launches are faster.
  4. The weights load into memory and run on the GPU. LiteRT-LM tries the GPU through OpenCL first and falls back to the CPU. Brello 1.0 does not use the NPU.
  5. The model writes the answer on the phone. Each reply sees up to 4,096 tokens of context and can run to 1,200 tokens, or 2,048 with Think harder.
Figure 1How a local model gets from a download to an answer, shown for Brello Core on a phone with 4 GB of RAM. The sizes, checks and limits are the ones Brello 1.0 uses; the question and answer are illustrative.

02

How much storage does a local model need?

A local model needs room for its file, for any cache the runtime writes on first launch, and for some free space to spare. For the three models in Brello 1.0, the files range from 977 MB to 3.66 GB, and the total including the cache from about 1.95 GB to 6.26 GB.

The file holds the model’s weights, the numbers it learned in training, so its size follows from the number of parameters and the bits used to store each one (section 06). It is usually downloaded once, ideally over Wi-Fi, from a model host such as Hugging Face. Size labels need care: the Qwen3 1.7B file that Brello Core uses is listed as 932 MB on its Hugging Face page 1, which is the same 977,184,032 bytes expressed in binary units (1 MiB is 1,048,576 bytes). Brello reports sizes in decimal units, where 1 GB is 1,000,000,000 bytes.

A multi-gigabyte download also has to survive interruptions. Brello runs it as an Android foreground service, so it keeps going if you leave the app, and retries automatically up to 10 times. The runtime may then build a cache on first launch (section 07), which for Brello’s models adds about 974 MB to 2.60 GB. Brello checks before it starts: if free space is less than the model plus its speed cache plus 300 MB of headroom, it shows “Not enough space” with the exact numbers.

03

How much RAM does a local model need?

A local model needs enough memory for the weights it computes with, for a key–value cache that grows with the length of the conversation, and for the runtime’s own working space. Measurements on Android phones, published on the model pages for the files Brello 1.0 uses, range from about 0.7 GB to 3.3 GB of process memory, depending on the model and on whether it runs on the GPU or the CPU 1 2 3.

The weights come first. Their size is roughly the number of parameters times the bits stored per parameter, divided by eight: 1.7 billion parameters take about 3.4 GB at 16 bits and about 0.85 GB at 4 bits. Runtimes can avoid holding every weight in RAM. Google’s Gemma 4 E-series models keep a large share of their parameters in embedding tables, and LiteRT-LM memory-maps these, reading them from storage as needed; the model page for Gemma 4 E4B says this “enables significant working memory savings on some platforms”, and that the vision and audio parts are loaded only when needed 3.

The second consumer is the key–value (KV) cache. So that it does not recompute the whole conversation for every new token, a transformer keeps the keys and values it has already computed for earlier tokens, in every layer, and this cache grows with every token of context 4. Its size follows from the model’s shape. Qwen3 1.7B has 28 layers, each with 8 key–value heads of 128 dimensions 5, so at 16-bit precision each token adds 2 × 28 × 8 × 128 × 2 bytes, about 115 KB. A 4,096-token context therefore needs about 470 MB, and the model’s full 32,768-token context would need about 3.76 GB. The exact figure depends on the precision the runtime stores the cache in, but the scaling explains why phone apps run models with far shorter contexts than they support, and why a context window is a memory decision as much as a capability one.

Android and other apps share the same memory, and phones report slightly less than their advertised RAM. No single rule says how much RAM a model needs: it depends on the file, the runtime, the context length and the processor. Brello reads the phone’s total memory and treats a model as a fit if the phone has at least 90% of the model’s recommended RAM: 12 GB for Brello Pro, 6 GB for Brello Vision and 4 GB for Brello Core. On a smaller phone it warns, for example, that “Brello Pro may be too big… It may run slowly or fail to start.” You can still choose “Download anyway”.

04

Which chip runs the model: GPU, CPU or NPU?

Any of the three can. The GPU suits the parallel arithmetic of a language model and is the usual first choice, the CPU works on every phone and serves as the fallback, and the neural processing unit (NPU) in recent chips can be faster still but needs a model built for that chip.

Generating text has two phases with different demands. Prefill reads the whole prompt at once, which means many independent multiplications and suits a GPU’s parallelism. Decode then produces the answer one token at a time; each new token needs every weight read from memory again, so memory bandwidth limits it as much as arithmetic does.

Published measurements show how much the choice can matter, and that it varies by file. A measurement on the Hugging Face page for the 4-bit Qwen3 1.7B file that Brello Core uses, taken on a Samsung Galaxy S26, recorded about 41 tokens per second of decoding on the GPU through OpenCL and about 8 on the CPU, with peak memory of 1,036 MB on the GPU against 2,523 MB on the CPU; the page calls the file “a GPU specialist” 1. For Gemma 4 E2B on a Galaxy S26 Ultra, the GPU decoded 52.1 tokens per second against 46.9 on the CPU, but produced its first token in 0.3 seconds instead of 1.8 2.

LiteRT, the layer beneath LiteRT-LM, accelerates the CPU with XNNPack and the GPU with its ML Drift engine 2. Brello 1.0 loads each model on the GPU through OpenCL and silently falls back to the CPU if that fails; its Settings row shows “Running on GPU” or “Running on CPU”. Running on an NPU requires models built for a specific chip, and Brello 1.0 doesn’t ship any, so it leaves the NPU libraries out of the app, which saves about 58 MB.

05

What do runtimes and model files do?

A runtime is the software that loads a model file into memory, runs it on the phone’s processors and turns its output into text. Each runtime reads its own file format: Google’s LiteRT-LM reads .litertlm files, and llama.cpp reads GGUF files.

LiteRT-LM is open source, and its repository describes it as “Google’s production-ready orchestration layer to run LLMs with LiteRT”, with support for Android, iOS, the web, desktop and IoT devices 6. LiteRT supplies the hardware acceleration, and LiteRT-LM adds what language models need on top, such as KV-cache management, prompt templating and function calling 2. Converted models are published by the litert-community organisation on Hugging Face, often in more than one variant: Qwen3 1.7B is offered both as an 8-bit file of 2.1 GB and as the 4-bit, 977 MB file Brello Core uses, which is configured for a 4,096-token context 1.

llama.cpp is an independent open-source project whose stated goal is “to enable LLM (and VLM) inference with minimal setup” on a wide range of hardware, locally and in the cloud 7. It supports integer quantization from 1.5 to 8 bits and reads models in GGUF, which its specification describes as “a binary format that is designed for fast loading and saving of models”, built to contain “all the information needed to load a model” 8.

Android also offers a model as part of the system. Gemini Nano runs in AICore, a system service that apps call through APIs instead of shipping a model themselves. Google states that AICore “does not have direct internet access”, routes requests such as model downloads through its Private Compute Services companion app, and “doesn’t store any record of the input data or the resulting outputs after processing them” 9 (checked 5 October 2026). A system model saves each app a download; an app that brings its own model chooses the model, its version and the settings it runs with.

06

Why are model files smaller than you’d expect?

Model files are small because their weights are quantized: stored as 8-, 4- or even 2-bit integers instead of the 16- or 32-bit floating-point numbers used in training. Going from 16 to 4 bits cuts the size of the weights by a factor of four, at some cost in accuracy.

The arithmetic is direct. Qwen3 1.7B is published with 16-bit weights 5, about 3.4 GB for 1.7 billion parameters. The file Brello Core runs is “a dynamic INT4 variant (block-32 weights, FP32 activations)” 1: weights are stored as 4-bit integers in blocks of 32, and the computation runs in 32-bit floating point. At 977 MB, the file averages about 4.6 bits per parameter, a little above 4 because block-wise schemes store a scale factor for every block of weights.

Rounding weights to fewer bits introduces error, and quantization research is about keeping it small. LLM.int8() halved the memory needed for inference with 8-bit matrix multiplication, running a 175-billion-parameter model “without performance degradation” 10. GPTQ reduced models of that size to 3 or 4 bits per weight “with negligible accuracy degradation” 11, and AWQ found that protecting about 1% of salient weights “can greatly reduce quantization error” 12. Production schemes often mix precisions: the Gemma 4 E2B file for LiteRT-LM uses “a mixture of 2bit, 4bit and 8 bit weights” 2.

Fewer bits also speed up decoding, since each new token means reading the weights from memory again. The accuracy cost varies by model, method and task, so judge a quantized model on its own results, not its parent’s; our explainer on quantization covers the methods.

07

What happens the first time a local model starts?

The first time a model loads, the runtime prepares its weights for the phone’s processor and saves the result, so the first start is slower than later ones. In Brello 1.0 this step is labelled “Optimize for this phone” and takes up to a minute.

Brello’s runtime builds a weight cache tuned to the phone’s chip on first load, using XNNPack’s weight cache on the CPU path. Published benchmarks reflect the same effect: the LiteRT-LM model pages note that their figures were taken “with caches enabled and initialized”, and that “During the first run, the latency and memory usage may differ” 2. The Qwen3 1.7B page adds that starting the GPU engine took 4.4 to 8.3 seconds per process in its measurements 1.

The cache costs storage: about 974 MB for Brello Core, almost the size of its model file, 1.01 GB for Brello Vision and 2.60 GB for Brello Pro. Removing a model in Brello frees both the file and its cache. After the first launch, Brello warms the model up while the app draws its first frame, so it is usually ready by the time you type.

First launch is also when memory problems show. A model too large for the phone’s memory can crash the app while it loads. Brello remembers when that happens and, on the next launch, switches to another installed model with a notice such as “Brello Pro couldn’t start on this phone · Using Brello Vision”, instead of crashing again.

08

What should you check in a local AI app?

Check six things before installing a local AI app: which model it runs and under what licence, the download and cache sizes, the RAM it recommends, which processor it uses, what it sends over the network, and its context and answer limits.

  1. Model and licence. The app should name the base model, link to its owner’s model card and state the licence, for example Apache 2.0.
  2. Storage. Look for the download size and the cache written on first launch, in stated units, and for a free-space check before the download starts.
  3. Memory. Look for a recommended amount of RAM for each model, and for what the app does on a phone with less.
  4. Processor. Check whether it uses the GPU, falls back to the CPU, and tells you which one is running.
  5. Network. Check what leaves the phone after the download: web search, analytics, crash reports or nothing. “Works offline” should mean the model answers with no connection at all.
  6. Limits. Look for the context window and the maximum answer length in tokens. A short context means long chats lose their early detail.

For the network questions, our explainer on where AI assistants process and keep your questions has a fuller checklist.

09

How does Brello 1.0 fit three models to different phones?

Brello 1.0 offers three models of different sizes and recommends the most capable one that fits the phone’s memory: Brello Pro for phones with 12 GB of RAM, Brello Vision for 6 GB and Brello Core for 4 GB. Table 1 gives the storage and memory each needs.

The models page gives each model’s full specification, our research note on fitting a model to the phone explains how these limits were chosen, and the Brello 1.0 page describes the app as a whole. Terms used here are defined in the glossary.

References

Reviewed . Web pages were checked on that date.

  1. litert-community. “Qwen3-1.7B.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/Qwen3-1.7B
  2. litert-community. “gemma-4-E2B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/gemma-4-E2B-it-litert-lm
  3. litert-community. “gemma-4-E4B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/gemma-4-E4B-it-litert-lm
  4. Kwon, W. et al. (2023). “Efficient Memory Management for Large Language Model Serving with PagedAttention.” Proceedings of the 29th Symposium on Operating Systems Principles (SOSP 2023). arxiv.org/abs/2309.06180
  5. Qwen. “Qwen3-1.7B.” Model card and configuration, Hugging Face. Accessed 5 October 2026. huggingface.co/Qwen/Qwen3-1.7B
  6. Google AI Edge. “LiteRT-LM.” GitHub repository. Accessed 5 October 2026. github.com/google-ai-edge/LiteRT-LM
  7. ggml-org. “llama.cpp.” GitHub repository. Accessed 5 October 2026. github.com/ggml-org/llama.cpp
  8. ggml-org. “GGUF.” File format specification, ggml repository. Accessed 5 October 2026. github.com/ggml-org/ggml/blob/master/docs/gguf.md
  9. Google. “Gemini Nano.” Android Developers. Accessed 5 October 2026. developer.android.com/ai/gemini-nano
  10. Dettmers, T. et al. (2022). “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/abs/2208.07339
  11. Frantar, E. et al. (2023). “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.” International Conference on Learning Representations (ICLR 2023). arxiv.org/abs/2210.17323
  12. Lin, J. et al. (2024). “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.” Proceedings of Machine Learning and Systems (MLSys 2024). arxiv.org/abs/2306.00978

Version history

  1. 1.0First published.