Three models at a glance
Brello 1.0 offers three models that trade capability against download size and the memory they need: Brello Pro is the most capable, Brello Vision is balanced and Brello Core is the fastest. Brello recommends one from the phone’s memory, and you can switch between installed models at any time in Settings.
| Property | Brello Pro | Brello Vision | Brello Core |
|---|---|---|---|
| Tier | Most capable | Balanced | Fastest |
| In the app | The sharpest answers and photo understanding. Made for high-end phones. | Understands photos and reasons well, with a lighter download. | Quick, sharp text answers and the smallest download. |
| Based on | Gemma 4 E4B by Google | Gemma 4 E2B by Google | Qwen3 1.7B by Alibaba |
| Input | Text and photos | Text and photos | Text |
| Download | 3.66 GB3,659,530,240 bytes | 2.59 GB2,588,147,712 bytes | 977 MB977,184,032 bytes |
| Speed cache | ≈2.60 GB | ≈1.01 GB | ≈974 MB |
| Total storage | ≈6.26 GB | ≈3.60 GB | ≈1.95 GB |
| Free space needed | ≈6.56 GB | ≈3.90 GB | ≈2.25 GB |
| Recommended RAM | 12 GB | 6 GB | 4 GB |
| Fits from | 10.8 GB | 5.4 GB | 3.6 GB |
| Think harder | Yes | Yes | Yes |
| Licence | Apache 2.0 | Apache 2.0 | Apache 2.0 |
Inside Brello 1.0 all three work within the same limits: a 4,096-token context window and answers of up to 1,200 tokens, or 2,048 with Think harder. Each model card gives the base model’s published specifications, the exact file Brello downloads, the settings Brello uses and the model’s limitations.
Which model fits your phone
Brello recommends the most capable model that fits the phone’s memory, and a model fits when the phone reports at least 90% of the model’s recommended RAM. Brello reads the phone’s total memory from the system (/proc/meminfo) each time it starts.
The 10% allowance exists because phones report slightly less memory than they are sold with. A 12 GB phone reports about 11.2 GB, an 8 GB phone about 7.4 GB, a 6 GB phone about 5.6 GB and a 4 GB phone about 3.7 GB, so the thresholds sit at 10.8 GB for Brello Pro, 5.4 GB for Brello Vision and 3.6 GB for Brello Core. Figure 1 applies the rule step by step, and to any reading you choose.
Try a phone
A 12 GB phone reads about 11.2 GB. Brello recommends Brello Pro. Brello Vision and Brello Core also fit.
- Each model has a recommended memory size. Brello Pro is recommended for 12 GB of RAM, Brello Vision for 6 GB and Brello Core for 4 GB.
- A model fits at 90% of its recommended memory. Phones report slightly less memory than they are sold with, so Brello allows 10%: the thresholds are 10.8, 5.4 and 3.6 GB.
- Brello recommends the most capable model that fits. An 8 GB phone reads about 7.4 GB. Brello Vision and Brello Core both fit, so Brello Vision gets the blue “Best” badge.
- If the memory can’t be read, Brello recommends Brello Vision. The recommendation is preselected during setup, and you can still choose another model.
- Every phone gets one recommendation. 4 GB phones get Brello Core, 6 and 8 GB phones Brello Vision, and 12 GB phones Brello Pro. Below 3.6 GB nothing fits, and Brello still recommends Brello Core.
Show data
| Model | Recommended RAM | Fits from | Recommended for |
|---|---|---|---|
| Brello Pro | 12 GB | 10.8 GB | 12 GB phones (≈11.2 GB) |
| Brello Vision | 6 GB | 5.4 GB | 8 GB phones (≈7.4 GB), 6 GB phones (≈5.6 GB) and unknown memory |
| Brello Core | 4 GB | 3.6 GB | 4 GB phones (≈3.7 GB) and phones below every threshold |
The recommended model is preselected during setup and carries a blue “Best” badge wherever models are listed. It is advice, not a limit. If you choose a model that needs more memory than the phone has, Brello warns that it “may run slowly or fail to start” and offers “Download anyway”. If a model then crashes the app while loading, Brello remembers, switches to another installed model at the next launch and says so, instead of crashing again. The reasoning behind the thresholds is in ‘Fitting a model to the phone in your pocket’.
Text and photos, or text only
Brello Pro and Brello Vision answer questions about photos as well as text; Brello Core reads text only.
You attach a photo from the camera or the gallery through the + menu, one photo per message, with or without a question. Brello scales it to at most 1280 × 1280 pixels, at quality 88, and keeps a copy in its private storage so the photo stays with the chat. A message with a photo never triggers a web search, so photo questions are answered on the phone.
With Brello Core active, tapping Camera or Photos opens a sheet titled “Photos need Brello Vision”, which offers to switch to an installed vision model or to download one; Brello Core keeps working while the download runs. Google’s model card lists audio as an input for Gemma 4 E2B and E4B,1 but Brello 1.0 has no voice input or output, so it uses text and photos only.
Storage: download, speed cache and headroom
Each model needs room for three things: the one-time download, a speed cache that the runtime writes the first time the model loads, and 300 MB of headroom. Brello checks for all three before it starts downloading.
The download comes straight from Hugging Face, with no account or token, and Wi-Fi is recommended. It runs as a foreground service, so it continues if you leave the app, and it retries automatically up to 10 times. The last setup step, “Optimize for this phone”, builds a weight cache tuned to the phone’s chip; it happens on first launch only and takes up to a minute. If free space is short, Brello shows “Not enough space” with the exact numbers instead of starting. Removing a model in Settings frees both the model and its cache.
Sizes on this site follow the app and use decimal units: 1 GB is 1,000,000,000 bytes and 1 MB is 1,000,000 bytes. Some tools count in binary units instead, which is why the 977,184,032-byte file behind Brello Core is listed as 932 MB on its Hugging Face page5 and as 977 MB here.
Where the models come from
The three models are built on open models from Google and Alibaba, released under the Apache 2.0 licence. Brello downloads ready-made LiteRT-LM builds of them from the litert-community organisation on Hugging Face.
| Brello model | Base model | Released by | Hugging Face repository and file |
|---|---|---|---|
| Brello Pro | Gemma 4 E4B | litert-community/gemma-4-E4B-it-litert-lmgemma-4-E4B-it.litertlm | |
| Brello Vision | Gemma 4 E2B | litert-community/gemma-4-E2B-it-litert-lmgemma-4-E2B-it.litertlm | |
| Brello Core | Qwen3 1.7B | Alibaba (Qwen team) | litert-community/Qwen3-1.7BQwen3-1.7B_dynamic_wi4b32_afp32.litertlm |
LiteRT-LM is Google’s open-source runtime for language models on devices, and it reads these .litertlm files directly. Brello 1.0 runs it through the flutter_gemma plugin, on the phone’s GPU through OpenCL where it can and on the CPU otherwise; the Settings row shows “Running on GPU” or “Running on CPU”. ‘Running a language model on an Android phone’ explains the general approach.
Brello isn’t affiliated with or endorsed by Google or Alibaba. The licence notices are on the licences page.
Context window and answer length
Brello 1.0 gives every model a 4,096-token context window and limits answers to 1,200 tokens, or 2,048 tokens with Think harder.
Each reply starts a fresh model session that holds a short system prompt, up to six earlier messages and, if you have chosen to search the web, the passages Brello selected from the pages it read. Earlier messages are clipped, yours to 300 characters and Brello’s to 600, so long chats lose their early detail. Web passages are capped at 3,400 characters for the Gemma 4 models and 2,600 for Qwen3. When Brello searches, the search text goes directly from your phone to a search engine, which sees the request and your IP address as it would for any web request.
The base models accept far longer inputs. Google’s model card gives Gemma 4 E2B and E4B a 128K-token context,1 and Qwen’s gives Qwen3 1.7B 32,768 tokens.2 The Gemma 4 LiteRT-LM builds support up to 32k tokens,3 and the INT4 build of Qwen3 1.7B that Brello uses is listed with a 4,096-token context.5 Brello 1.0 sets 4,096 tokens for all three. ‘Context windows, explained’ covers what the window holds.
Think harder works on all three
Think harder, Brello’s reasoning mode, works with all three models: the model reasons before it answers, and the reasoning appears in a “Thought process” panel above the answer.
You turn it on from the + menu. While the model reasons, the status line reads “Reasoning” and the panel shows the last few lines as they are written; when it finishes, the panel collapses to a row that expands to show the full reasoning. Answers can run to 2,048 tokens instead of 1,200, and sampling changes to temperature 0.6 and top-p 0.95. Replies are slower and more careful. ‘Reasoning modes, explained’ describes the technique in general.
What small on-device models can’t do
The models in Brello 1.0 are far smaller than frontier cloud models: they can be wrong, their knowledge stops at a training cutoff, and they are weaker at long or complex reasoning.
- They make mistakes. Brello’s system prompt tells the model, “If you are not sure about something, say so instead of guessing.” The instruction is no guarantee.
- Their knowledge has a cutoff. Google’s model card gives January 2025 as the cutoff of Gemma 4’s training data.1 Brello looks up newer facts on the web only if you allow it, and the search text then leaves the phone.
- Speed and quality depend on the phone. Brello Pro is recommended only for 12 GB phones, and the smaller models are faster but less capable. The first launch of each model takes up to a minute.
- Memory within a chat is short. Only the last six messages, clipped, are carried into each answer.
- Inputs are limited. One photo per message; no PDFs, documents or other files; no voice. The interface is in English only, and the models handle other languages to varying degrees.
We have not published quality benchmarks for these models. Benchmark scores in Google’s and Qwen’s model cards describe the base models under their own test conditions, not the builds Brello runs on a phone, and we don’t present them as Brello’s. The Brello 1.0 system card lists the known limitations of the whole app.
Model cards
Each model has a card with its base model’s published specifications, the exact file Brello downloads, the settings Brello 1.0 uses and the model’s limitations.
Brello Super Intelligence is in development and has no model card until its specifications are published. Read how Brello Super Intelligence is being designed.
Sources
Facts about the base models come from their publishers’ own pages, checked on 5 October 2026. Facts about Brello describe Brello 1.0, version 1.0.0 (build 1).
- Google. “Gemma 4 model card.” Google AI for Developers. ai.google.dev/
gemma/ . Checked 5 October 2026.docs/ core/ model_card_4 - Qwen. “Qwen3-1.7B.” Model card, Hugging Face. huggingface.co/
Qwen/ . Checked 5 October 2026.Qwen3-1.7B - litert-community. “gemma-4-E4B-it-litert-lm.” Model repository, Hugging Face. huggingface.co/
litert-community/ . Checked 5 October 2026.gemma-4-E4B-it-litert-lm - litert-community. “gemma-4-E2B-it-litert-lm.” Model repository, Hugging Face. huggingface.co/
litert-community/ . Checked 5 October 2026.gemma-4-E2B-it-litert-lm - litert-community. “Qwen3-1.7B.” Model repository, Hugging Face. huggingface.co/
litert-community/ . Checked 5 October 2026.Qwen3-1.7B