01
System overview
Brello 1.0 is a private AI assistant for Android and iPhone, made by Stuvio, that runs its language model entirely on the phone. This card describes the Android app, version 1.0.0 (build 1), labelled ‘Brello 1.0 · Prototype’ in the app, as built on 4 October 2026.
- Maker
- Stuvio (app ID
co.stuvio.brello) - Status
- On Google Play and the App Store
- Platform
- Android 8.0 (API 26) or newer, 64-bit ARM (arm64-v8a) only
- Models
- Brello Pro, Brello Vision and Brello Core, run on the phone with LiteRT-LM
- Network use
- A one-time model download; optional web search, off by default
- Servers
- None: no Brello backend, accounts, analytics, crash reporting or ads
- Evaluation
- No benchmark results are published (section 10)
Brello is not a cloud chatbot: there is no server-side model and no fallback that sends a question elsewhere. Nor is it a developer tool; users choose between three curated models and cannot load other model files or change sampling settings. This card follows the way AI developers document deployed systems in model and system cards, and says plainly where no evaluation exists. For a plain-language overview, see the Brello 1.0 product page.
02
Models and configuration
Brello 1.0 ships three models, each a LiteRT-LM build of an open model, downloaded from Hugging Face when the user chooses it and run on the phone’s GPU or CPU (Table 1). Sizes are decimal (1 GB = 1,000,000,000 bytes).
| Table 1 | Brello Pro | Brello Vision | Brello Core |
|---|---|---|---|
| In the app | Most capable | Balanced | Fastest |
| Base model | Gemma 4 E4B, Google | Gemma 4 E2B, Google | Qwen3 1.7B, Alibaba |
| Licence | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| Inputs | Text, photos | Text, photos | Text |
| Download | 3.66 GB 3,659,530,240 bytes | 2.59 GB 2,588,147,712 bytes | 977 MB 977,184,032 bytes |
| Speed cache | ≈2.60 GB | ≈1.01 GB | ≈974 MB |
| Total storage | ≈6.26 GB | ≈3.60 GB | ≈1.95 GB |
| Free space needed | ≈6.56 GB | ≈3.90 GB | ≈2.25 GB |
| Recommended RAM | 12 GB | 6 GB | 4 GB |
| Sampling, for reference | Temperature 1.0, top-k 64, top-p 0.95 | Temperature 1.0, top-k 64, top-p 0.95 | Temperature 0.7, top-k 20, top-p 0.8 |
All three share a 4,096-token context window and a maximum answer of 1,200 tokens, or 2,048 in Think harder, which samples at temperature 0.6 and top-p 0.95. Google’s LiteRT-LM runtime, called through flutter_gemma, uses the GPU through OpenCL and falls back to the CPU with an XNNPack weight cache. NPU libraries are left out of the build, saving about 58 MB. The files download from Hugging Face’s litert-community organisation without an account or token (Table 2).
| Table 2 | File | Hugging Face repository |
|---|---|---|
| Brello Pro | gemma-4-E4B-it. | litert-community/ |
| Brello Vision | gemma-4-E2B-it. | litert-community/ |
| Brello Core | Qwen3-1.7B_ | litert-community/ |
How Brello recommends a model
On launch, Brello reads the phone’s total memory from /proc/meminfo. Phones report a little less than their advertised RAM, so a model fits when the phone’s RAM is at least 0.9 times the model’s recommended RAM. Brello recommends the most capable model that fits and marks it ‘Best’ (Table 3). ‘Fitting a model to the phone in your pocket’ describes the recommendation, the storage check, GPU loading and crash-loop recovery in more detail.
| Table 3 · Phone | Reported memory | Recommended model |
|---|---|---|
| 12 GB phone | ≈11.2 GB | Brello Pro |
| 8 GB phone | ≈7.4 GB | Brello Vision |
| 6 GB phone | ≈5.6 GB | Brello Vision |
| 4 GB phone | ≈3.7 GB | Brello Core |
| Below every minimum | n/a | Brello Core |
| Memory unknown | n/a | Brello Vision |
Download and lifecycle
Before downloading, Brello checks for free space equal to the model, its speed cache and 300 MB of headroom, and compares the phone’s memory with the model’s recommendation. Downloads run as a foreground service, one at a time, with up to 10 automatic retries. On first load the runtime builds a weight cache tuned to the phone’s chip (‘Optimize for this phone’), which takes up to a minute. The engine then moves through the states none, installed, loading, ready or failed, trying the GPU first and falling back to the CPU without asking. Brello isn’t affiliated with or endorsed by Google or Alibaba.
03
Inputs and outputs
Brello 1.0 takes text and at most one photo per message, and returns a streamed Markdown answer, with numbered citations when web results were used (Table 4).
| Table 4 | Specification |
|---|---|
| Text | Typed questions, trimmed on send. No voice input or output. |
| Photos | One per message, from Camera or Photos, on Brello Pro and Brello Vision; downscaled to at most 1280 × 1280 at quality 88. Sent alone, a photo becomes ‘Describe this image in detail.’ |
| Not accepted | PDFs, documents, other files; more than one photo |
| Conversation context | Up to the 6 most recent earlier messages, clipped to 300 characters (user) and 600 (Brello) |
| Answer | Streamed Markdown, up to 1,200 tokens, or 2,048 with Think harder |
| Reasoning | With Think harder, a ‘Thought process’ panel, stored with the chat |
| Citations | Inline [1], [2] links to source cards showing a letter mark, domain, number and page title |
| Answer details | Time taken and route, such as ‘2.4s · On-device’, and the model that wrote it |
04
How a reply is generated, step by step
Every reply runs the same nine steps on the phone, from sending to saving. Only step 5 can use the network, and only when web search is allowed. Figure 1 traces one reply, a follow-up question about the weather with web search on.
- The message is sent. The text is trimmed, and an empty reply with the status ‘Thinking’ joins the chat. A photo would be copied into private storage.
- Six earlier messages are carried forward. Older ones are dropped, your turns are clipped to 300 characters and Brello’s to 600, citations are removed, and a photo becomes ‘[shared a photo]’.
- Brello decides whether to search. No photo, web search on, and not small talk, maths or a creative task, so the research pipeline runs while the reply reads ‘Searching the web’, then ‘Reading 4 sources’. It returns numbered web results: up to 3,400 characters for Brello Vision.
- The prompt is assembled. The status returns to ‘Thinking’, and the four-sentence system prompt gains the date and the citation instruction because web results are present. The identity line is added only when you ask who Brello is.
- A fresh session writes the answer. Every 48 characters, a check looks for a block repeated 3 or more times over at least 120 characters; a loop is cut after its first copy. Stray markup is removed or routed to the thought panel.
- The reply is saved on the phone. The time taken and the model that wrote it are recorded, and the chat is written to conversations.json 400 ms after the last change, atomically.
- Send. The text is trimmed and any photo is copied into private storage. A new chat is titled with its first 42 or so characters, cut at a word boundary, or ‘Photo’. The user message and an empty reply with the status ‘Thinking’ are added.
- History. Up to the 6 most recent earlier messages are included: user turns clipped to 300 characters, Brello’s to 600, citation markers removed, photos noted as ‘[shared a photo]’.
- Photo path. If there is a photo and the model can see, the image goes to the model, an empty question becomes ‘Describe this image in detail.’, and the web is skipped.
- Freshness check. With web search off, no photo and a question that looks time-sensitive, the reply pauses at ‘Needs the web’ and shows ‘Search the web for this?’.
- Search decision. If web search is allowed and the question is not small talk, maths or a creative task, the research pipeline in section 7 runs.
- System prompt. The prompt is short on purpose, because small models copy the shape of their instructions. Three lines are added only under the stated conditions:
System prompt, verbatim
You are Brello, a helpful assistant that runs privately on the user's phone. Reply to the user's message directly and naturally, like a knowledgeable friend. Keep it clear and to the point. If you are not sure about something, say so instead of guessing.
Only when the user asks about Brello’s identity: You are {Model}, based on {base model}.
Only when the question is about time or web results are present: Current date: {weekday, Month day, year}.
Only with web results: Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2]. If the results do not answer the question, say so briefly and answer from general knowledge.
- Generation. Each reply opens a fresh model session, so earlier web context never crowds the 4,096-token window. Reasoning streams to the ‘Thought process’ panel.
- Quality guards. The repetition stopper, markup clean-up, thought fallback and empty-reply message run on the output (section 8).
- Finish. The time taken is recorded, a haptic tap plays, and the chat is saved with a 400 ms debounce and an atomic write. The reply records which model wrote it, so old chats keep the right name and orb.
05
Data flows: what stays on the phone and what leaves
Questions, photos, answers, reasoning, chats and settings stay on the phone (Table 5). Three kinds of request leave it, each started by the user: a model download, a web search and the page requests that follow the search (Table 6).
| Table 5 · Data | Where it lives | Notes |
|---|---|---|
| Questions to the AI | Processed by the local model | The prompt goes only to the on-device runtime |
| Brello’s answers | Generated locally | |
| Photos | Brello’s private app storage, images/ | Never uploaded; photo questions never trigger web search |
| Chat history | One JSON file, conversations.json, in private app storage | Written atomically: a temporary file, then a rename |
| Reasoning | Stored with the chat | The ‘Thought process’ text |
| Settings | Android SharedPreferences, private | Theme, web search, Think harder, GPU, chosen model |
Backups are disabled: allowBackup="false", and Android’s data-extraction rules exclude every domain from cloud backup and device-to-device transfer, so chats are not copied to Google Drive or to a new phone. The app has no analytics, crash-reporting, advertising or account SDKs, and source cards use letter marks drawn on the phone rather than downloaded site icons.
| Table 6 · When | What is sent | To whom | Controlled by |
|---|---|---|---|
| Downloading a model, once per model | A standard file download request | Hugging Face, litert-community repositories | The user starting a download |
| Web search, only if on or approved for that question | The search text: a cleaned-up version of the question, sometimes with up to 10 words of the previous question for context on follow-ups | Each search engine tried, in order, until one returns relevant results: DuckDuckGo, Bing, Brave Search, Google News (RSS) or Wikipedia | The web search setting, or the ‘Search the web for this?’ card |
| Reading results, at the same moment | Ordinary page requests for up to 4 result pages | The websites in the results | As above |
| Tapping a source card or citation | Opens the page | The user’s browser | The user |
Requests go directly from the phone, with no Brello server, proxy or relay, so Brello has no logs to keep. Search and page requests send DNT: 1 and Sec-GPC: 1. The search browser is torn down after 3 idle minutes and its cookies, cache and storage are wiped; results stay in memory for 15 minutes. Search engines and websites see what any web request shows, including the phone’s IP address. The privacy page explains these flows for a general reader.
06
Android permissions and why
Brello 1.0 asks for network, download-service and notification permissions; it reaches the camera and photos only through the system picker, when the user chooses them (Table 7).
| Table 7 · Permission | Why |
|---|---|
| Internet, Network state | Model download and optional web search |
| Foreground service (data sync), Post notifications, Wake lock | Lets the one-time model download continue, with a notification, when the app is in the background |
| Camera, photos | Requested by the system picker only when the user chooses Camera or Photos |
It requests no contacts, location, microphone, SMS, calendar or storage-wide permissions.
07
Web retrieval: providers, reading and ranking
When web search runs, the phone tries providers in a fixed order until one returns relevant results, reads up to four pages, ranks passages with BM25 and hands the model a numbered context block (Table 8). No step runs on a server.
| Table 8 · Order | Provider | Method | Timeout |
|---|---|---|---|
| 1 | DuckDuckGo, full results page | Invisible on-device browser (system WebView) | 14 s |
| 2 | Bing | Invisible on-device browser | 14 s |
| 3 | DuckDuckGo Lite | Direct HTTP | 7 s |
| 4 | DuckDuckGo HTML | Direct HTTP | 7 s |
| 5 | Google News RSS, last 7 days | Direct HTTP, only for time-sensitive questions | 7 s |
| 6 | Brave Search | Direct HTTP | 7 s |
| 7 | Bing | Direct HTTP | 7 s |
| 8 | Wikipedia search API | Direct HTTP | 7 s |
A provider that fails or is blocked is rested for 2 minutes, except Wikipedia. Results are de-duplicated and must pass a relevance guard: at least a third of the top hits must mention the question’s key terms. Results are cached in memory for 15 minutes per query.
Reading and ranking
- The top 6 results become sources [1] to [6].
- Up to 4 pages are fetched in parallel, each with a 7-second timeout, HTML or plain text only, at most 1.5 MB. Google News results use the headline and date.
- Reader-mode extraction keeps up to about 24,000 characters of article text per page.
- Pages are split into passages of about 420 to 700 characters.
- Passages are ranked with BM25 (k1 = 1.2, b = 0.75), with bonuses for covering more of the question’s terms and for lead paragraphs.
- At most 2 passages per source are kept, within 3,400 characters for the Gemma 4 models and 2,600 for Qwen3; sources without a good passage fall back to their snippets, 900 characters in total.
- The block handed to the model is labelled ‘Web results (retrieved {date})’.
Query handling and the invisible browser
Search is skipped for messages under 4 characters, arithmetic, small talk and creative or writing tasks, unless they mention something fresh. With web search off, words such as ‘latest’, ‘today’, ‘price’, ‘weather’ or ‘score’, or the current year, bring up the card that asks first. Short or referential follow-ups get the first 10 words of the previous question prepended. Queries are refined: answer-style instructions are removed, contractions expanded, ‘What’s new in X’ becomes ‘latest X news’, and time-relative questions get the current month and year, except for live data such as weather, scores and prices.
Brello searches in a single hidden WebView tab, one search at a time, which warms up on a blank page, with no network request, only when web search is on. If a provider returns a challenge page, Brello moves on to the next. After 3 idle minutes the browser is destroyed and its storage wiped. Parsing and ranking run on background isolates. The method is described in ‘Answering from the open web, without a server’.
08
Safeguards around the model
Brello 1.0’s safeguards are instructions, checks and recovery paths around the model. None changes the model’s weights, and none guarantees a correct answer.
- Uncertainty instruction. The system prompt ends: “If you are not sure about something, say so instead of guessing.”
- Short prompt, fresh session. The date and identity lines are added only when relevant, and each reply starts a new session, so earlier web context cannot pile up in the window.
- Asking first, photos offline. With web search off, a time-sensitive question waits for ‘Search the web’ or ‘Answer offline’. A message with a photo never triggers a search.
- Relevance guard and citations. Results must mention the question’s key terms, and the model is told to cite sources as [1], [2] and to say when they do not answer the question.
- Repetition stopper. Every 48 characters, Brello checks for a block repeated 3 or more times over at least 120 characters. If it finds one, it cuts the reply after the first copy and stops the model.
- Markup clean-up. Stray
<think>,<|im_end|>,<end_of_turn>,<eos>,/no_thinkand Gemma channel markers are removed or routed to the thought panel. - Thought fallback and empty replies. If the model reasoned but never answered while Think harder was off, the reasoning becomes the answer. An empty reply shows ‘No response was generated. Try rephrasing.’
- Pre-download checks. Too little free space shows ‘Not enough space’; too little memory shows ‘{Model} may be too big’, though the user can still choose ‘Download anyway’.
- Crash-loop recovery. If a model crashes the app while loading, the next launch switches to another installed model and shows ‘{Model} couldn’t start on this phone · Using {Other}’.
- GPU fallback. GPU acceleration is on by default, with automatic fallback to the CPU; Settings shows ‘Running on GPU’ or ‘Running on CPU’ and ‘Turn off if answers fail.’
The safety page sets these guards in Brello’s wider approach.
09
Known limitations
Brello 1.0 has the limits of a prototype and of models small enough to run on a phone.
- Platform. The Android app needs a 64-bit ARM phone with Android 8.0 or newer.
- Size. A 977 MB to 3.66 GB download plus a speed cache of similar size; some phones won’t have room for Brello Pro.
- Hardware. Speed and quality depend on the phone. Brello Pro wants 12 GB of RAM, and a model’s first launch takes up to a minute.
- Accuracy. The models are far smaller than frontier cloud models: they can be wrong, have a knowledge cutoff and are weaker at long or complex reasoning.
- Memory and inputs. Only the last 6 messages, clipped, reach each answer. One photo per message, no files, no voice.
- Web search. It depends on third-party engines, can take up to about 14 seconds per provider attempt or return nothing, and Brello then answers from its own knowledge.
- Language. The interface is English only; other languages work to varying degrees and are not offered as a feature.
- No sync, backup or export. If the phone is lost or reset, the chats are gone, by design.
10
Testing and evaluation status
No benchmark or evaluation results are published for Brello 1.0, and this card reports none. Nor does it present the published scores of Gemma 4 or Qwen3 as Brello’s: those describe the base models, not the models as Brello runs them, with its runtime, prompt, context limits and guards.
The app’s automated tests, written with flutter_test, cover model recommendation, the repetition guard, thought-channel splitting, search heuristics, query refinement and the relevance guard, plus a live test of the whole web research pipeline. They exercise the mechanisms described here; they are not benchmarks, and no scores from them are published. If evaluation results are published, this card will add them with their method, sample size, date and which direction is better, and its version will change. ‘AI safety evaluations, explained’ describes what such work involves.
11
Dependencies and licences
Brello 1.0 is a Flutter app that runs its models through flutter_gemma on Google’s LiteRT-LM engine; its models and bundled assets are openly licensed (Tables 9 and 10).
| Table 9 · Layer | Technology |
|---|---|
| App framework | Flutter (Dart SDK ^3.12.2; Flutter 3.44+), Material 3 with Cupertino dialogs and switches |
| On-device inference | flutter_gemma ^1.11.3 and flutter_gemma_litertlm ^1.8.5, running .litertlm files on Google’s LiteRT-LM engine |
| Acceleration | GPU through OpenCL (Android 12+ native library hooks); CPU fallback with an XNNPack weight cache |
| Web search | webview_flutter (headless system WebView), http and html; Dart isolates for parsing and ranking |
| Persistence | shared_preferences for settings; a JSON file through path_provider for chats; a private images/ folder |
| Interface | flutter_markdown_plus for answers, url_launcher for sources, image_picker for camera and gallery |
| Native (Kotlin) | A MethodChannel, brello/device, exposing free storage through StatFs |
| Table 10 · Component | Source | Licence |
|---|---|---|
| Gemma 4 E4B and E2B (Brello Pro, Brello Vision) | Google, through Hugging Face litert-community | Apache 2.0 |
| Qwen3 1.7B (Brello Core) | Alibaba (Qwen team), through litert-community | Apache 2.0 |
| LiteRT-LM runtime | Open source | |
| 3D renders | 3dicons by Vijay Verma | CC0 |
| 3D emoji | Microsoft Fluent Emoji | MIT |
| Icons | Phosphor Icons | MIT |
| Typeface | Inter and Inter Display by Rasmus Andersson | SIL Open Font License |
Brello isn’t affiliated with or endorsed by Google or Alibaba. Full notices are on the licences page.
12
Version history
Brello 1.0 was built in two commits on 4 October 2026, and this card was first published on 5 October 2026. The card is updated whenever the facts it describes change, and each change is recorded here (Table 11).
| Table 11 · Date | Version | Change |
|---|---|---|
| 4 October 2026 | Commit 82cc662 | Initial build: on-device inference through LiteRT-LM with GPU and CPU fallback; Brello Vision, Brello Core and a smaller third model; photo understanding; Think harder; opt-in web search from an on-device WebView with fallbacks, BM25 ranking, a relevance guard and citations. |
| 4 October 2026 | Commit 18dc2f5 | Brello Pro added and the smaller third model retired; models retiered. RAM-based recommendation, oversize warnings, a free-storage check and crash-loop recovery. Licence labels corrected to Apache 2.0. A shorter system prompt, the repetition stopper, leaked-thought routing and per-reply model attribution. |
| 5 October 2026 | Card 1.0 | First published. Describes Brello 1.0, version 1.0.0 (build 1). |
Cite this card. Brello Research. “Brello 1.0 system card.” Stuvio, 5 October 2026. https://brello.ai/brello/system-card/