Brello Research · Explainers
Learn how private, on-device AI works
Learn is Brello Research’s library of explainers: plain, sourced accounts of how private and on-device AI works, from small language models and quantization to retrieval, BM25 and prompt injection.
How each explainer is written. It answers its question in the first paragraph, then defines the concept, explains the mechanism and its limits, and shows how Brello 1.0 uses it, with the app’s real numbers. Each one lists its sources and the date it was last reviewed.
Foundations: models and their limits
What on-device and small language models are, and the limits they work within, from context windows to reasoning modes.
- On-device AI, explained
- Small language models, explained
- Context windows, explained
- Reasoning modes, explained
- Model cards and system cards, explained
Privacy: where your questions go
Where an assistant processes and keeps what you ask it, and how confidential computing protects information while a server works on it.
Show explainers on privacyHow it works: running and grounding a model
How a phone runs a language model, and how an assistant answers from documents it fetches at question time.
- Running a language model on an Android phone
- Quantization, explained
- Retrieval-augmented generation, explained
- BM25, explained
Safety: failure modes and testing
Why models state falsehoods fluently, how hidden instructions can redirect them, and how AI systems are tested before release.
Show explainers on safetyAll explainers

On-device AI, explained
What is on-device AI? How a private AI assistant runs a local LLM on an Android phone, why it works without internet, and where it still falls short.
Read
Where AI assistants process and keep your questions
Where your prompts go when you use an AI assistant: processing, storage, model training, human review and accounts, plus a checklist for judging privacy claims.
Read
Running a language model on an Android phone
What an Android phone needs to run a language model locally: storage for the download, RAM to load it, a GPU or CPU to run it, and a runtime like LiteRT-LM.
Read
Small language models, explained
What makes a language model small, what parameters and effective parameters mean, why small models fit on phones, and what they give up in capability.
Read
Quantization, explained
Quantization stores a model’s weights in fewer bits so it fits in a phone’s memory. How INT8 and INT4 work, what accuracy you trade and real on-device examples.
Read
Context windows, explained
A context window is how much text a language model can consider at once, counted in tokens. Why it limits memory in a chat, and how on-device apps budget it.
Read
Retrieval-augmented generation, explained
Retrieval-augmented generation lets a language model answer from documents fetched at question time and cite them. How RAG works, where it fails and why.
Read
BM25, explained
BM25 scores how well a passage matches a query from term frequency, term rarity and length. The formula, what k1 and b do, and how it ranks passages on a phone.
Read
Why AI makes things up
Language models predict likely text, not verified facts, so they can state falsehoods fluently. Why hallucinations and knowledge cutoffs happen, and what helps.
Read
Prompt injection, explained
Prompt injection hides instructions in text a language model reads, such as a web page. How direct and indirect attacks work, why they’re hard and the defences.
Read
AI safety evaluations, explained
How AI systems are tested before release: capability and safety evaluations, red-teaming, calibration and release gates, and what each can and can’t show.
Read
Confidential computing for AI, explained
Confidential computing protects data while it’s processed, using hardware isolation and remote attestation. How it applies to AI inference, and its limits.
Read
Model cards and system cards, explained
Model cards document what a model is for, how it was evaluated and its limits; system cards cover a deployed system. What each contains and how to read one.
Read
Reasoning modes, explained
What happens when an AI model ‘thinks’ before answering: reasoning tokens, chain of thought, when it helps, what it costs in time and length, and its limits.
ReadGlossary
The glossary of private and on-device AI terms defines the words these explainers use, from attestation and BM25 to context windows and quantization, each in one sentence, with how the term applies in Brello.
Research notes on how Brello 1.0 applies these ideas
The explainers cover general concepts. Brello Research’s publications show how Brello 1.0 applies them, with its exact parameters: ‘Answering from the open web, without a server’ on retrieval and BM25, and ‘Fitting a model to the phone in your pocket’ on memory and storage.