Brello Research · Explainers

Learn how private, on-device AI works

Learn is Brello Research’s library of explainers: plain, sourced accounts of how private and on-device AI works, from small language models and quantization to retrieval, BM25 and prompt injection.

How each explainer is written. It answers its question in the first paragraph, then defines the concept, explains the mechanism and its limits, and shows how Brello 1.0 uses it, with the app’s real numbers. Each one lists its sources and the date it was last reviewed.

Foundations: models and their limits

What on-device and small language models are, and the limits they work within, from context windows to reasoning modes.

Show explainers on foundations

Privacy: where your questions go

Where an assistant processes and keeps what you ask it, and how confidential computing protects information while a server works on it.

Show explainers on privacy

How it works: running and grounding a model

How a phone runs a language model, and how an assistant answers from documents it fetches at question time.

Show explainers on how it works

Safety: failure modes and testing

Why models state falsehoods fluently, how hidden instructions can redirect them, and how AI systems are tested before release.

Show explainers on safety

All explainers

Foundations12 min

On-device AI, explained

What is on-device AI? How a private AI assistant runs a local LLM on an Android phone, why it works without internet, and where it still falls short.

Read

Privacy12 min

Where AI assistants process and keep your questions

Where your prompts go when you use an AI assistant: processing, storage, model training, human review and accounts, plus a checklist for judging privacy claims.

Read

How it works12 min

Running a language model on an Android phone

What an Android phone needs to run a language model locally: storage for the download, RAM to load it, a GPU or CPU to run it, and a runtime like LiteRT-LM.

Read

Foundations11 min

Small language models, explained

What makes a language model small, what parameters and effective parameters mean, why small models fit on phones, and what they give up in capability.

Read

How it works11 min

Quantization, explained

Quantization stores a model’s weights in fewer bits so it fits in a phone’s memory. How INT8 and INT4 work, what accuracy you trade and real on-device examples.

Read

Foundations10 min

Context windows, explained

A context window is how much text a language model can consider at once, counted in tokens. Why it limits memory in a chat, and how on-device apps budget it.

Read

How it works12 min

Retrieval-augmented generation, explained

Retrieval-augmented generation lets a language model answer from documents fetched at question time and cite them. How RAG works, where it fails and why.

Read

How it works11 min

BM25, explained

BM25 scores how well a passage matches a query from term frequency, term rarity and length. The formula, what k1 and b do, and how it ranks passages on a phone.

Read

Safety11 min

Why AI makes things up

Language models predict likely text, not verified facts, so they can state falsehoods fluently. Why hallucinations and knowledge cutoffs happen, and what helps.

Read

Safety11 min

Prompt injection, explained

Prompt injection hides instructions in text a language model reads, such as a web page. How direct and indirect attacks work, why they’re hard and the defences.

Read

Safety11 min

AI safety evaluations, explained

How AI systems are tested before release: capability and safety evaluations, red-teaming, calibration and release gates, and what each can and can’t show.

Read

Privacy12 min

Confidential computing for AI, explained

Confidential computing protects data while it’s processed, using hardware isolation and remote attestation. How it applies to AI inference, and its limits.

Read

Foundations9 min

Model cards and system cards, explained

Model cards document what a model is for, how it was evaluated and its limits; system cards cover a deployed system. What each contains and how to read one.

Read

Foundations9 min

Reasoning modes, explained

What happens when an AI model ‘thinks’ before answering: reasoning tokens, chain of thought, when it helps, what it costs in time and length, and its limits.

Read

Glossary

The glossary of private and on-device AI terms defines the words these explainers use, from attestation and BM25 to context windows and quantization, each in one sentence, with how the term applies in Brello.

Research notes on how Brello 1.0 applies these ideas

The explainers cover general concepts. Brello Research’s publications show how Brello 1.0 applies them, with its exact parameters: ‘Answering from the open web, without a server’ on retrieval and BM25, and ‘Fitting a model to the phone in your pocket’ on memory and storage.