Glossary

Glossary of private and on-device AI terms

This glossary from Brello Research defines 40 terms used in private and on-device AI, from tokens and context windows to BM25, prompt injection and remote attestation, and gives the Brello 1.0 fact wherever a term applies.

Brello ResearchReviewed 5 October 2026

Contents6 sections

Each entry gives a definition in general terms. Where the term applies to Brello 1.0, our app for Android and iPhone, a line headed ‘In Brello 1.0’ gives the fact as built. Links then lead to a longer explainer, where there is one, and to related terms. In Brello’s research notes and explainers, the first mention of a term links to its entry here.

Brello 1.0 facts describe version 1.0.0 of the app, dated 4 October 2026. Definitions that rest on a standard, specification or paper name it and link to it, and all of them are listed under Sources. For longer explanations, start with the Learn library.

Brello Super Intelligence is in development. Where an entry mentions it, the text is marked ‘Brello SI: intended, not shipped’ and describes design intent, not results. The design is set out in ‘Private compute you can verify: the design space’.

Foundations

Apache License 2.0

A permissive open-source licence, published by the Apache Software Foundation in 2004, that allows use, modification and redistribution, including commercial use, provided the licence and notices are kept. It also grants a patent licence from contributors.

In Brello 1.0Gemma 4 and Qwen3, the open models that Brello’s three models are built on, are released under Apache 2.0. Credits are listed on the open-source licences page.

Cloud AI

Artificial intelligence that runs on a provider’s servers: the app sends each request over the internet and receives the answer. Servers can hold far larger models than a phone, and the provider’s policies decide what is kept.

In Brello 1.0Brello 1.0 uses no cloud AI and has no Brello server: questions, photos and answers are processed on the phone. With web search on, the search text goes directly from the phone to a search engine, which sees the request and the phone’s IP address, as it would for any web request.

Context window

The maximum number of tokens a model can consider at once, including its instructions, the conversation so far, any retrieved text and the answer it is writing.

In Brello 1.0Brello 1.0 runs its models with a 4,096-token window and carries only the last six messages into each answer (user messages clipped to 300 characters, Brello’s to 600), so long chats lose their early detail.

Effective parameters

For Google’s Gemma E2B and E4B models, a parameter count that leaves out large per-layer embedding tables, which are used only for quick lookups. The ‘E’ stands for effective; Google’s Gemma 4 model card lists E2B at 2.3 billion effective parameters (5.1 billion with embeddings) and E4B at 4.5 billion (8 billion).

In Brello 1.0Brello Vision is built on Gemma 4 E2B and Brello Pro on Gemma 4 E4B.

Large language model (LLM)

A neural network trained on very large amounts of text to predict the next token, which lets it read, summarise and write language. ‘Large’ refers to its number of parameters, often billions.

In Brello 1.0Brello 1.0 offers three language models, each running entirely on the phone: Brello Pro, built on Gemma 4 E4B; Brello Vision, built on Gemma 4 E2B; and Brello Core, built on Qwen3 1.7B.

On-device AI

Artificial intelligence whose model runs on the user’s own device, such as a phone or laptop, instead of on a provider’s servers. Questions are processed locally, so it can work offline once the model is installed.

In Brello 1.0Brello 1.0 runs its models on the phone and works offline after a one-time download of 977 MB to 3.66 GB, depending on the model. Web search is off by default; when it is used, the search text and page requests go directly from the phone, and the search engine and websites see an ordinary web request.

Open-weight model

A model whose trained weights are published for anyone to download and run, under a licence that sets the terms. Its training data and code may stay private; such releases are often called open models.

In Brello 1.0Brello Pro and Brello Vision are built on Google’s Gemma 4, and Brello Core on Alibaba’s Qwen3, all released under Apache 2.0. Brello is not affiliated with or endorsed by Google or Alibaba.

Parameter

A number inside a neural network whose value is set during training; together, these numbers, often called weights, encode what the model has learned. Model size is usually quoted as a count of them.

In Brello 1.0Brello Core is built on Qwen3 1.7B, where ‘1.7B’ means about 1.7 billion parameters.

Small language model (SLM)

A language model compact enough to run on constrained hardware such as a phone. There is no agreed cut-off; models of a few billion parameters or fewer are usually called small.

In Brello 1.0Brello Core, built on Qwen3 1.7B, is a 977 MB download recommended for phones with 4 GB of RAM. Like any on-device model, it is far smaller than frontier cloud models and can be wrong.

Token

The unit of text a language model reads and writes, often a word, part of a word or a punctuation mark, as defined by the model’s vocabulary. Context windows and answer lengths are counted in tokens.

In Brello 1.0Brello 1.0 caps each answer at 1,200 tokens, or 2,048 with Think harder, within a context window of 4,096 tokens.

Running models on a phone

GPU acceleration

Running a model’s arithmetic on the graphics processing unit (GPU), which performs many calculations in parallel, instead of on the central processing unit (CPU). It usually makes inference faster.

In Brello 1.0Brello 1.0 loads each model on the GPU first, through OpenCL, and falls back to the CPU if that fails. GPU acceleration is on by default, and Settings shows ‘Running on GPU’ or ‘Running on CPU’.

Inference

Running a trained model to produce an output, such as the next token of an answer, as distinct from training it. It can happen on a provider’s servers or on the user’s own device.

In Brello 1.0Brello 1.0 runs inference on the phone with Google’s LiteRT-LM runtime, on the GPU where it can and on the CPU otherwise.

LiteRT-LM

Google’s open-source framework for running large language models on phones and other edge devices, built on LiteRT, the runtime formerly called TensorFlow Lite.

In Brello 1.0Brello 1.0 runs all three of its models with LiteRT-LM, using builds published in the litert-community repositories on Hugging Face and downloaded without an account.

Quantization

Storing a model’s weights in fewer bits, such as 4- or 8-bit integers instead of 16- or 32-bit floating point, so it needs less memory and runs faster. Activations can be stored the same way, and the saving usually costs some accuracy.

In Brello 1.0Brello Core’s model has about 1.7 billion parameters yet downloads as a 977 MB file. At 16 bits (2 bytes) each, 1.7 billion parameters would take about 3.4 GB.

Sideloading

Installing an Android app from an APK (Android package) file instead of through an app store such as Google Play. Android first asks the user to allow the app doing the installing, such as a browser, to install unknown apps.

In Brello 1.0Brello 1.0 is not sideloaded: it installs from Google Play on Android and the App Store on iPhone.

Weight cache

Files a runtime writes the first time it loads a model, holding the weights rearranged for a particular processor so that later loads start faster.

In Brello 1.0Brello 1.0 builds this cache in a step labelled ‘Optimize for this phone’, and its model pages list it as the speed cache. The step runs once per model, takes up to a minute and writes about 974 MB for Brello Core, 1.01 GB for Brello Vision and 2.60 GB for Brello Pro. Read ‘Fitting a model to the phone in your pocket’.

Retrieval and answers

BM25

A ranking function that scores how well a passage matches a query, using how often each query term appears, how rare that term is and how long the passage is. Its parameters k1 and b control term saturation and length normalisation; Robertson and Zaragoza (2009) give the standard account.

In Brello 1.0Brello 1.0 ranks web passages with BM25 using k1 = 1.2 and b = 0.75, plus a bonus for covering more of the question’s terms and a small one for lead paragraphs.

Citation

A reference that points to the source a statement relies on, so that a reader can check it. In AI answers, it is usually a numbered marker linked to a retrieved page.

In Brello 1.0Brello 1.0 numbers its web sources [1] to [6]. Markers such as [1] in the answer link to the matching page, and source cards above the answer show each site’s domain and page title.

Grounding

Tying a model’s answer to specific source material supplied when the answer is written, so that each claim can be traced and checked. Retrieval and citations are the usual means.

In Brello 1.0With web results, Brello 1.0 instructs the model: ‘Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2].’ An instruction is not a guarantee.

Hallucination

A fluent, confident statement from a language model that is false or not supported by its sources. It happens because models generate likely text rather than look up verified facts.

In Brello 1.0Brello 1.0’s system prompt tells the model: ‘If you are not sure about something, say so instead of guessing.’ That is no guarantee, and like any on-device model, Brello’s models can be wrong.

Knowledge cutoff

The point at which a model’s training data ends. The model knows nothing of later events unless they are supplied when the question is asked, for example from a web search.

In Brello 1.0Brello 1.0’s models have a knowledge cutoff. When web search is off and a question looks time-sensitive, Brello asks first, with a card titled ‘Search the web for this?’ and the buttons ‘Search the web’ and ‘Answer offline’. Searching sends the search text from the phone directly to a search engine, which sees an ordinary web request. Read ‘Asking before going online: consent for web search’.

Reasoning mode

A setting in which a model writes out its working step by step before giving a final answer, trading speed for more careful answers to multi-step questions.

In Brello 1.0Brello 1.0 calls it Think harder. The reasoning appears in a ‘Thought process’ panel above the answer, answers can run to 2,048 tokens instead of 1,200, and all three models support it.

Retrieval-augmented generation (RAG)

A method that retrieves passages relevant to a question, from documents or the web, and gives them to the model so it can answer from them and cite them. The name comes from Lewis et al. (2020).

In Brello 1.0With web search on, Brello 1.0 sends the search text from the phone to a search engine and fetches up to four result pages directly; the search engine and those websites see ordinary web requests. It splits the pages into passages of about 420–700 characters, ranks them with BM25 and gives the model at most two per source, within 3,400 characters for the Gemma 4 models or 2,600 for Qwen3. Read ‘Answering from the open web, without a server’.

System prompt

Instructions an application gives a model ahead of the user’s message, setting its role, tone and rules. Users usually don’t see it, but it shapes every answer.

In Brello 1.0Brello 1.0’s base system prompt is four sentences long, because small models tend to copy the shape of long instructions. It adds the current date only when a question is about time or uses web results. Read ‘Small models copy the shape of their instructions’.

Safety and security

AI safety evaluation

A structured test of how an AI system could fail or be misused, such as through harmful outputs or privacy leaks, used to decide whether and how to release it.

Calibration

How closely a model’s confidence matches how often it is right. A well-calibrated assistant expresses uncertainty when, and only when, it is likely to be wrong.

In Brello 1.0Brello 1.0’s system prompt asks the model to say when it is not sure instead of guessing. That is an instruction, not a measurement of calibration.

Indirect prompt injection

An attack that hides instructions in content a model reads on the user’s behalf, such as a web page, email or document, instead of typing them into the chat. It matters most for assistants that browse or take actions.

In Brello 1.0Text from web pages reaches Brello 1.0’s model only when web search is used, as a numbered block of passages within a budget of 3,400 characters for the Gemma 4 models or 2,600 for Qwen3. Questions with a photo never use the web.

Model card

A short document published with a model that sets out its intended uses, how it was evaluated and its known limitations. The format was proposed by Mitchell et al. (2019).

In Brello 1.0Brello’s model pages follow this format for its three models, giving each one’s base model, licence, download size, memory needs and limitations.

Prompt injection

An attack that places instructions in text a model processes, aiming to override the instructions its developer or user gave it. OWASP lists it first, as LLM01:2025, among risks to applications built on large language models.

Red-teaming

Deliberately trying to make an AI system fail or behave harmfully, whether by people or by other models, to find problems before release. The name comes from military and security exercises in which a ‘red team’ plays the adversary.

System card

Documentation for a deployed AI system as a whole, covering its models, safeguards, data flows, evaluations and limitations, rather than a single model.

In Brello 1.0The Brello 1.0 system card covers the three on-device models, their context and output limits, every data flow, the safeguards around the model and known limitations.

Privacy and verifiable compute

Confidential computing

Protecting data while it is being processed by running the computation in a hardware-based, attested trusted execution environment. The wording follows the Confidential Computing Consortium.

In Brello 1.0Brello 1.0 uses no cloud AI: its AI runs entirely on the phone.
Brello SI: intended, not shipped. Work that needs more than the phone is intended to go to sealed compute: hardware-isolated, stateless servers that the device verifies before sending anything. Read ‘Private compute you can verify: the design space’.

Data retention

How long a service keeps what it collects, such as conversations, before deleting it, including any copies it keeps after the user deletes them.

In Brello 1.0There is no Brello server, so Brello has no logs to keep. Chats stay in private storage on the phone, excluded from Android cloud backups and device-to-device transfer, so losing or resetting the phone loses them. Search results are held only in memory, for 15 minutes. The search engine and websites Brello reads receive ordinary web requests, and their own policies decide what they keep. Compare how AI assistants handle your questions.

Do Not Track (DNT)

An HTTP request header, DNT: 1, asking websites not to collect and share data about the user’s activity across different sites. It was never widely honoured, and the W3C specification ended as a Working Group Note in January 2019.

In Brello 1.0Brello 1.0 still sends it with its search and page requests, alongside Global Privacy Control.

Global Privacy Control (GPC)

A signal, sent as the Sec-GPC: 1 request header, telling a website that the person does not want their personal data sold or shared. California’s Attorney General lists it as a valid way to opt out of the sale or sharing of personal information under the California Consumer Privacy Act.

In Brello 1.0Brello 1.0 sends it with its search and page requests, alongside Do Not Track. It is a request: what a site does with it depends on the site and the laws that apply to it. Read what Brello 1.0 sends when it searches the web.

Oblivious HTTP

A protocol (RFC 9458) that sends encrypted requests through a relay: the relay sees who is asking but not what, and the server sees what but not who. This holds only if relay and server don’t collude.

In Brello 1.0Not used: Brello 1.0 sends search and page requests directly from the phone, with no proxy or relay, so search engines and websites see the phone’s IP address as they would for any web request.
Brello SI: intended, not shipped. Work for sealed compute is intended to travel through an independent relay that forwards it without the device’s IP address and cannot read it.

Remote attestation

A process in which a system produces signed evidence of the software it is running, so a remote party can check that software before trusting it with data. The IETF’s RATS architecture (RFC 9334) defines its core roles: attester, verifier and relying party.

In Brello 1.0Not used: Brello 1.0 has no Brello server, proxy or relay, so there is nothing to verify.
Brello SI: intended, not shipped. Before sending a sealed-compute server anything, the device is intended to send it a fresh random value and check the signed evidence it returns against a public transparency log.

Reproducible build

A build process in which the same source code, build environment and instructions produce bit-for-bit identical output, so anyone can confirm that a published binary matches its source. The definition follows the Reproducible Builds project.

Transparency log

An append-only public record, built on a Merkle tree, that can prove an entry is included and that earlier entries have not been changed. Certificate Transparency for web certificates (RFC 6962, revised as RFC 9162) is the best-known example.

In Brello 1.0Not used: Brello 1.0 has no server software to record.
Brello SI: intended, not shipped. Each release of the sealed-compute software is intended to be built reproducibly, with its measurement, a hash of the exact software, appended to a public transparency log. Devices are designed to refuse any release that is not in the log.

Trusted execution environment (TEE)

A hardware-isolated area of a processor whose memory the rest of the machine, including the operating system and hypervisor, cannot read or alter.

In Brello 1.0Brello 1.0 runs its models on the phone’s own GPU or CPU and has no server of its own.
Brello SI: intended, not shipped. Sealed compute for Brello SI is intended to run in hardware-isolated environments of this kind. No hardware has been chosen.

Sources

  1. Apache Software Foundation. “Apache License, Version 2.0.” January 2004. apache.org/licenses/LICENSE-2.0
  2. H. Birkholz, D. Thaler, M. Richardson, N. Smith and W. Pan. “Remote ATtestation procedureS (RATS) Architecture.” RFC 9334, IETF, January 2023. rfc-editor.org/rfc/rfc9334
  3. California Department of Justice, Office of the Attorney General. “California Consumer Privacy Act (CCPA).” oag.ca.gov/privacy/ccpa (accessed 5 October 2026).
  4. Confidential Computing Consortium. “A Technical Analysis of Confidential Computing.” Version 1.3. confidentialcomputing.io (accessed 5 October 2026).
  5. “Global Privacy Control (GPC).” Specification draft. w3c.github.io/gpc (accessed 5 October 2026).
  6. Google. “Gemma 4 model card.” Google AI for Developers. ai.google.dev/gemma/docs/core/model_card_4 (accessed 5 October 2026).
  7. Google AI Edge. “LiteRT-LM.” GitHub repository. github.com/google-ai-edge/LiteRT-LM (accessed 5 October 2026).
  8. B. Laurie, A. Langley and E. Kasper. “Certificate Transparency.” RFC 6962, IETF, June 2013. rfc-editor.org/rfc/rfc6962
  9. B. Laurie, E. Messeri and R. Stradling. “Certificate Transparency Version 2.0.” RFC 9162, IETF, December 2021. rfc-editor.org/rfc/rfc9162
  10. P. Lewis et al. “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” Advances in Neural Information Processing Systems 33 (NeurIPS 2020). arXiv:2005.11401
  11. M. Mitchell et al. “Model Cards for Model Reporting.” Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19), 2019. arXiv:1810.03993
  12. OWASP Gen AI Security Project. “LLM01:2025 Prompt Injection.” OWASP Top 10 for LLM Applications 2025. genai.owasp.org/llmrisk/llm01-prompt-injection (accessed 5 October 2026).
  13. Reproducible Builds. “Definitions.” reproducible-builds.org/docs/definition (accessed 5 October 2026).
  14. S. Robertson and H. Zaragoza. “The Probabilistic Relevance Framework: BM25 and Beyond.” Foundations and Trends in Information Retrieval 3(4), 333–389, 2009. doi:10.1561/1500000019
  15. M. Thomson and C. A. Wood. “Oblivious HTTP.” RFC 9458, IETF, January 2024. rfc-editor.org/rfc/rfc9458
  16. W3C. “Tracking Preference Expression (DNT).” W3C Working Group Note, 17 January 2019. w3.org/TR/tracking-dnt