Safety · Brello 1.0 and Brello SI

Safety at Brello: what Brello 1.0 does and what we intend for Brello SI

Brello 1.0, our on-device AI assistant for Android and iPhone, guards each reply with checks in code and two instructions to its model. Brello Super Intelligence is in development; this page sets out how we intend to evaluate it before release, and what we will publish.

Updated 5 October 2026Brello 1.0, version 1.0.0Brello SI in development

Contents10 sections

Our approach: evaluate first, then ship

Before Brello can do something new, we will test how it could fail or be misused, hold the release until it meets criteria written before testing began, and publish what we found. This is commitment 04 of the Brello Charter, version 1.0, and it covers Brello Super Intelligence and every Brello release after 1.0.0. Brello 1.0, version 1.0.0, predates the charter and has no published evaluation. It has the safeguards below, each with its real parameters, and none is a guarantee.

In Brello 1.0
Shipped in version 1.0.0, 4 October 2026.
Planned
Design intent for Brello SI. No results exist yet.

Safeguards in Brello 1.0 today

In Brello 1.0

Brello 1.0 runs eight safeguards. Six are checks written in code; two are instructions to the model, which it may not follow.

Table 1The safeguards in Brello 1.0, version 1.0.0, with the parameters the app uses.
SafeguardWhat it doesKind
Uncertainty instructionThe system prompt is short on purpose, because small models copy the shape of long instructions. It ends: “If you are not sure about something, say so instead of guessing.”Instruction
Inline citationsWith web results, the model is told to “cite the sources you rely on inline using their numbers, like [1] or [2]”. Each citation in the answer becomes a link to its source page.Instruction, with links in code
Asking before searchingWeb search is off by default. With it off, a question that looks time-sensitive pauses at “Needs the web” and shows the “Search the web for this?” card. A question with a photo never searches.Check in code
Repetition stopperEvery 48 characters, Brello checks whether the end of the reply is a block repeated 3 or more times over at least 120 characters. If it is, Brello cuts the reply after the first copy and stops the model.Check in code
Leaked-markup clean-upStray <think>, <|im_end|>, <end_of_turn>, <eos>, /no_think and Gemma “channel” markers are removed, or routed to the “Thought process” panel.Check in code
Empty repliesIf the model reasoned but never answered while Think harder was off, the reasoning becomes the answer. A reply with no text shows “No response was generated. Try rephrasing.”Check in code
Storage and memory checksBefore a download, Brello requires free space for the model, its speed cache and 300 MB of headroom, or shows “Not enough space” with exact numbers. On a phone with less memory than the model needs, it warns: “It may run slowly or fail to start.” You can still choose “Download anyway”.Check in code
Crash-loop recoveryIf a model crashes the app while loading, usually by running out of memory, Brello remembers. On the next launch it switches to another installed model and says why: “{Model} couldn't start on this phone · Using {Other}”.Check in code

Figure 1 follows one reply through the safeguards that act on its text: the two instructions, the repetition stopper and the markup clean-up.

One reply in Brello 1.0, checked as it streams At the top, the system prompt asks the model to say when it is not sure and, with web results, to cite sources inline. Below, the model’s raw output streams in rows of 48 characters, and a check runs at the end of each row. A leaked control token, think tags, is removed. The checks at 48, 96, 144 and 192 characters find no loop. At 240 characters the same 49-character sentence has appeared three times, 147 characters in all, so the reply is cut after the first copy and the model is stopped. On the right, the reply as the person sees it: the question, source cards for sources 1 and 2 above the answer, the answer with the first sentence and one copy of the second and citations 1 and 2 linked, and the meta line, 3.1 seconds, web plus on-device. Below a divider, what was removed: the think tags and 122 characters of repeats. System prompt Quoted from Brello 1.0 … Keep it clear and to the point. If you are not sure about something, say so instead of guessing. With web results: “Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2].” Model output, as it streams Check <think></think>Air molecules scatter sunlight in every direction [1]. Blue light is scattered mo re than red light [2]. Blue light is scattered m ore than red light [2]. Blue light is scattered more than red light [2]. Blue light is scattered 48 96 144 192 240 No loop No loop No loop No loop Loop Copy 1 · kept Copy 2 Copy 3 Copy 4, partial ChecksEvery 48 characters, the end of the reply is tested for a loop Markup<think></think> removed before display Loop49-character block × 3 = 147 characters (rule: 3+ copies, 120+) CutCut after copy 1; the model is stopped at character 240 What you see Why is the sky blue? W wikipedia.org 1 N noaa.gov 2 Air molecules scatter sunlight in every direction [1]. Blue light is scattered more than red light [2]. 3.1s · Web + on-device Removed from the reply Markup: <think></think> Repeats: 122 characters One reply in Brello 1.0, checked as it streams At the top, the system prompt asks the model to say when it is not sure and, with web results, to cite sources inline. Below, the model’s raw output streams in rows of 48 characters, each shown as two lines of 24, and a check runs at the end of each row. A leaked control token, think tags, is removed. The checks at 48, 96, 144 and 192 characters find no loop. At 240 characters the same 49-character sentence has appeared three times, so the reply is cut after the first copy and the model is stopped. Then the raw output gives way to the reply as the person sees it: the question, source cards for sources 1 and 2 above the answer, the answer with citations 1 and 2, and the meta line, 3.1 seconds, web plus on-device. Below a divider, what was removed: the think tags and 122 characters of repeats. System prompt Quoted “…If you are not sure about something, say so instead of guessing.” With web results: “…cite the sources you rely on inline using their numbers…” Model output Check <think></think>Air molec ules scatter sunlight in every direction [1]. Bl ue light is scattered mo re than red light [2]. B lue light is scattered m ore than red light [2]. Blue light is scattered more than red light [2]. Blue light is scattered 48 96 144 192 240 Loop Copy 1 Copy 2 Copy 3 What you see Why is the sky blue? W wikipedia.org 1 N noaa.gov 2 Air molecules scatter sunlight in every direction [1]. Blue light is scattered more than red light [2]. 3.1s · Web + on-device Removed from the reply Markup: <think></think> Repeats: 122 characters One reply in Brello 1.0, checked as it streams On the left, the system prompt asks the model to say when it is not sure and, with web results, to cite sources inline. On the right, the model’s raw output streams in rows of 48 characters, each shown as two lines of 24, and a check runs at the end of each row. A leaked control token, think tags, is removed. The checks at 48, 96, 144 and 192 characters find no loop. At 240 characters the same 49-character sentence has appeared three times, so the reply is cut after the first copy and the model is stopped. Then the raw output gives way to the reply as the person sees it: the question, source cards for sources 1 and 2 above the answer, the answer with citations 1 and 2, and the meta line, 3.1 seconds, web plus on-device. Below a divider, what was removed: the think tags and 122 characters of repeats. System prompt “…If you are not sure about something, say so instead of guessing.” With web results: “…cite the sources you rely on inline using their numbers…” Quoted Model output Check <think></think>Air molec ules scatter sunlight in every direction [1]. Bl ue light is scattered mo re than red light [2]. B lue light is scattered m ore than red light [2]. Blue light is scattered more than red light [2]. Blue light is scattered 48 96 144 192 240 Loop Copy 1 Copy 2 Copy 3 What you see Why is the sky blue? W wikipedia.org 1 N noaa.gov 2 Air molecules scatter sunlight in every direction [1]. Blue light is scattered more than red light [2]. 3.1s · Web + on-device Removed from the reply Markup: <think></think> Repeats: 122 characters
  1. The reply starts from a short instruction. The model is asked to admit uncertainty and to cite web results, but it can ignore both, so the next steps are checks in code.
  2. The reply is checked as it streams. Every 48 characters, Brello looks for a block repeated 3 or more times over 120 or more characters. Four checks pass.
  3. Leaked markup is removed. Control tokens such as <think> or <end_of_turn> are stripped or routed to the Thought process panel, so they never reach the answer.
  4. A loop is caught at the next check. One 49‑character sentence has now appeared three times, so Brello cuts the reply after the first copy and stops the model.
  5. You see the cleaned reply, with its sources. Source cards sit above the answer, and each citation links to its page. The checks make a reply tidier, not correct.
Figure 1One reply in Brello 1.0, checked as it streams. The instruction text, the 48-character interval, the three-copy rule and the 120-character minimum are the app’s own; the question, model output, sources and timing are illustrative. Notice that the instructions work only if the model follows them, while the checks below them act on the text whatever the model does.

None of these makes a small model correct. On-device models are far smaller than frontier cloud models: they can be wrong, have a knowledge cutoff and are weaker at long or complex reasoning. Why AI makes things up explains the underlying problem, and section 2 of Evaluate first, then ship maps where each safeguard acts in the life of a reply.

How we intend to evaluate Brello Super Intelligence

Planned

We intend to evaluate Brello SI in six areas, each re-tested with every release. Because Brello is designed so that we don’t see conversations, the evidence has to be gathered before release, from task sets, synthetic profiles and attacks written for the purpose.

Table 2The six areas in which we intend to evaluate Brello SI. Design intent; no results exist.
AreaThe question it answersHow we intend to test it
Capability on real tasksDoes it do the work people bring to it, at the quality it implies?Task sets built around research with sources, writing and multi-step tasks, written for the purpose and never taken from conversations. Checked automatically where possible, and by people otherwise.
Honesty and calibrationDoes it say it isn’t sure when it should, and only then?Comparing its confidence with its accuracy, and checking that each citation supports the sentence it is attached to.
Privacy leakageDoes it reveal personal information it shouldn’t?Privacy audits that seed synthetic profiles with made-up personal details, then search every output and every outbound request for them.
Misuse and harmCould someone use it to hurt other people, or themselves?Red-teaming by people and by language models, before release.
Indirect prompt injectionDoes it follow instructions planted in what it reads?Pages seeded with instructions, with any change in behaviour counted as a failure.
Safety of actionsDoes every consequential action wait for you?Checking that the confirmation appears every time, shows exactly what will happen and can’t be bypassed by any phrasing or web page.

When we publish a result, it will give the method, the sample size, the date and which direction is better. AI safety evaluations, explained introduces these methods in general terms, and section 3 of Evaluate first, then ship sets out each area in full.

Release gates

Planned

Each Brello SI release, and each new capability within one, is intended to pass four gates in order, with pass criteria written before testing starts. A failure at any gate will send the release back to the first gate, not to the one it failed, because a fix for one problem can cause another.

Table 3The four release gates intended for Brello SI, in order.
GateTo pass, the release will need
1 Internal evaluationsAutomated suites and human review across all six areas, with every threshold met, no unexplained regression and no hard stop.
2 Adversarial red-teamingEvery finding from people and automated methods trying to break it resolved: fixed, mitigated, or accepted as a known limit with a written reason.
3 Invited early accessA small group who know they are using an early system, with a clear way to report problems. Serious reports resolved and known limits written down.
4 Wider availabilityIts evaluation summary published. Reports from people using Brello SI and from independent researchers continue to be reviewed.

The criteria will be of three types: a threshold, the minimum acceptable result in an area; no regressions, meaning nothing gets worse than in the previous release without a stated reason; and hard stops, failures that block a release however good everything else is. Completing a consequential action without your confirmation will be a hard stop. Because the criteria are fixed before testing, the bar can’t drift towards whatever a release already does.

Indirect prompt injection: Brello 1.0 and the Brello SI design

Brello SI is being designed to treat text from outside as data, never as instructions. We intend to make that hold with four layers, because no single defence against indirect prompt injection is reliable alone.

Table 4Four layers against indirect prompt injection: what Brello 1.0 does today and what Brello SI is being designed to do.
LayerIn Brello 1.0Planned for Brello SI
Label retrieved text as dataWeb passages reach the model as a numbered block labelled “Web results (retrieved {date})”, each with its source number.The same: each passage in a marked, numbered block.
Keep data out of the instruction channelNot in Brello 1.0.Prompt assembly designed so retrieved text can’t reach the place instructions go.
Let people checkInline citations link each claim to the page it came from.The same, so a distorted answer can be traced to the page that caused it.
Confirm before actingBrello 1.0 can only write an answer.Sending, spending or deleting will wait for your confirmation, a check designed to sit outside the model.

In Brello 1.0 a page cannot make Brello take an action, because Brello 1.0 can only write an answer. A page can still distort the answer that draws on it and mislead the person who reads it. That answer, clipped to 600 characters, stays in the context of later replies in the same chat for as long as it is among the six most recent messages. The web passages themselves are used for that reply only: each reply opens a fresh model session. No defence against prompt injection is complete, and we don’t claim these layers remove the risk. Prompt injection, explained introduces the attack, and section 4 of Evaluate first, then ship gives the design in full.

Confirmation before consequential actions: Brello 1.0 and the Brello SI design

Brello SI is being designed so that it will not complete an action that spends money, sends a message on your behalf or deletes something until you have confirmed it, shown in plain terms. This is commitment 05 of the charter.

Table 5Consequential actions in Brello 1.0, and as planned for Brello SI.
ActionBrello 1.0Planned for Brello SI
Spending moneyNot possible.Will wait until you have seen the amount and confirmed it.
Sending a message on your behalfNot possible.Will wait until you have seen who receives what, and confirmed it.
DeletingSwiping left deletes one chat. “Delete all chats” asks first: “This permanently removes every conversation from this device.”Will wait until you have seen what will be deleted, and confirmed it.

We are designing the confirmation as a check outside the model, so that a model that has been misled could not complete the action on its own. A confirmation protects only people who read it, and one that appears too often invites approval by habit, so how often it appears is part of what we intend to evaluate.

Saying when it isn’t sure

Brello 1.0 tells its model to admit uncertainty and to cite its web sources, and we intend to evaluate whether Brello SI does both. This is commitment 06 of the charter.

Table 6The two instructions in Brello 1.0’s system prompt that bear on accuracy, quoted exactly.
InstructionText in the system promptWhen it is included
Uncertainty“If you are not sure about something, say so instead of guessing.”Every reply
Citations“Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2]. If the results do not answer the question, say so briefly and answer from general knowledge.”Replies that use web results

An instruction is not a guarantee, and a small model can still be wrong. For Brello SI, each evaluation summary will report how often it says it isn’t sure when it should.

What we will publish

We will publish an evaluation summary with each Brello SI release and with each new capability in any Brello release, and we will publish Brello SI’s architecture before anyone outside the team uses it.

Table 7What we publish about safety, and its status on 5 October 2026.
PublicationWhat it containsStatus
Evaluation summary, with each releaseWhat was tested and how, the pass criteria set before testing, the results including failures, and what the release still gets wrong. Anything withheld, such as a test that would work as instructions for an attack, is named with the reason.None yet: no Brello SI release exists.
Brello SI architectureWhat runs where, and what each layer can and cannot see.Before anyone outside the team uses Brello SI (Charter, section 4).
Brello 1.0 system cardWhat the app does, the requests it makes and the permissions it holds.Published
Research notesOur methods and designs, with references.Published: Evaluate first, then ship; Asking before going online

Reporting a vulnerability

Report safety and security issues through our security and disclosure page, which sets out what to include and the safe harbour for research done in good faith. A machine-readable contact is published at /.well-known/security.txt.

Table 8What to report, and how.
If you findDo this
A vulnerability in Brello 1.0 or this websiteEmail our security contact with “Security report” in the subject.
A request from Brello 1.0 that our privacy page doesn’t listReport it as a security issue, the same way.
A safeguard on this page that doesn’t behave as describedReport it the same way, with the steps that reproduce it.

Further reading

Changelog

  1. Version 1.0. First published, for Brello 1.0, version 1.0.0.