01
Overview
Language models small enough to run on a phone copy the shape of their instructions, not only their content. A sentence that describes the parts of a good answer tends to come back as the answer’s headings.
Brello 1.0 runs three such models entirely on the phone. Brello Pro and Brello Vision are built on Gemma 4 E4B and E2B by Google, and Brello Core on Qwen3 1.7B by Alibaba. All three run locally through LiteRT-LM, Google’s on-device inference runtime.9 While we were building the app, a system prompt that asked for “the direct answer, then useful detail” produced replies under the headings Direct Answer: and Useful Detail: instead of plain prose.
This paper describes what we changed in response, with the parameters Brello 1.0 ships with. The system prompt is four sentences that set a register rather than a layout (section 4). The context is rebuilt from capped parts for every reply, inside a 4,096-token window that the answer shares (section 5). The rules that must always hold, such as stopping a repetition loop, run in code on the model’s output (section 7).
02
Background: how a model reads a prompt
A language model reads its prompt as one sequence of tokens to continue, so the wording and format of the prompt are evidence about what should come next. The system prompt, the conversation and the question have no separate channels.
Research on in-context learning shows how strongly format carries. Brown et al. showed that a large language model can perform a new task from a few examples placed in the prompt, without any change to its weights.1 Min et al. found that much of what those examples contribute is the label space, the distribution of inputs and the overall format of the sequence, rather than whether their labels are correct.2 Sclar et al. measured differences of up to 76 accuracy points from formatting changes alone in few-shot settings, using LLaMA-2-13B. The sensitivity remained with larger models, more examples and instruction tuning.3
An instruction that names the parts of an answer is a description of a format. Our working explanation for what we saw is that a small model treats such a description much as it treats an example: as a template to fill in. We have not tested that explanation directly, and section 10 lists it as an open question.
03
Instruction echo: what we observed
We saw small models turn descriptive instructions into literal structure. We call this instruction echo: labels, headings or phrasing in a reply that come from the system prompt rather than from the question.
The clearest case came from a prompt that asked the model to “lead with the direct answer, then useful detail”. Replies came back as Direct Answer: … Useful Detail: … In a narrow sense the instruction worked, because the direct answer did come first. But the model had turned a description of a good reply into a pair of headings and filled them in, so the reply read like a completed form.
Brello’s prompt asks the model to answer “like a knowledgeable friend”, and a form is not how a friend answers, even when every fact in it is right. Figure 1 reconstructs the effect with example text. On the left is a prescriptive prompt written as a list of rules. On the right is the prompt Brello 1.0 ships with, quoted exactly.
- A prescriptive prompt names the parts of an answer. This illustrative prompt is a list of rules, and two of them describe sections: “the direct answer” and “useful detail”.
- The person asks a simple question. Nothing in it asks for headings or sections.
- The reply copies the prompt’s shape. The two phrases come back as headings, so a correct answer reads like a completed form.
- Brello 1.0’s prompt names no parts of an answer. Its four sentences, quoted exactly, set a name, a register, brevity and honesty about uncertainty.
- The same question gets a plain answer. With no layout to copy, the reply takes the shape the question needs.
The lesson we took is that, with a small model, every word in the context is a candidate for the output. The rest of this paper follows from that: say less in the prompt, add context only when a question needs it, and keep hard rules out of the prompt altogether.
04
A four-sentence system prompt
Brello 1.0’s system prompt is four sentences and 254 characters long, and it is short on purpose. It sets a name, a register, a call for brevity and a rule about uncertainty. It names no parts of an answer.
// Brello 1.0 system prompt, verbatim
You are Brello, a helpful assistant that runs privately on the user's phone. Reply to the user's message directly and naturally, like a knowledgeable friend. Keep it clear and to the point. If you are not sure about something, say so instead of guessing.
Each sentence has one job. The first gives the model a name and a setting. The second sets a register: “directly and naturally” describes how a reply should sound, and names nothing that could become a heading. The third asks for brevity. The fourth asks the model to admit uncertainty, which is an instruction, not a guarantee.
What the prompt leaves out matters as much. It has no formatting rules, no list of things to avoid and no example replies. Answers render as Markdown, so the model can still use a list or bold text when the content suits it. Figure 2 shows a real reply from Brello 1.0 that does exactly that.

Lines added only when needed
Some questions need context that four sentences can’t give. Brello adds it one plain sentence at a time, and only when the question calls for it:
// Only when the person asks about Brello’s identity
You are {Model}, based on {base model}.
// Only for questions about time, or when web results are present
Current date: {weekday, Month day, year}.
// Only when web results are present
Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2]. If the results do not answer the question, say so briefly and answer from general knowledge.
The identity line is the clearest case. Someone who asks which model they are talking to should get an accurate answer, such as “Brello Core, based on Qwen3 1.7B”. Most questions aren’t about identity, though, and a name and lineage repeated in every prompt is one more thing a small model could bring up unprompted.
The date follows the same rule. A model has no clock, so a question about time needs the date supplied, and so does an answer built on web results, which are labelled with the date they were retrieved. The citation instruction appears only with web results, since without sources it has nothing to refer to. Its last sentence tells the model to say briefly when the results don’t answer the question, and then to answer from general knowledge. How those results are found and ranked on the phone is covered in Answering from the open web, without a server.
05
Assembling the context for each reply
Every reply in Brello 1.0 opens a fresh model session, and Brello rebuilds the context from parts it controls: the system prompt, a clipped slice of the conversation, web results when there are any, and the new question. Nothing carries over inside the model from one reply to the next.
The reason is the size of the window. Brello’s models run with a context window of 4,096 tokens, where a token is a short piece of text, often part of a word. The answer is generated into the same window as the prompt: up to 1,200 tokens normally, or 2,048 with Think harder, the mode in which the model reasons before it answers. A long-lived session would let earlier web results pile up until they crowded the window. A fresh session means old web context never takes up space. Figure 3 builds the window for one reply, part by part.
- Each reply opens a fresh session, starting with the system prompt. It is four sentences and 254 characters, with an identity, date or citation line added only when needed.
- Up to six earlier messages follow, clipped. Yours are cut to 300 characters and Brello’s to 600, so the conversation’s share of the window is bounded.
- Web results are added only after a search. The ranked excerpts are capped at 3,400 characters on the Gemma 4 models and 2,600 on Qwen3.
- The new question goes in last. On Brello Pro and Brello Vision, it can come with a photo.
- The answer is written into the same window. It may use up to 1,200 tokens, 29% of the 4,096, drawn here to scale.
- With Think harder, the answer may use up to 2,048 tokens. That is half the window, shared with everything above it.
Show data
| Part of the window | Limit in Brello 1.0 | Unit | In Figure 3 |
|---|---|---|---|
| System prompt | 4 sentences (254 characters), plus an identity, date or citation line only when needed | Characters | Schematic |
| Earlier messages | Up to 6. Yours clipped to 300, Brello’s to 600 | Characters | Schematic; 300 and 600 to scale with each other |
| Web results | Only after a search. Up to 3,400 (Gemma 4) or 2,600 (Qwen3) | Characters | Schematic; 3,400 and 2,600 to scale with each other |
| Question | The new message, with a photo on Brello Pro and Brello Vision | – | Schematic |
| Answer | Up to 1,200 (29% of the window) | Tokens | To scale |
| Answer with Think harder | Up to 2,048 (half the window) | Tokens | To scale |
| Context window | 4,096 | Tokens | To scale |
History is carried as plain text, under fixed rules:
| History rule | Brello 1.0 |
|---|---|
| Earlier messages included | Up to the 6 most recent |
| Your messages | Clipped to 300 characters |
| Brello’s replies | Clipped to 600 characters |
| Citation markers such as [1] | Removed |
| Photos from earlier turns | Noted as “[shared a photo]” |
Each rule removes something the model doesn’t need. A citation marker in an old reply would point at a source that is no longer in the window, so it is removed rather than left as a pattern to copy. A photo from an earlier turn is reduced to a short note, so the model knows one was shared. Clipping bounds the conversation’s share of the window: with three messages of each kind, the history comes to at most 2,700 characters.
A lean window has a second benefit. Long contexts aren’t used evenly: Liu et al. found that language models use information in the middle of a long input less reliably than information near its start or end.4 The cost is a short memory. Six clipped messages won’t hold a detail from much earlier in a long chat, and we treat that as a known limit of Brello 1.0.
06
Sampling settings, as shipped
Brello 1.0 samples Brello Pro and Brello Vision at temperature 1.0, top-k 64 and top-p 0.95, and Brello Core at temperature 0.7, top-k 20 and top-p 0.8. With Think harder on, it uses temperature 0.6 and top-p 0.95.
A model scores every possible next token, and a sampler chooses one. Temperature rescales the scores before they become probabilities. Top-k keeps only the k most likely tokens.5 Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least p.6 Figure 4 gives both in standard notation.
We record the values below for reference, as shipped, not as recommendations:
| Models | Temperature | Top-k | Top-p |
|---|---|---|---|
| Brello Pro and Brello Vision (Gemma 4) | 1.0 | 64 | 0.95 |
| Brello Core (Qwen3) | 0.7 | 20 | 0.8 |
Brello Core samples more narrowly than the Gemma 4 models on all three settings, and Think harder lowers the temperature further. Which model Brello recommends for a given phone depends on its memory, as we explain in Fitting a model to the phone in your pocket.
Narrow sampling favours the most likely continuation, and that has a known failure mode. Holtzman et al. showed that decoding which maximises likelihood tends towards bland, repetitive text, and proposed sampling from the nucleus instead.6 Sampling settings alone don’t stop a small model from looping, so Brello 1.0 also watches what the model writes.
07
Guards that act on the output
Brello 1.0 enforces its hard rules in code, on the text the model produces, rather than by asking the model to behave. There are three: a repetition-loop guard, a cleanup for leaked markup, and fallbacks for replies that come back empty.
Repetition loops
Small models sometimes get stuck. A sentence comes round again, then again, and the reply becomes a loop that would run on towards the length limit. Xu et al. describe why loops persist: the more times a sentence has already been repeated in the context, the more likely the model is to generate it again.7
Brello 1.0 checks the stream as it arrives. Every 48 characters, it asks whether the tail of the text is looping: a block repeated three or more times over at least 120 characters. If it is, Brello cuts the reply after the first copy of the block and stops the model. Figure 5 follows one reply through the guard.
- The reply streams in and is checked every 48 characters. Each row here is one interval, and at each check Brello asks whether the tail of the text is looping.
- The model starts repeating itself. From 240 characters the same block comes round again, but fewer than three copies don’t meet the rule, so the checks pass.
- At 528 characters, the third copy is complete. A block repeated three or more times over at least 120 characters meets the rule, so this check finds a loop.
- Brello cuts the reply after the first copy and stops the model. The reply keeps one clean copy, and generation ends instead of running on towards the length limit.
The rule’s shape decides what it catches. Requiring three copies means a single deliberate repeat passes. The 120-character floor means a word or a short phrase repeated for effect passes too. Checking every 48 characters, rather than after every token, keeps the work small, and a loop that meets the rule is caught within 48 characters. Cutting after the first copy, rather than where the loop was found, keeps one copy of what the model was repeating, so the reply ends where the repetition began.
The guard is covered by unit tests in the app, alongside the code that separates reasoning from the reply.
Leaked markup and empty replies
Chat models are trained with special markers that delimit turns, separate reasoning from the reply and end a sequence. They are meant for the runtime, not the reader. A small model sometimes writes them into a reply as ordinary text, so Brello 1.0 removes them, or uses them to route reasoning to the right place. The markers it handles are:
<think>, which opens a block of reasoning in Qwen3’s thinking mode;8/no_think, a flag that Qwen3 recognises for switching that mode off;8<|im_end|>,<end_of_turn>and<eos>, which end a turn or a sequence in the models’ chat formats;- Gemma “channel” markers.
Text that the markers identify as reasoning goes to the Thought process panel instead of the answer. With Think harder off, a model can still produce a thought and then stop without answering. When that happens, Brello shows the thought as the answer rather than an empty reply. If the model produced nothing at all, the reply says so plainly: “No response was generated. Try rephrasing.” The design of the panel is covered in Designing an intelligence you can see working.
| What the model produced | What the person sees |
|---|---|
| A block repeated three or more times over at least 120 characters | The reply up to the end of the first copy. The model is stopped. |
| Stray control markup | Nothing: it is removed, or routed to the Thought process panel |
| Reasoning, with Think harder on | The Thought process panel, then the answer |
| A thought but no answer, with Think harder off | The thought, shown as the answer |
| Nothing | “No response was generated. Try rephrasing.” |
None of these rules depends on the model’s cooperation. They run in code on every reply, whichever of the three models wrote it.
08
Privacy and safety properties
The context assembly, generation and guards described here all run on the phone. The system prompt, the clipped history, any web excerpts, the question and the answer are assembled and processed by the on-device runtime, with no Brello server at any step.
- The context stays on the phone. Each reply’s context goes only to the model running on the device, so your questions, photos and answers are processed locally.
- History stays on the phone, with one exception. Earlier messages come from the conversation itself, which is stored only in Brello’s private storage and excluded from backups. When web search is on, a short or referential follow-up adds the first ten words of the previous question to the search text sent to the search engine.
- Web excerpts need a search you allowed. They enter the context only after the person turns on web search in Settings or the + menu, or chooses “Search the web” on the card, which runs that search and turns web search on for later questions. Even then, only the search text goes to a search engine, and result pages are requested directly from their websites, which see those requests as they would any other. Reading and ranking happen on the phone.
- Identity is supplied when asked. If a person asks which model they are talking to, the identity line gives the model its name and the open model it is based on.
- Hard rules don’t rely on instructions. Loop cutting, markup cleanup and empty-reply handling run in code, whatever the model does.
The full account of what Brello 1.0 keeps on the phone, and what leaves it, is in the privacy policy.
09
Limitations
This paper reports observations and design decisions, not measurements. Several limits follow from that, and from the design itself.
- No controlled study. We have not measured how often instruction echo occurs on each of the three models, or how much the shorter prompt reduces it. Figure 1 is a reconstruction with example text.
- A short memory. Only the six most recent earlier messages are carried into each reply, clipped to 300 or 600 characters. Details from earlier in a long chat are lost.
- Limits in characters, a window in tokens. The history and web limits are set in characters, while the window is measured in tokens. The number of tokens a passage takes depends on the tokenizer and on the text, so the prompt’s share of the window varies.
- The guard reads shape, not purpose. Repetition that a person asked for, such as a long refrain repeated three times, can meet the rule and be cut. A loop whose copies differ from one another may not be caught.
- Instructions are not guarantees. The prompt asks the model to say when it isn’t sure. A small model can still be wrong without saying so.
10
Open questions for Brello Super Intelligence
Brello Super Intelligence is in development. This section describes our intent, not results.
The main lesson from Brello 1.0 is that, with a small model, the prompt is part of the model’s behaviour rather than a neutral wrapper around it. We intend to carry that into Brello Super Intelligence, and to study these questions rather than assume answers to them:
- Why does echo happen, and does it fade with scale? Our working explanation, that a small model treats a description of a format as a template, is untested. Sclar et al. found that sensitivity to formatting remained as model size increased.3 We don’t yet know how much larger models echo the wording of their instructions, and we intend to measure it.
- How should echo be evaluated? We intend to treat every change to a system prompt the way we treat a change of model, with checks for labels, headings or phrasing that come from the prompt rather than the question. Our approach is set out in Evaluate first, then ship.
- How can a model get richer context without a template? Longer memory and more sources mean more text in the window, and more wording a model could copy. We are designing Brello Super Intelligence to add context only when a question needs it, and to check what each addition does to the shape of replies before it ships.
- Which rules belong in code? We are designing Brello Super Intelligence to keep guards such as loop cutting and markup cleanup in code that acts on the output, whichever model is running.
If some work moves beyond the phone, the privacy constraints we are designing for are described in Private compute you can verify.
11
References
- T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, et al. “Language Models are Few-Shot Learners.” Advances in Neural Information Processing Systems 33 (NeurIPS 2020). arXiv:2005.14165
- S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi and L. Zettlemoyer. “Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?” Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022). arXiv:2202.12837
- M. Sclar, Y. Choi, Y. Tsvetkov and A. Suhr. “Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.” International Conference on Learning Representations (ICLR 2024). arXiv:2310.11324
- N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni and P. Liang. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics 12 (2024): 157–173. arXiv:2307.03172
- A. Fan, M. Lewis and Y. Dauphin. “Hierarchical Neural Story Generation.” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL 2018): 889–898. arXiv:1805.04833
- A. Holtzman, J. Buys, L. Du, M. Forbes and Y. Choi. “The Curious Case of Neural Text Degeneration.” International Conference on Learning Representations (ICLR 2020). arXiv:1904.09751
- J. Xu, X. Liu, J. Yan, D. Cai, H. Li and J. Li. “Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arXiv:2206.02369
- Qwen Team (A. Yang et al.). “Qwen3 Technical Report.” arXiv:2505.09388, 2025. arxiv.org/
abs/ 2505.09388 - Google AI Edge. “LiteRT-LM.” Source code repository. github.com/
google-ai-edge/ LiteRT-LM
Cite this work
Brello Research. “Small models copy the shape of their instructions.” Stuvio, 5 October 2026. https://brello.ai/research/small-models-shape-of-instructions/
@misc{brello2026smallmodels,
title = {Small models copy the shape of their instructions},
author = {{Brello Research}},
year = {2026},
month = {oct},
url = {https://brello.ai/research/small-models-shape-of-instructions/},
note = {Stuvio}
}
Version history
- 1.0First published.



