Engineering · System description · Brello 1.0

Answering from the open web, without a server

How Brello 1.0 finds, reads and ranks web pages entirely on an Android phone and cites them, described stage by stage with the parameters the app uses.

Brello Research14 min readVersion 1.0

Abstract

Brello 1.0 runs its language models on an Android phone and has no server, so when a question needs fresh facts the phone must search the web itself. It sends only the search text, trying up to eight providers in order until one returns relevant results. It reads up to four result pages, keeps their article text, splits it into passages of about 420–700 characters and ranks them with BM25 (k₁ = 1.2, b = 0.75). At most two passages per source, within 3,400 characters (2,600 for Brello Core), reach the on-device model, which cites them inline. Search engines and sites still see each request. We describe the design and its parameters, not measurements of answer quality.

  • Only the search text and ordinary requests for up to four result pages leave the phone; there is no Brello server, proxy or relay, and the requests carry Do Not Track and Global Privacy Control headers.
  • Eight search providers are tried in a fixed order, two in an invisible on-device browser with 14-second timeouts and six over direct HTTP with 7-second timeouts; a provider that fails is rested for 2 minutes, and Wikipedia is never rested.
  • A relevance guard accepts a set of results only when at least a third of the top hits mention the question’s key terms.
  • Passages of about 420–700 characters are ranked with BM25 (k₁ = 1.2, b = 0.75), and at most two per source are kept, within 3,400 characters for Brello Pro and Brello Vision or 2,600 for Brello Core.
  • Web search is off by default: with it off, a time-sensitive question brings up a card titled “Search the web for this?”, and a question with a photo never triggers a search.
Contents13 sections

01

Overview: cited web answers with no server

Brello 1.0, the app for Android and iPhone made by Stuvio, runs its language models entirely on the phone. When a question needs fresh facts, it searches the open web from the phone as well, with no Brello server in between.

This paper describes that pipeline as it ships in version 1.0.0, stage by stage, with the parameters the app uses. The phone rewrites the question as a search query and tries up to eight search providers in turn until one returns relevant results. It reads up to four of the result pages, ranks short passages from them, and gives the best to the on-device model, which cites them inline as [1], [2] and so on.

We call this private AI web search, and the phrase needs a precise scope. The question, the text of the pages and the answer are processed on the phone. The search text and the page requests do leave it: they go directly to a search engine and to the sites in the results, which see a normal web request, including the phone’s IP address. Section 9 lists each privacy property and its mechanism. We describe the design and its parameters here; we have not published measurements of answer quality.

02

The constraint: a small model, a phone and no broker

Three constraints shape every stage of the pipeline: the model is small, all the work runs on a phone, and there is no server to broker the search.

The model is small. Brello Pro and Brello Vision are based on Gemma 4 E4B and E2B by Google, and Brello Core on Qwen3 1.7B by Alibaba. All three are released under Apache 2.0 and run locally with Google’s LiteRT-LM runtime.1 Brello is not affiliated with or endorsed by Google or Alibaba. Each model has a context window of 4,096 tokens, shared by the instructions, the conversation so far, any web results, the question and the answer. Like any language model, it also has a knowledge cutoff, so fresh facts have to come from somewhere else.

The work runs on a phone. Parsing and ranking pages compete with the interface for the same processor, so each step has to be cheap, and the heavy ones have to stay off the thread that draws the screen.

There is no broker. Retrieval-augmented generation, giving a language model retrieved text alongside the question, is an established technique.2 Brello 1.0 has no backend, so the phone uses the web’s ordinary interfaces instead, the same results pages and article pages a browser requests, and does the reading and ranking itself. Two older ideas from information retrieval make that practical. Boilerplate removal separates an article from the page around it,3 and BM25, a ranking function developed for the Okapi system at TREC-3 in 1994, scores how well a passage matches a query from word counts alone, without a neural network.4

03

The pipeline at a glance

Six stages run between a question and a cited answer. Two of them contact other computers, and in both the phone makes the request directly: the search text goes to a search engine, and ordinary page requests go to the sites in the results. Figure 1 follows one question through all six.

  1. You ask a question. Web search is switched on, so Brello can use fresh sources. Everything inside the dashed line runs on the phone.
  2. The phone writes a search query. Instructions meant for the answer, such as “Keep it short”, are removed before anything is sent.
  3. The search text leaves the phone. It goes straight to a search engine, with no Brello server in between. The top six results become sources [1] to [6].
  4. Up to four pages are requested and read in parallel. The requests go directly to the sites, and reader mode keeps the article text, split into passages of about 420 to 700 characters.
  5. Every passage is scored with BM25. At most two passages per source are kept, so one long page can’t crowd out the others.
  6. The on-device model writes the answer. The kept passages, up to 3,400 characters for Brello Pro and Brello Vision, go into its prompt, and the answer cites them by number.
Figure 1How Brello 1.0 answers a question from the web, in six stages. Only the search text and the page requests leave the phone. The sources, scores and answer are illustrative; the steps, limits and parameters are the ones the app uses.

Parsing pages and ranking passages are the heaviest stages, so both run on background isolates: threads with their own memory in Dart, the language Brello is written in.5 Running them off the interface thread is meant to keep the interface at the display’s 60 or 120 frames per second; we have not published frame-rate measurements. The reply’s status line moves through “Searching the web”, “Reading 4 sources” and “Thinking”, and the finished reply records its time and route, for example “3.1s · Web + on-device”.

04

Deciding whether to search

Brello 1.0 searches only after the person turns web search on in Settings or the + menu, or chooses “Search the web” on the card described below, which runs that search and turns web search on for later questions. Even then, it skips questions that don’t need the web. Web search is off by default.

With search on, Brello skips it for messages under four characters, pure arithmetic, short small talk such as “hi” or “thanks”, and writing tasks such as “translate” or “brainstorm”, unless they mention something fresh. With search off, a question that looks time-sensitive, for example one about news, prices, weather, scores, schedules or exchange rates, or one that says “latest” or “this week”, pauses the reply at “Needs the web”. A card titled “Search the web for this?” appears, with two buttons, “Search the web” and “Answer offline”, and choosing search turns web search on. A message with a photo never triggers a search. Asking before going online describes these rules and the card in detail.

When a search goes ahead, the question is rewritten as search text (Table 1). Instructions about the answer’s form are removed, contractions are expanded, and “What’s new in X” becomes “latest X news”. A time-relative question without a year gains the current month and year, except for live data such as weather, scores, prices, stocks, rates and traffic. A follow-up of four words or fewer, or one that refers back with “he”, “it” or “that”, gets the first ten words of the previous question in front of it.

05

Finding results without a broker: the provider chain

Brello 1.0 has no search API of its own. The phone tries up to eight providers in a fixed order and stops at the first that returns relevant results.

The first two providers load a full results page in an invisible browser on the phone: Android’s system WebView,6 used as a single hidden tab for one search at a time. A browser engine receives the same results a person would, where a bare HTTP request might be refused. The browser warms up on a blank page, with no network request, three seconds after a chat opens and when the person starts typing, but only while web search is on. After three idle minutes it is destroyed, and its cookies, cache and storage are wiped. The other six providers are direct HTTP requests. Figure 2 follows a hypothetical search through the chain.

Show data
#ProviderMethodTimeout
1DuckDuckGo (full results page)Invisible browser on the phone (system WebView)14 s
2BingInvisible browser on the phone14 s
3DuckDuckGo LiteDirect HTTP7 s
4DuckDuckGo HTMLDirect HTTP7 s
5Google News RSS (last 7 days; time-sensitive questions only)Direct HTTP7 s
6Brave SearchDirect HTTP7 s
7BingDirect HTTP7 s
8Wikipedia (search API; never rested)Direct HTTP7 s

Four rules keep the chain fast and predictable:

  • A provider that fails is rested. A provider that fails or is blocked is skipped for 2 minutes, so later questions don’t wait on it. Wikipedia is never rested, so the chain always ends with a provider it can ask.
  • Irrelevant results are rejected. Results are de-duplicated, and at least a third of the top hits must mention the question’s key terms. This rejects, for example, dictionary pages for “why” when the question was “why is the sky blue”.
  • Human checks end the attempt. When a provider responds with a check such as “unusual traffic” or “are you a robot”, Brello moves on to the next provider rather than trying to get past it.
  • News feeds are used only for news. Google News RSS, limited to the last seven days, is tried only for time-sensitive questions.

Results are cached in memory for 15 minutes per query, so a repeated search isn’t sent again. The cache is never written to storage and is gone when the app closes.

06

Reading pages: reader-mode extraction

Brello reads each result page the way a browser’s reader view does: it keeps the article text and discards everything around it. Figure 3 follows one page from the fetch to the passages that reach the model.

The top six results become the sources numbered [1] to [6]. The phone fetches up to four of those pages in parallel, each with a 7-second timeout. It accepts only HTML or plain text and stops reading a page at 1.5 MB, so one heavy page can’t stall the phone. Google News results are redirects, so for those Brello uses the headline and date rather than following the link.

Reader-mode extraction then strips navigation, headers and footers, adverts, cookie banners, share widgets, comments, reference lists and Wikipedia edit links. It keeps the main article text, removes duplicated blocks and stops at about 24,000 characters per page. Finally, the text is split into passages of about 420 to 700 characters. At those lengths, between four and eight passages fit in the 3,400-character budget described in Section 8.

07

Ranking passages with BM25

Every passage is scored against the search text with BM25, using and , values within the range that standard textbooks report as reasonable.7 BM25 rewards passages that contain the query’s words, especially rare ones, without letting any one word dominate.

For a passage D and a query Q, BM25 adds up a contribution from each query term q (Equation 1).8 The contribution grows with the number of times the term appears in the passage, f(q, D), with diminishing returns. It is weighted by the term’s inverse document frequency, IDF(q), so rare words count for more than common ones.9

Equation 1BM25, with k1 = 1.2 and b = 0.75 in Brello 1.0. f(q, D) is the number of times term q appears in passage D, |D| is the passage’s length, and avgdl is the average length of the passages being ranked. The fraction is the term-frequency factor, TF(q, D), and the bracketed term in its denominator is the length normalisation, L(D). Implementations differ in the exact form of IDF(q); the term-frequency fraction is the same in all of them.

Two parameters shape the fraction on the right. k1 sets how quickly repeated mentions stop adding to the score: however many times a term appears, it can’t contribute more than k1 + 1 times its IDF. b sets how strongly the passage’s length counts, from 0 (ignored) to 1 (fully normalised). Figure 4 computes the fraction for the values Brello uses.

Show data
Mentions, f(q, D)Short (0.75 × average)Average lengthLong (1.25 × average)Added by this mention (average)
00.000.000.00–
11.111.000.91+1.00
21.481.381.28+0.38
31.661.571.49+0.20
41.771.691.62+0.12
51.841.771.71+0.08
61.891.831.78+0.06
71.931.881.83+0.04
81.961.911.87+0.03
91.981.941.90+0.03
102.001.961.93+0.02
Limit as f grows2.202.202.200

Saturation matters most for questions with several key terms. Take two passages of average length, and assume “sky” and “blue” have the same IDF. A passage that mentions “sky” ten times scores 1.96 IDF for that term. A passage that mentions each word once scores 1.00 + 1.00 = 2.00. Before any adjustment, BM25 already prefers the passage that covers more of the question.

Brello adds two adjustments to the BM25 score: a bonus for passages that cover more of the question’s distinct terms, and a small bonus for lead paragraphs. BM25 itself matches words, not meanings, so a passage that answers the question in different words scores only on the words it shares with the query.

08

Writing a cited answer within the context budget

Brello keeps at most two passages from each source, within a budget of 3,400 characters for Brello Pro and Brello Vision, which are built on Gemma 4, and 2,600 characters for Brello Core, built on Qwen3. The on-device model then writes the answer from those passages and cites them by number.

The per-source cap means one long page can’t crowd out the others, and the budget is the tighter limit: four pages at two passages of up to 700 characters could supply 5,600 characters, more than either budget allows. A source with no good passage falls back to its search snippet, with 900 characters shared among all such snippets. The kept passages reach the model as a compact, numbered block labelled Web results (retrieved …), each tagged with its source number.

The budget is small because everything shares one context window (Table 2). Every reply also opens a fresh model session, so web results from an earlier question never take up room in the next.

The citation instruction is the only change web results make to the system prompt, apart from the date. It reads, verbatim:

The system prompt is otherwise short on purpose, because small models tend to copy the shape of long instructions. Small models copy the shape of their instructions explains why. As the answer streams in, citation markers become tappable links, and source cards above the answer show each site’s letter mark, domain, citation number and page title (Figure 5). Tapping a card opens the page in the phone’s browser.

Brello 1.0 running Brello Vision, answering a question about New York politics from the web: source cards for cbsnews.com and nydailynews.com sit above an answer with numbered citations, and the composer shows the Web search chip.
Figure 5A web answer in Brello 1.0, running Brello Vision. The source cards above the answer carry citation numbers, and the same numbers appear as blue links in the text. The coloured letters on the cards are drawn on the phone, and the glow around the composer shows that the model is still writing.

The letter marks are a privacy decision. Fetching each site’s icon would mean another request to every site in the results, whether or not the person opens it. Brello draws a monogram on the phone instead, in a colour derived from the domain name, so no icon requests leave the phone.

09

Privacy properties

The pipeline sends two kinds of request, the search text to a search engine and page requests to up to four sites, and nothing to Stuvio. Table 3 lists each property and the mechanism that provides it.

These properties have limits. The search engine and the sites in the results still receive each request, including the phone’s IP address, as they would for any ordinary web request. The two headers state a preference; they don’t stop a site from receiving the request. Brello doesn’t hide the request, and it doesn’t add anything of its own. Privacy at Brello covers the rest of the app, and the privacy policy is the formal statement.

10

Limitations

The pipeline trades breadth for privacy and speed, and it fails in ways we can name.

  • It depends on third-party search engines. A search can take up to 14 seconds for each browser-based attempt, or come back empty. When it does, Brello answers from the model’s own knowledge.
  • The evidence is small. A few thousand characters stand in for whole documents, so anything outside the kept passages is invisible to the model.
  • Ranking is lexical. BM25 matches words, not meanings. A passage that answers in different words can rank below one that merely repeats the query.
  • The relevance guard is a heuristic. Requiring a third of the top hits to mention the key terms filters obvious mismatches. It says nothing about whether a page is accurate.
  • Small models make mistakes. The model can misread a passage, cite the wrong source or state something no source supports. The prompt tells it to say when the results don’t answer the question, which is an instruction, not a guarantee. We have not published measurements of how often this happens.
  • Pages are untrusted input. A page can contain text written to steer a model, known as indirect prompt injection.12 Brello 1.0 labels retrieved text as web results and tells the model to base its answer on them, but a small model may still follow instructions it reads.
  • News results are thin. Google News links are redirects, so only a headline and a date reach the model.

11

What comes next

Brello Super Intelligence is in development. This section describes intent, not results.

Brello Super Intelligence (Brello SI) is being designed to keep the open web as a layer it uses only on request, as Brello 1.0 does, and to run heavier work on sealed compute that the device verifies before sending anything. Private compute you can verify: the design space sets out that design.

Questions this paper leaves open carry over to Brello SI: how large an evidence budget a model can use well; how ranking can match meaning as well as words, as dense retrieval models do,13 within a phone’s memory; and how a reader can check that a cited passage supports the sentence that cites it. We don’t yet know the answers. The Brello Charter commits us to evaluate new capability before it ships, and to publish what we evaluate and what we find.

12

References

  1. Google AI Edge. “LiteRT-LM.” GitHub repository. github.com/google-ai-edge/LiteRT-LM (accessed 5 October 2026).
  2. P. Lewis et al. “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” Advances in Neural Information Processing Systems 33 (NeurIPS 2020). arXiv:2005.11401
  3. C. Kohlschütter, P. Fankhauser and W. Nejdl. “Boilerplate Detection Using Shallow Text Features.” Proceedings of the Third ACM International Conference on Web Search and Data Mining (WSDM ’10), pp. 441–450, 2010. doi:10.1145/1718487.1718542
  4. S. E. Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu and M. Gatford. “Okapi at TREC-3.” Proceedings of the Third Text REtrieval Conference (TREC-3), NIST Special Publication 500-226, 1994. trec.nist.gov/pubs/trec3
  5. Dart. “Concurrency in Dart.” Dart documentation. dart.dev/language/concurrency (accessed 5 October 2026).
  6. Android Developers. “WebView.” Android API reference. developer.android.com/reference/android/webkit/WebView (accessed 5 October 2026).
  7. C. D. Manning, P. Raghavan and H. Schütze. Introduction to Information Retrieval, section 11.4.3, “Okapi BM25: a non-binary model.” Cambridge University Press, 2008. nlp.stanford.edu/IR-book
  8. S. Robertson and H. Zaragoza. “The Probabilistic Relevance Framework: BM25 and Beyond.” Foundations and Trends in Information Retrieval, 2009. doi:10.1561/1500000019
  9. K. Spärck Jones. “A Statistical Interpretation of Term Specificity and Its Application in Retrieval.” Journal of Documentation 28(1):11–21, 1972. doi:10.1108/eb026526
  10. W3C. “Tracking Preference Expression (DNT).” W3C Working Group Note, 17 January 2019. w3.org/TR/tracking-dnt
  11. W3C. “Global Privacy Control (GPC).” W3C Working Draft, 24 September 2026. w3.org/TR/gpc (accessed 5 October 2026).
  12. K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz and M. Fritz. “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23), pp. 79–90, 2023. doi:10.1145/3605764.3623985
  13. V. Karpukhin et al. “Dense Passage Retrieval for Open-Domain Question Answering.” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 6769–6781, 2020. doi:10.18653/v1/2020.emnlp-main.550

Cite this work

Brello Research. “Answering from the open web, without a server.” Stuvio, 5 October 2026. https://brello.ai/research/on-device-web-answers/

@misc{brello2026ondevice,
  title  = {Answering from the open web, without a server},
  author = {{Brello Research}},
  year   = {2026},
  month  = {oct},
  url    = {https://brello.ai/research/on-device-web-answers/},
  note   = {Stuvio}
}

Version history

  1. 1.0First published.