How it works

Retrieval-augmented generation, explained

How a language model answers from passages retrieved at question time: the three stages, sparse and dense retrieval, what citations prove, where it fails, and RAG on a phone.

Brello Research12 min readVersion 1.0

Summary

Retrieval-augmented generation (RAG) connects a language model to an outside source of text at the moment a question is asked. A retriever finds passages relevant to the question, the best of them are added to the model’s prompt, and the model writes an answer from them, ideally citing each one. The approach, named by Lewis and colleagues in 2020, lets an answer draw on information newer than the model’s training and lets readers check its sources. It fails when retrieval misses, when sources are wrong or hostile, or when the model misuses what it is given. Brello 1.0 runs the whole pipeline on the phone, ranking web passages with BM25.

  • Retrieval-augmented generation answers from passages retrieved when the question is asked, so an answer can use sources newer than the model’s training data and cite them.
  • Sparse retrievers such as BM25 match exact words and need no training; dense retrievers match meaning with a trained encoder and an index of vectors.
  • A citation shows where a claim should come from, not that the source supports it: a 2023 study found only 51.5% of four generative search engines’ sentences fully supported by their citations.
  • Brello 1.0 keeps at most two passages of about 420–700 characters per source, within 3,400 characters for its Gemma 4 models and 2,600 for Qwen3, and asks the model to cite them as [1] or [2].
  • In Brello 1.0, ranking and generation run on the phone; only the search text and the requests for up to four result pages leave it, and the sites that receive them see a normal web request.
Contents8 sections

01

What is retrieval-augmented generation?

Retrieval-augmented generation (RAG) is a way of making a language model answer from documents fetched at the moment a question is asked. A retriever finds passages relevant to the question, the system adds them to the model’s prompt, and the model writes an answer grounded in them, usually citing each source.

The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, University College London and New York University 1. They paired a pre-trained sequence-to-sequence model, which they called parametric memory because its knowledge is stored in its weights, with a dense vector index of Wikipedia: a non-parametric memory that can be searched, and replaced, without retraining the model. REALM, published the same year by researchers at Google, built retrieval into the pre-training of a language model 2.

Today the term covers any system that retrieves text and places it in a model’s prompt, whether the retriever and the model were trained together, as in the original paper, or are separate parts joined only by the prompt. The text can come from a company’s documents, a database or the open web.

Retrieval addresses three weaknesses of a model on its own. Its knowledge stops at its training cut-off. It can produce fluent statements that no source supports, often called hallucinations (Why AI makes things up). And it cannot show where a claim came from. Retrieval supplies current text, and in knowledge-grounded dialogue it has been shown to reduce hallucinated facts 3. Citations let the reader check the rest.

02

Retrieve, rank, then generate

A RAG system answers in three stages: it retrieves candidate passages for the question, ranks them and keeps the best that fit the model’s prompt, then has the model generate an answer from them. Figure 1 follows one question through all three.

01Retrieve3 of 5 kept

Why is the sky blue?

Candidate passages, ranked by BM25

  • [1]Sunlight is scattered by the molecules in air, and blue light far more than red.
  • [1]Shorter wavelengths scatter more strongly, an effect named after Lord Rayleigh.
  • [1]The same scattering makes distant hills look faintly blue.Not kept: third from [1]
  • [2]Violet scatters even more, but sunlight has less of it and our eyes are less sensitive to it.
  • [3]Sky blue paint, ready to ship for any room.Not kept: low score

02Augment3 passages in the prompt

Instruction, quoted from Brello 1.0

“Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2]. If the results do not answer the question, say so briefly and answer from general knowledge.”

Web results (retrieved {date})

  1. [1]Sunlight is scattered by the molecules in air, and blue light far more than red.
  2. [1]Shorter wavelengths scatter more strongly, an effect named after Lord Rayleigh.
  3. [2]Violet scatters even more, but sunlight has less of it and our eyes are less sensitive to it.

Passage budget: 3,400 characters

Question

Why is the sky blue?

03Generate

Language modelIn Brello 1.0, on the phone

Air molecules scatter blue light far more than red [1]. Violet scatters even more, but sunlight has less of it and our eyes are less sensitive to it [2].

  1. A question arrives. On its own, a language model can answer only from what it learned in training, which may be out of date or wrong.
  2. The retriever finds and ranks candidate passages. Brello 1.0 ranks them with BM25 and keeps at most two per source, so the third passage from source [1] is dropped despite its score.
  3. The kept passages enter the prompt, numbered by source. An instruction, quoted from Brello 1.0, tells the model to use and cite them. For its Gemma 4 models, Brello allows up to 3,400 characters of passages.
  4. The model writes the answer from the passages. Each claim carries the number of the source it came from, like [1] or [2].
  5. Every citation leads back to a source. The reader can open source [1] and check the claim. A citation shows where a claim should come from; it does not prove that the source supports it.
Figure 1Retrieval-augmented generation in three stages. The passages, scores and answer are illustrative and shortened. The instruction and the label on the block of passages, ‘Web results (retrieved {date})’, are quoted from Brello 1.0, where {date} stands for the date the results were retrieved; the two-per-source limit and 3,400-character budget are the ones it applies with its Gemma 4 models.

Retrieve

Retrieval starts from the question, or from a search query written from it. Systems over a fixed collection, such as a company’s manuals, split the documents into passages ahead of time and index them; the original RAG system used Wikipedia split into passages of 100 words, about 21 million in all 1 4. Systems that answer from the web search at question time and split the pages they fetch. Passages, rather than whole documents, are the unit because a model’s context window is limited (Context windows, explained), and because a passage that answers the question directly is more useful to the model than a long page that mentions it once.

Rank

Candidates are scored against the question and only the best are kept. Sparse methods such as BM25 score the words a passage shares with the question; dense methods compare learned vectors (section 03). Systems typically keep a fixed number of passages or fill a fixed budget, and some cap how many may come from one source, so that a single long page cannot crowd out the rest. A second, slower model can re-rank the top candidates; in the BEIR benchmark, re-ranking gave the best zero-shot results, at a higher computational cost 5.

Augment and generate

The kept passages are inserted into the prompt, usually numbered, together with an instruction on how to use them and the question itself. The model then generates its answer token by token, with the passages in its context. Nothing in this step forces the model to rely on them: whether it does depends on its training and on the instruction, which is why citations, and checking them, matter (section 04).

03

Sparse and dense retrieval

Sparse retrieval matches the words of the question against the words of each passage; dense retrieval maps both into vectors with a neural encoder and matches them by meaning.

Sparse methods represent text as a vector with one dimension per word of the vocabulary, almost all of them zero, hence the name. BM25 is the standard example: it weighs each shared word by how rare it is, with diminishing returns for repetition and a correction for length 6. It needs no training, finds exact names, numbers and codes reliably, and its scores can be explained word by word. Its weakness is vocabulary mismatch: a passage that answers the question in other words scores nothing. BM25, explained works through the formula with an interactive example.

Dense methods train an encoder, typically a transformer, to map questions and passages into the same vector space, so that a question lands near the passages that answer it. Dense Passage Retrieval, the retriever behind the original RAG system, trained two such encoders on pairs of questions and passages 4. Dense retrieval can match paraphrases that share no words, at the cost of a trained model, a vector index and the computation to embed every passage. It can also generalise poorly: across the 18 datasets of the BEIR benchmark, BM25 was a robust zero-shot baseline that dense retrievers often failed to beat outside their training domain 5.

Hybrid systems run both and merge the results, using sparse scores for exact terms and dense scores for meaning. Brello 1.0 uses sparse retrieval only: BM25 runs on the phone’s background threads, alongside the page parsing already done there, with no second model and no index of vectors.

04

Why citations matter

Citations connect each claim in an answer to the passage it came from, so a reader can check it. They turn an answer from something to trust into something to verify.

Brello 1.0 asks for them directly. When web results are present, its system prompt adds: “Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2]. If the results do not answer the question, say so briefly and answer from general knowledge.” In the app, each number in the answer becomes a tappable link to its source, and a row of source cards above the answer shows each site’s letter mark, domain, citation number and page title. The letter marks are drawn on the phone, so no request for a site’s icon leaves it.

A citation shows where a claim is supposed to come from; it does not prove that the source supports it. A model can attach the right number to a claim the passage never makes, or cite one source for a claim drawn from another. A 2023 study of four generative search engines found that only 51.5% of their generated sentences were fully supported by their citations, and that 74.5% of citations supported the sentence they were attached to 7. Small on-device models are not exempt. A citation’s value is that it makes such errors checkable: open the source and find the sentence.

05

Where RAG fails

RAG fails in three places: retrieval can miss the passage that answers the question, the passages it finds can be wrong or hostile, and the model can misuse what it is given. Each failure has partial remedies and none has a complete one. Table 1 lists the common failures.

Table 1Common failure modes of retrieval-augmented generation and the usual mitigations. General properties; individual systems vary.
FailureWhat goes wrongCommon mitigations
Missed retrievalThe passage that answers is never retrieved, often because it uses different words from the question.Rewriting the query; hybrid sparse and dense retrieval; retrieving more candidates.
Off-topic resultsThe top results share words with the question but not its meaning.Relevance checks before ranking; re-ranking.
Wrong or stale sourcesA passage is retrieved correctly but is itself false or out of date, and the answer repeats it.Showing sources and dates so readers can judge them.
Injected instructionsA retrieved page contains text written to instruct the model that reads it.Treating passages as data, not commands; limiting what the model can do.
Overflow and positionThe passages exceed the context budget, or the model overlooks those in the middle of a long context.Fixed budgets; fewer, better passages.
Unsupported answersThe answer makes claims the passages do not, or attributes them to the wrong source.Citation instructions; checking claims against sources.

Indirect prompt injection is the failure most specific to RAG. Because retrieved text enters the prompt, anyone who can publish a page that gets retrieved can try to steer the model that reads it; Greshake and colleagues demonstrated such attacks against real applications built on large language models in 2023 8. Prompt injection, explained covers the defences. Position matters too: Liu and colleagues found that language models used relevant information best when it appeared at the beginning or end of a long context, and worst when it sat in the middle 9, one reason to give a model fewer, better passages.

Brello 1.0 addresses several of these. It rewrites the question into a search query and tries up to eight search providers in order, moving on from any provider unless at least a third of its top results mention the question’s key terms. It keeps at most two passages per source within a fixed budget, and it tells the model to say so when the results do not answer the question. When search fails or comes back empty, which can take up to about 14 seconds for each provider attempt, Brello answers from the model’s own knowledge. None of these checks can establish that a source is true.

06

RAG on a phone: a worked example

In Brello 1.0, a web answer is a complete RAG pipeline: the search engine and the websites supply the text, and everything else, from writing the query to writing the answer, happens on the phone. Here is what happens, in order, when someone asks “Why is the sky blue? Keep it short.” with web search on.

  1. Decide whether to search. The question is not small talk, arithmetic or a writing task, so it merits a lookup. With web search off, a time-sensitive question would instead bring up a card titled ‘Search the web for this?’, with the buttons ‘Search the web’ and ‘Answer offline’.
  2. Write the query. Instructions meant for the answer are removed, so the search text is ‘Why is the sky blue’. For a short or referential follow-up, such as ‘how old is he?’, the first ten words of the previous question are added. The search engine receives this text and sees the phone’s IP address, as it would for any web request.
  3. Retrieve. The first provider in the chain that returns relevant results supplies them, and the top six become sources [1] to [6]. Up to four pages are fetched in parallel, each with a 7-second timeout and a 1.5 MB cap, and reduced to their main text. Meanwhile the status line reads ‘Searching the web’, then ‘Reading 4 sources’.
  4. Rank. Each page is cut into passages of about 420–700 characters, and BM25 scores every passage against the query. At most two passages per source are kept: up to 3,400 characters for Brello Pro or Brello Vision, or 2,600 for Brello Core.
  5. Augment. The kept passages become a numbered block labelled ‘Web results (retrieved {date})’. Because web results are present, the system prompt also gains the current date and the instruction to cite.
  6. Generate. The on-device model starts a fresh session, so web context from earlier replies cannot crowd its 4,096-token window, and the status line reads ‘Thinking’. The answer streams in with inline citations.
  7. Check. Source cards sit above the answer and every citation is a link. Tapping one opens the page in the phone’s browser, which requests it from the site like any other visit.

The answer is only as good as its sources: if a passage is wrong, a faithful answer repeats the error, with a citation that lets the reader find it. The research note ‘Answering from the open web, without a server’ describes the system in full, and ‘Asking before going online: consent for web search’ explains when Brello asks first.

07

Privacy: who sees the query?

In any RAG system, whoever runs the retriever sees the query, and whoever runs the model sees the question, the retrieved passages and the answer. Running the model on the device changes the second part, not the first.

When retrieval and generation both run on a provider’s servers, the full question and the answer pass through them, and what is kept, and for how long, is set by that provider’s policy. When the model runs on the phone, the question and the answer can stay there. A search engine is still needed to find pages, so the search text has to leave, and the sites that serve the pages see ordinary requests for them.

Brello 1.0 works this way. Questions, photos and answers are processed on the phone. Web search is off by default, and Brello asks before going online. With search on, the phone sends the search text directly to the search engine that answers, which may be DuckDuckGo, Bing, Brave Search, Google News (RSS) or Wikipedia, and requests up to four result pages from their sites. Those requests carry ‘Do Not Track’ and Global Privacy Control headers, and they show what any web request shows, including the phone’s IP address. There is no Brello server, proxy or relay, so Brello keeps no logs of them. On Android, the hidden browser used for searching is torn down after 3 idle minutes, with its cookies, cache and storage wiped. Search results are held in memory for 15 minutes only. Photo questions never trigger web search.

How Brello handles your questions, photos and answers sets out what stays on the phone and what leaves it. For a sourced comparison with a cloud answer engine, see ‘Brello vs Perplexity: two ways to answer from the web’.

References

Reviewed .

  1. Lewis, P. et al. (2020). “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” Advances in Neural Information Processing Systems 33 (NeurIPS 2020). arxiv.org/abs/2005.11401
  2. Guu, K. et al. (2020). “REALM: Retrieval-Augmented Language Model Pre-Training.” Proceedings of the 37th International Conference on Machine Learning (ICML 2020). arxiv.org/abs/2002.08909
  3. Shuster, K. et al. (2021). “Retrieval Augmentation Reduces Hallucination in Conversation.” Findings of the Association for Computational Linguistics: EMNLP 2021. arxiv.org/abs/2104.07567
  4. Karpukhin, V. et al. (2020). “Dense Passage Retrieval for Open-Domain Question Answering.” Proceedings of EMNLP 2020. arxiv.org/abs/2004.04906
  5. Thakur, N. et al. (2021). “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models.” NeurIPS 2021 Datasets and Benchmarks Track. arxiv.org/abs/2104.08663
  6. Robertson, S. and Zaragoza, H. (2009). “The Probabilistic Relevance Framework: BM25 and Beyond.” Foundations and Trends in Information Retrieval, 3(4), 333–389. doi.org/10.1561/1500000019
  7. Liu, N. F., Zhang, T. and Liang, P. (2023). “Evaluating Verifiability in Generative Search Engines.” Findings of the Association for Computational Linguistics: EMNLP 2023. arxiv.org/abs/2304.09848
  8. Greshake, K. et al. (2023). “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec 2023). arxiv.org/abs/2302.12173
  9. Liu, N. F. et al. (2024). “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics, 12. arxiv.org/abs/2307.03172

Version history

  1. 1.0First published.