01
What is a small language model?
A small language model (SLM) is a language model compact enough to run on everyday hardware, such as a phone or laptop, instead of a data centre. There is no agreed cut-off: models with a few billion parameters or fewer are usually called small, and one 2024 survey uses a range of 100 million to 5 billion.
Small and large language models work the same way. Both are neural networks, almost always transformers, trained to predict the next token of text, and both are adapted after pre-training to follow instructions; the survey above covers decoder-only transformers of exactly this kind 1. The difference is scale. A phone-sized model has a few billion parameters, while frontier cloud models are far larger, so a small model holds less knowledge and has less capacity for reasoning (section 05). Small models are what make on-device AI practical.
Size is not the only lever on quality. Kaplan and colleagues found that a language model’s loss falls as a power law as its parameters, its training data and its training compute grow 2, and Hoffmann and colleagues showed that many large models had been trained on too little data: for a fixed compute budget, model size and training tokens should grow together 3. A small model’s quality therefore depends on its training as much as its size. It can learn from a larger model through distillation, in which it is trained to match the larger model’s outputs 4, and it gains from carefully chosen data: phi-1, a 1.3-billion-parameter code model trained largely on “textbook quality” data, reached 50.6% pass@1 on the HumanEval coding benchmark 5. Qwen’s technical report describes using its flagship models’ knowledge to “significantly reduce the computational resources required to build smaller-scale models” 6.
02
What do parameters and ‘effective parameters’ mean?
Parameters are the numbers a model learns during training, and their count is the usual measure of its size: the “1.7B” in Qwen3 1.7B means about 1.7 billion. “Effective parameters”, the “E” in Google’s Gemma 4 E2B and E4B, counts the parameters the model computes with and leaves out large embedding tables that it only looks up.
Most parameters sit in the weight matrices of a transformer’s attention and feed-forward layers, which every token passes through. A separate block, the embeddings, maps each token in the vocabulary to a vector, and in a small model it can be a large share of the total. Qwen3 1.7B’s embedding table has 151,936 rows (its configured vocabulary, padded beyond the tokeniser’s 151,669 entries) of 2,048 dimensions, about 311 million embedding parameters, consistent with its model card, which lists 1.7 billion parameters in all and 1.4 billion outside the embeddings 7. Scaling studies often count only the non-embedding parameters; Kaplan and colleagues found that performance follows a cleaner trend when embeddings are excluded 2.
Gemma 4’s E-series takes the idea further. Google’s model card explains that ‘The “E” in E2B and E4B stands for “effective” parameters’: the models use Per-Layer Embeddings (PLE), which give “each decoder layer its own small embedding for every token”, and “These embedding tables are large but are only used for quick lookups, which is why the effective parameter count is much smaller than the total” 8. The card lists Gemma 4 E2B at 2.3 billion effective parameters, 5.1 billion with embeddings, and E4B at 4.5 billion effective, 8 billion with embeddings.
The distinction matters on a phone because lookups need memory but little computation, and a runtime can leave the tables in storage. For text-only use of Gemma 4 E2B, the LiteRT-LM model page states that “the weight footprint in memory can be as low as 0.8 GB while the runtime uses memory mapping to support the 1.12GB of embedding parameters” 9. “Effective” is different again from “active”: Gemma 4 26B A4B is a mixture-of-experts model with 25.2 billion parameters, of which 3.8 billion are active during inference 8.
03
Why do small models fit on phones?
A model’s weights take roughly its parameter count times the bits stored per parameter, divided by eight. At 4 bits, 1.7 billion parameters need about 0.85 GB, which leaves room for Android and other apps on a phone with 4 GB of RAM, the amount Brello recommends for Brello Core.
Parameter count and precision are the two levers. Figure 1 works through the arithmetic for Qwen3 1.7B, the model behind Brello Core, from 32-bit floating point down to the file Brello Core actually downloads.
- Stored as 32-bit floats, 1.7 billion parameters take 6.8 GB. That is more than all the memory in a phone with 4 GB of RAM, the amount Brello recommends for Brello Core.
- At 16 bits, the precision Qwen3 1.7B is published in, they take 3.4 GB. That would leave little room for Android and other apps on a 4 GB phone.
- At 8 bits, they take 1.7 GB. Each halving of the bits per weight halves the memory the weights need.
- At 4 bits, about 0.85 GB. Half a byte per weight is what makes a model of this size practical on a 4 GB phone.
- Brello Core’s actual file is 977 MB. It stores the weights as 4-bit integers in blocks of 32, each block with its own scale, which averages about 4.6 bits per parameter.
Show data
| Precision | Bits per parameter | Weight memory |
|---|---|---|
| 32-bit float | 32 | 6.8 GB |
| 16-bit float | 16 | 3.4 GB |
| 8-bit integer | 8 | 1.7 GB |
| 4-bit integer | 4 | 0.85 GB |
| Brello Core file (INT4, blocks of 32) | ≈4.6 on average | 977 MB |
At 16 bits, the precision Qwen3 1.7B is published in 7, the same weights take about 3.4 GB, most of a 4 GB phone’s memory before the operating system takes its share. The file Brello Core runs is “a dynamic INT4 variant (block-32 weights, FP32 activations)” 10, which is why it comes to 977 MB. Phone-sized models combine a few billion parameters with 4-bit or mixed-precision weights for this reason; our explainer on quantization covers the methods.
Weights are not the whole footprint. A running model also needs memory for its key–value cache, which grows with the length of the context window, and for the runtime itself; our guide to running a language model on an Android phone works through each requirement. Speed follows the same arithmetic: each new token requires reading the model’s weights from memory, so fewer and smaller weights mean faster answers on the same phone.
04
What do small models do well?
Small models are good at tasks where what they need is in the prompt or in common knowledge: rewriting, summarising, drafting, explaining and answering from supplied passages. Because they run on the device, they also answer without a network connection and without sending the prompt to a server.
Their developers describe broad abilities. Google’s model card calls the Gemma 4 family “well-suited for reasoning, agentic workflows, coding, and multimodal understanding”, with E2B and E4B accepting text, image and audio input 8, and Qwen’s technical report says Qwen3 supports 119 languages and dialects 6. These are the developers’ own descriptions and benchmarks; how well a particular task works on a particular phone still needs testing.
Supplying the right text helps with what a small model doesn’t know. Retrieval-augmented generation adds passages fetched at question time, so the model works from sources rather than memory; Brello 1.0 uses it for web answers, with numbered citations. Small models are also predictable to run: there is no per-question server cost, no queue and no dependence on a connection after the download.
05
Where do small models fall short?
Small models hold less knowledge, are weaker at long or multi-step reasoning and handle long contexts less well than larger models, and like all language models they can state false things fluently.
The pattern shows within a single family. On Google’s model card, Gemma 4 E2B scores 60.0% on the MMLU Pro knowledge and reasoning benchmark and E4B 69.4%, against 85.2% for the 31-billion-parameter Gemma 4 model; on GPQA Diamond, a set of graduate-level science questions, the scores are 43.4%, 58.6% and 84.3% 8. These are Google’s results for its models, not measurements of Brello.
Small models can also get stuck repeating themselves. Qwen’s model card warns that greedy decoding “can lead to performance degradation and endless repetitions” 7. Brello 1.0 guards against loops: every 48 characters it checks whether the end of the answer repeats a block 3 or more times over at least 120 characters, and if it does, cuts the answer after the first copy and stops the model.
In Brello 1.0, the model also sees only part of a long chat: each answer uses the last six messages, clipped, within a 4,096-token context, so long chats lose their early detail. The system prompt asks the model to say when it is unsure; we have not measured how often it does, and the instruction is no guarantee. Our explainer on AI hallucinations covers why models state false things with confidence.
06
Why do small models copy the shape of their instructions?
Small models tend to imitate the form of a prompt as well as follow its content, so a long, structured system prompt can come back as structure in every answer. We found this while building Brello 1.0, and kept its system prompt to four plain sentences as a result.
A prompt that told the model to “lead with the direct answer, then useful detail” came back as answers with bold labels reading “Direct Answer:” and “Useful Detail:”. The instruction was meant as advice about content; the model treated it as a template. Brello 1.0’s system prompt now reads, in full:
Brello 1.0 system prompt, verbatim
You are Brello, a helpful assistant that runs privately on the user's phone. Reply to the user's message directly and naturally, like a knowledgeable friend. Keep it clear and to the point. If you are not sure about something, say so instead of guessing.
Everything else is added only when needed: the model’s name and base model when you ask about Brello itself, the date when a question is about time or uses web results, and an instruction to cite numbered sources when web results are present. Our research note on small models and the shape of instructions describes what we tried.
Settings matter as much as wording. Brello 1.0 uses each model developer’s recommended sampling settings: for the Gemma 4 models, temperature 1.0, top-k 64 and top-p 0.95 8; for Qwen3, temperature 0.7, top-k 20 and top-p 0.8, the values Qwen recommends for non-thinking mode 7.
07
How do thinking modes work in small models?
A thinking mode makes the model write out intermediate reasoning before its final answer, spending more time and tokens for more careful results on multi-step problems. Qwen3 and Gemma 4 both offer it as a switch within one model rather than as a separate reasoning model.
The idea comes from chain-of-thought prompting. Wei and colleagues showed in 2022 that reasoning abilities “emerge naturally in sufficiently large language models” when they are prompted with worked, step-by-step examples 11, which at the time meant very large models. Newer small models are trained to reason, in Qwen3’s case partly by learning from its larger models: Qwen’s report describes integrating “thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework” 6.
The switch is part of the prompt. Qwen3 1.7B thinks by default and writes its reasoning inside a <think>…</think> block; developers turn thinking off with an enable_thinking setting, and users can add /think or /no_think to a message to change mode from turn to turn 7. Gemma 4 enables thinking when its system prompt starts with a <|think|> token 8.
Thinking costs time and context, because reasoning tokens are part of the model’s output and share its window, which matters at Brello 1.0’s 4,096 tokens. Brello’s Think harder mode raises the answer limit from 1,200 to 2,048 tokens and samples at temperature 0.6 and top-p 0.95, in line with Qwen’s recommendation for thinking mode 7. The reasoning appears in a “Thought process” panel above the answer, and if a model reasons but never answers while Think harder is off, Brello shows the reasoning as the answer. Our explainer on reasoning models covers the technique in more depth.
08
How do Qwen3 1.7B, Gemma 4 E2B and Gemma 4 E4B compare?
The three differ in size, inputs and context length. Qwen3 1.7B has 1.7 billion parameters and takes text only; Google lists Gemma 4 E2B and E4B at 2.3 and 4.5 billion effective parameters, with text, image and audio input. All three are released under Apache 2.0, and Brello 1.0 runs one of each (Table 1).
| Qwen3 1.7B | Gemma 4 E2B | Gemma 4 E4B | |
|---|---|---|---|
| Developer | Alibaba (Qwen team) | Google DeepMind | Google DeepMind |
| Parameters | 1.7B, of which 1.4B outside the embeddings | 2.3B effective, 5.1B with embeddings | 4.5B effective, 8B with embeddings |
| Layers | 28 | 35 | 42 |
| Context length | 32,768 tokens | 128K tokens | 128K tokens |
| Input | Text | Text, image, audio | Text, image, audio |
| Thinking mode | Yes, on by default | Yes | Yes |
| Licence | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| In Brello 1.0 | Brello Core, 977 MB | Brello Vision, 2.59 GB | Brello Pro, 3.66 GB |
The models page gives each Brello model’s full specification, and the glossary defines the terms used here.
References
Reviewed . Model cards and model pages were checked on that date.
- Lu, Z. et al. (2024). “Small Language Models: Survey, Measurements, and Insights.” arXiv preprint. arxiv.org/
abs/ 2409.15790 - Kaplan, J. et al. (2020). “Scaling Laws for Neural Language Models.” arXiv preprint. arxiv.org/
abs/ 2001.08361 - Hoffmann, J. et al. (2022). “Training Compute-Optimal Large Language Models.” arXiv preprint. arxiv.org/
abs/ 2203.15556 - Hinton, G., Vinyals, O. and Dean, J. (2015). “Distilling the Knowledge in a Neural Network.” NIPS 2014 Deep Learning Workshop. arxiv.org/
abs/ 1503.02531 - Gunasekar, S. et al. (2023). “Textbooks Are All You Need.” arXiv preprint. arxiv.org/
abs/ 2306.11644 - Qwen Team (2025). “Qwen3 Technical Report.” arXiv preprint. arxiv.org/
abs/ 2505.09388 - Qwen. “Qwen3-1.7B.” Model card and configuration, Hugging Face. Accessed 5 October 2026. huggingface.co/
Qwen/ Qwen3-1.7B - Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Accessed 5 October 2026. ai.google.dev/
gemma/ docs/ core/ model_card_4 - litert-community. “gemma-4-E2B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/
litert-community/ gemma-4-E2B-it-litert-lm - litert-community. “Qwen3-1.7B.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/
litert-community/ Qwen3-1.7B - Wei, J. et al. (2022). “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/
abs/ 2201.11903
Version history
- 1.0First published.



