Foundations

Model cards and system cards, explained

What model cards and system cards contain, how to read one critically, and two worked examples: Google’s Gemma 4 model card and the Brello 1.0 system card.

Brello Research9 min readVersion 1.0

Summary

A model card is a short document published with a machine-learning model that states what it is for, how it was evaluated and where it falls short. Mitchell et al. proposed the format in 2019 with nine sections; on Hugging Face, a model card is the README.md file in a model’s repository, with a YAML metadata header. A system card documents a deployed system: the models plus the instructions, tools, safeguards and data flows around them. Reading either critically means checking whether results describe the model as you will run it, whether they are broken down by group and condition, and what is missing. The Brello 1.0 system card reports no benchmark results, because none are published.

  • Mitchell et al. (2019) proposed model cards with nine sections, from model details and intended use to quantitative analyses, ethical considerations, and caveats and recommendations.
  • On Hugging Face, a model card is the README.md file in a model repository, with a YAML header for metadata such as the licence, base model, datasets and evaluation results.
  • A 2024 analysis of 32,111 model cards on Hugging Face found that the sections on environmental impact, limitations and evaluation were the least often filled out.
  • A system card documents a deployed system, including its instructions, safeguards and data flows, because a model’s published scores don’t describe the model as a product runs it.
  • The Brello 1.0 system card covers three on-device models with a 4,096-token context window and states that no benchmark or evaluation results are published for Brello 1.0.
Contents7 sections

01

What is a model card?

A model card is a short document published with a machine-learning model that says what the model is for, how it was built and evaluated, and where it falls short. Its job is to let people decide whether a model suits their purpose before they rely on it.

The format was proposed by Mitchell et al. in “Model Cards for Model Reporting”, presented at the FAT* conference on fairness, accountability and transparency in January 2019 1. The paper describes model cards as short documents accompanying trained models that provide “benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups”, and that disclose the context in which a model is meant to be used. The authors gave two examples: a model that detects smiling faces in images and one that detects toxic comments in text.

Model cards have a companion for data. “Datasheets for Datasets” proposes that every dataset come with a record of why it was created, how it was collected and what it should and shouldn’t be used for 2. Both rest on the same observation: a model’s behaviour depends on choices its users can’t see, so those choices have to be written down.

On Hugging Face, which hosts many open models, a model card is the README.md file in a model’s repository. Its documentation puts it plainly: “model cards are simple Markdown files with additional metadata” 3.

02

What should a model card contain?

A model card should say what the model is, how well it works and for whom, and what to watch for. Mitchell et al. set out nine sections that between them answer those three questions, as Figure 1 shows.

The anatomy of a model card: nine sections that answer three questions A model card with the nine sections proposed by Mitchell et al. in 2019. Model details and intended use answer what the model is and what it is for. Factors, metrics, evaluation data, training data and quantitative analyses answer how well it works and for whom. Ethical considerations, and caveats and recommendations, answer what to watch for. An illustrative chart beside quantitative analyses shows an overall error rate of 6% hiding a rate of 13% for one group. Model card Nine sections · Mitchell et al. 2019 Three questions Our grouping of the paper’s nine sections 01Model details 02Intended use What is it, and what is it for? Who made it, which version and when,its type, licence and how to cite it. What it is for and who it is for, andthe uses it was not designed for. 03Factors 04Metrics 05Evaluation data 06Training data 07Quantitative analyses How well does it work, and for whom? Groups and conditions that may changeresults, such as age group or lighting. Which measures were used and why, withthresholds and their uncertainty. The test datasets, why they werechosen and how they were prepared. The same detail for the trainingdata, where it can be shared. Results for each group and condition,not only one overall average. Error · illustrative All Group A Group B 6% 4% 13% 08Ethical considerations 09Caveats and recommendations What should you watch for? Sensitive data, risks to people, andthe mitigations that were applied. What was not tested, and advice foranyone deploying the model. Model card Mitchell et al. 2019 What is it, and what is it for? 01Model details 02Intended use How well does it work, and for whom? 03Factors 04Metrics 05Evaluation data 06Training data 07Quantitative analyses Error · illustrative All Group A Group B 6% 4% 13% What should you watch for? 08Ethical considerations 09Caveats and recommendations
  1. A card opens by saying what the model is and what it is for. Model details give the developer, version, date, type and licence. Intended use names the uses the model was built for, and the uses it was not.
  2. Next it fixes how performance will be judged. Factors are the groups and conditions that could change the results; metrics are the measures used, with their thresholds and uncertainty.
  3. Then it names the data. Which datasets the model was tested on, why they were chosen and how they were prepared, and the same for its training data, where that can be shared.
  4. Results are reported for each group, not only on average. In these illustrative figures, an overall error rate of 6% hides a rate of 13% for one group.
  5. It closes with what to watch for. Ethical considerations and caveats record risks, mitigations and what was not tested, so a reader can judge whether the model fits their use.
Figure 1The nine sections of a model card proposed by Mitchell et al. (2019), grouped by the question each answers. The grouping is ours, and the error rates in step 4 are illustrative.

The central idea is disaggregation. A single accuracy figure averages over everyone a model was tested on, and can hide a group for whom it works badly. Model cards ask for results broken down by the factors relevant to the intended use, such as age group, camera type or lighting, and by combinations of them 1. The paper accepts that training data may not be shareable, and asks at least for how it is distributed across the same factors.

Cards on Hugging Face add a machine-readable header. A YAML block at the top of README.md records fields such as the licence, languages, datasets and base model, and can carry evaluation results in a structured model-index that the Hub displays on the model page 3. For a fine-tuned, adapted, merged or quantised derivative, base_model names the original, so a derived model can be traced to the card that describes its parent.

---
license: apache-2.0
base_model: org/base-model       # the model this one derives from
base_model_relation: quantized   # or finetune, adapter, merge
datasets:
  - org/training-data
model-index:                     # structured evaluation results
  - name: org/this-model
    results: [...]
---

Naming the sections is the easy part; filling them in is not. In 2024, Liang et al. analysed 32,111 model cards on Hugging Face. Most models with substantial downloads had a card, but the sections on environmental impact, limitations and evaluation were the least often filled out, while the training section was the most consistently completed 4. The parts a reader most needs in order to judge a model are the parts most often missing.

03

What is a system card?

A system card documents a deployed AI system rather than a single model: the models, plus the instructions, tools, retrieval, safeguards and data flows around them, and how the whole was tested. It describes the product people actually use.

The difference matters because the same model behaves differently in different systems. A system prompt changes what it says. Retrieval changes what it knows. A context limit changes how much of a conversation it sees, and safeguards stop some outputs before anyone reads them. None of this appears on a model card. OpenAI’s system card for GPT-4, published in March 2023, is one example: it describes safety challenges found in testing and a deployment process spanning “measurements, model-level changes, product- and system-level interventions (such as monitoring and policies), and external expert engagement” 5.

A useful system card answers a fixed set of questions: which models and versions run, and with what settings; what instructions they are given; which tools and sources they can use; what is stored or sent, where and for how long; which safeguards sit around the models; what the system is known to get wrong; and how it was tested. Each answer should be specific enough for a reader to check. Table 1 compares the two kinds of document.

Table 1Model cards and system cards compared.
AspectModel cardSystem card
DescribesOne trained modelA deployed product built on one or more models
Written byUsually the model’s developerThe organisation that runs the product
CoversIntended use, data, results by group, limitationsModels and settings, instructions, tools, safeguards, data flows, limitations, testing
Results describeThe model as its developer evaluated itThe system as people use it
ExampleGoogle’s Gemma 4 model cardThe Brello 1.0 system card

04

How do you read a model card critically?

Read a model card as evidence, not as a specification. Check what was measured, on what data and for which version of the model, and notice what the card leaves out. Five questions cover most of it.

  1. Is this the model you will run? Published results usually describe the developer’s own release. A quantised build, a shorter context window, a different runtime or a different system prompt makes a different system, and the card’s numbers don’t carry over automatically.
  2. Are results broken down? Look for results by group and condition, with sample sizes and some measure of uncertainty, rather than a single average.
  3. Was the test data kept separate? If benchmark questions leaked into the training data, a high score may reflect memory rather than skill. A careful card says how it guarded against this.
  4. Are limitations and out-of-scope uses stated? They are among the sections most often left empty 4, and they are the ones that tell you when not to use the model.
  5. Who wrote it, when, and on what terms? A card is a dated snapshot, usually written by the model’s developer. Check the version, the date and the licence, which governs what you may do with the model, and read independent evaluations alongside the card where they exist; ‘AI safety evaluations, explained’ describes what those involve.

05

What does Google’s Gemma 4 model card contain?

Google DeepMind’s Gemma 4 model card is a current example for an open-weight model family, and two Brello models are based on models it describes. It was last updated on 30 July 2026, and we checked it on 5 October 2026 6.

Its structure extends the original nine sections for a modern model. It opens with a models overview and a specification table for each size, then gives benchmark results, core capabilities and best practices for prompting and sampling. Model data covers the training dataset and its preprocessing. Ethics and safety covers the evaluation approach and its results. Usage and limitations covers intended usage, limitations, ethical considerations and risks, and benefits. The card names Google DeepMind as its author and Apache 2.0 as the licence.

It covers several sizes, from E2B and E4B, two small language models it describes as targeting mobile and edge devices, to larger models for consumer GPUs and workstations. For E4B it lists 4.5 billion effective parameters, 8 billion with embeddings; for E2B, 2.3 billion, or 5.1 billion with embeddings; both have a context length of 128K tokens. Three sentences show the habits of a careful card. On benchmarks: “Evaluation results marked in the table are for instruction-tuned models.” On safety testing: “All testing was conducted without safety filters to evaluate the model capabilities and behaviors.” And on factual accuracy, models “are not knowledge bases. They may generate incorrect or outdated factual statements.”

Read with the questions above, the card also shows what a model card can’t tell you. Brello Pro and Brello Vision run LiteRT-LM builds of Gemma 4 E4B and E2B, downloaded from Hugging Face, with a 4,096-token context window, answers of up to 1,200 tokens (2,048 with Think harder) and Brello’s own instructions and safeguards. The card’s benchmark scores and its 128K-token context describe Google’s instruction-tuned models, not those builds as Brello runs them, so Brello doesn’t present them as its own.

06

What does the Brello 1.0 system card cover?

The Brello 1.0 system card documents the app as a whole rather than a single model, because what people use is a system: three on-device models, their limits, the instructions and safeguards around them, and every route by which data leaves the phone.

Each model also has its own page in the model-card pattern. Brello Pro’s page, for example, gives its base model, licence, download size, recommended phone memory and runtime, and links to Google’s model card for Gemma 4. Brello Super Intelligence is in development and has no published specifications, so it has no model card or system card. The Brello Charter commits us to evaluate each new capability before release and to publish an evaluation summary with it (commitment 04).

References

Reviewed . Web pages were checked on that date.

  1. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D. and Gebru, T. (2019). “Model Cards for Model Reporting.” Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). arxiv.org/abs/1810.03993
  2. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H. and Crawford, K. (2021). “Datasheets for Datasets.” Communications of the ACM 64(12), 86–92. arxiv.org/abs/1803.09010
  3. Hugging Face. “Model Cards.” Hugging Face Hub documentation. Accessed 5 October 2026. huggingface.co/docs/hub/en/model-cards
  4. Liang, W., Rajani, N., Yang, X., Ozoani, E., Wu, E., Chen, Y., Smith, D. S. and Zou, J. (2024). “What’s documented in AI? Systematic Analysis of 32K AI Model Cards.” arXiv preprint. arxiv.org/abs/2402.05160
  5. OpenAI (2023). “GPT-4 System Card.” March 2023. Accessed 5 October 2026. cdn.openai.com/papers/gpt-4-system-card.pdf
  6. Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Last updated 30 July 2026. Accessed 5 October 2026. ai.google.dev/gemma/docs/core/model_card_4

Version history

  1. 1.0First published.