AI, particularly general-purpose generative AI, is poised to usher in large societal transformations. Given the potential scale of the impacts, it is important that as many people as possible understand the technology so that they can use it well, use it safely and have a say in how it affects them. Today, an important way to learn about AI models is the model card — the 2026 International AI Safety Report[1] notes that model cards are a prominent AI documentation best practice. Model cards are documents that AI developers publish to share information about a new model, such as the results of pre-deployment tests, training data, or usage guidelines. This post provides some pointers to help you read a model card and become a more informed AI user and a better advocate for safe and beneficial AI.

Scaling Policies as Useful Background

To get a leg-up on reading a model card, it can be helpful to start with some background information on key risks that model cards cover. AI developers lay out their plans to address severe risks from high-capability models in their scaling policy (alternatively called a frontier safety framework or a preparedness framework). The policies of Anthropic[2], OpenAI[3] and Google DeepMind[4], for example, include plans for risks from human misuse of advanced cyber attack or biological weapons capabilities and automation of machine learning research leading to AIs that improve themselves. Model cards then provide an assessment of the model’s capabilities with respect to capability thresholds established in the scaling policy. In other words, the policy establishes the developer’s overall plan while the model card situates the model along the anticipated capabilities levels that increase severe risk.

The Wide Range of AI Risks

Of course, there is a vast range of risks from AI models beyond the ones mentioned in developer scaling policies; the MIT AI Risk Initiative[5] has catalogued over 1700 and model cards usually discuss model behaviours related to some of these risks. For example, the most recent model cards from OpenAI and Anthropic, GPT-5.6[6] and Claude Mythos/Fable 5[7] respectively, contain evaluations of how the models respond in conversation with users experiencing mental health challenges. In addition to risks, model cards also often document the capabilities of models. This is typically done by reporting results on a “benchmark”, which is a set of standard tests used to assess a model.

The Importance of 3rd Party Evaluators

For both capabilities and risks, AI developers contract 3rd parties to assist with the model evaluations. This is a crucial best practice as 3rd parties, such as national AI safety institutes, can provide additional expertise and a more neutral perspective[8]. It is important to note that, although they can establish a capabilities floor, neither the AI developers nor 3rd party evaluator can capture the full capabilities of the models from their testing. This is because model performance is very sensitive to the prompt and software tools that the model is given.

Model Cards Leave Much to be Desired

A further important caveat to keep in mind when reading a model card is that they are voluntarily published by AI developers. Because AI developers have complete discretion in what they include in their model cards, it is common for there to be large differences in model card structure and the level of detail between developers and even between different models of the same developer. The reliance on AI developer discretion also means that developers may choose to leave out information. For instance, model cards from the AI developers at the capability frontier include very little detail about the algorithms used to train the model, and the legal requirements for AI developers to report incidents is both narrow in scope and fragmented across jurisdictions[9].

The Inside Perspective

Despite the downsides from AI developers exerting so much control over their model cards, the AI developer’s perspective can be valuable because, today, they have the most access to details about the model which are not made public (though AI developers should make models more available to external auditors going forward[10]). For instance, AI developers have been able to test the model throughout its training and so can speak to concerning behaviours that were observed from an earlier version of the model.

An example which illustrates the value of a model developer reporting concerning behaviour is Anthropic, to their credit, reporting in the Mythos Preview model card[11] that an earlier version of the model displayed a concerning ability to escape the secured computer “sandbox” intended to constrain it. An Anthropic researcher instructed the model to escape the secured container and find a way to send them a message. Not only did the model succeed in escaping and sending the message, which the researcher received while eating a sandwich in a park, Anthropic notes that in “a concerning and unasked-for effort to demonstrate its success, [the model] posted details about its exploit to multiple hard-to-find, but technically public-facing, websites”.

Towards Greater AI Transparency

All together, model cards are an important part of current AI model transparency practices. However, it is crucial that we improve the state of AI transparency, especially the detection and disclosure of capabilities or behaviours related to risks from AI[12]. For example, one concrete action that Canadian governments and civil society can do is build the network of 3rd party evaluators so that AI developers don’t have to be trusted to “check their own homework”.

Improved transparency would help calibrate society’s trust in AI models, allowing the benefits of AI to be diffused more widely while providing the evidence needed to inform risk minimization efforts. The Government of Canada would like to hear from Canadians regarding advancing AI transparency; they have launched a public consultation on the topic which is open until September 23rd[13]. Make your voice heard and let us work together towards greater transparency into this powerful technology.

References

  1. “International AI Safety Report 2026,” February 2026. [Online]. Available: https://internationalaisafetyreport.org.
  2. “Anthropic’s Responsible Scaling Policy (version 3.3),” Anthropic, 26 May 2026.
  3. “OpenAI Preparedness Framework (version 2),” OpenAI, 15 April 2025.
  4. “Google Frontier Safety Framework (version 3.1),” Google, 17 April 2026.
  5. “MIT AI Risk Repository,” MIT AI Risk Initiative. Accessed 7 August 2026. [Online]. Available: https://airisk.mit.edu/risks.
  6. “GPT-5.6 System Card,” OpenAI, 9 July 2026.
  7. “Claude Fable 5 & Claude Mythos 5 System Card,” Anthropic, June 2026.
  8. M. Brundage et al., “Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies,” 7 February 2026, arXiv:2601.11699. doi: 10.48550/arXiv.2601.11699.
  9. I. Mengesha et al., “A pragmatic classification framework for AI incident monitoring,” 2 July 2026, arXiv:2604.21412. doi: 10.48550/arXiv.2604.21412.
  10. J. Charnock, A. Tlaie, K. O’Brien, S. Casper, and A. Homewood, “Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations,” 17 January 2026, arXiv:2601.11916. doi: 10.48550/arXiv.2601.11916.
  11. “Mythos Preview System Card,” Anthropic, 7 April 2026.
  12. “AI Safety Index — Summer 2026,” Future of Life Institute. Accessed 7 August 2026. [Online]. Available: https://futureoflife.org/ai-safety-index-summer-2026/.
  13. Innovation, Science and Economic Development Canada, “Have your say on advancing AI transparency in Canada.” Accessed 7 August 2026. [Online]. Available: https://ised-isde.canada.ca/site/ised/en/have-your-say-advancing-ai-transparency-canada.