A system card is documentation for a deployed AI system, designed to provide insight into that system’s underlying architecture. It helps better explain how the AI operates, and it outlines the AI models that comprise an AI system. The sense covered here is AI documentation, scoped to a whole running system rather than to one trained model.
Two neighbors share the phrase. One is the paper filing card in a card index, which is where the bare term resolves in most general reference works. The other is the model card, which documents a single trained model rather than the system assembled around it. Meta published a prototype AI System Card tool as its next step in exploring increased explainability through model and system documentation.
Key Takeaways
- A system card describes a deployed AI system, giving insight into that system’s underlying architecture and outlining the AI models that comprise it.
- Meta’s stated reason for the format is that machine learning models are typically part of a larger AI system, so model cards do not paint a comprehensive picture of what an AI system does.
- The Claude Opus 4.5 system card covers tests of model safeguards, honesty and agentic safety, a comprehensive alignment assessment, a model welfare report, and evaluations mandated by Anthropic’s Responsible Scaling Policy.
- OpenAI documents GPT-5 as a unified system with a smart and fast model, a deeper reasoning model, and a real-time router that decides which model to use.
- An AI system card also carries security information: The intent and scope of the system’s security and safety posture, plus a link to the security and safety issues that have been fixed and when they occurred.
How Does a System Card Work?
Three moves produce one. The developer fixes the boundary of the system, the tests run against everything inside that boundary, and the card ships with the release.
1. The Developer Draws the Boundary Around the System
Many machine learning models are typically part of a larger AI system. That system is a group of ML models, AI and non-AI technologies that work together to achieve specific tasks. Drawing that outline comes before anything else. Everything the document says afterwards depends on where the edge sits.
An aircraft flight manual works the same way. The engine spec sheet states thrust, and the manual states the conditions the whole aircraft was cleared to fly in.
2. Testing Covers the System as Deployed, Safeguards Included
Anthropic’s system card for Claude Opus 4.5 provides a detailed assessment of the model’s capabilities. It then describes a wide range of safety evaluations: tests of model safeguards, honesty, and agentic safety, plus a comprehensive alignment assessment. That assessment includes investigations of sycophancy, sabotage capability, and evaluation awareness. The card also carries a model welfare report and a set of evaluations mandated by Anthropic’s Responsible Scaling Policy.
A building’s fire-safety certificate follows the same logic. It covers the exits, alarms, and sprinklers around the occupants, not the tensile strength of the steel.
3. The Card Ships With the Release and Becomes the Public Record
OpenAI keeps a standing deployment-safety page sharing the technical work it does to make its systems safe. The page covers how deployed models perform in evaluations, the risks it measures, and the steps it takes to improve over time.
The scope test: The unit of AI documentation moved from the model to the deployed system, and the safeguards live in that gap. A card earns the name only when it describes the wrapper: the routing, the mitigations, the deployment context and the release decision.
| Element | What it records | Where the example comes from |
|---|---|---|
| Capabilities assessment | A detailed assessment of the model’s capabilities | Anthropic |
| Safeguard tests | Tests of model safeguards, honesty, and agentic safety | Anthropic |
| Alignment assessment | Investigations of sycophancy, sabotage capability and evaluation awareness | Anthropic |
| Model welfare report | A welfare report filed alongside the safety evaluations | Anthropic |
| Policy evaluations | Evaluations mandated by the Responsible Scaling Policy | Anthropic |
| System architecture | The AI models that comprise the AI system | Meta |
| Components and data | Architecture and components, the models used, and the data used to train and augment them | Red Hat |
| Security and safety posture | The intent and scope of the posture, plus issues fixed and when they occurred | Red Hat |
Source: Anthropic, Meta, Red Hat
What Problem Do System Cards Solve?
Machine learning models do not always work in isolation to produce outcomes. Models may also interact differently depending on what systems they are a part of. Because of that, model cards do not paint a comprehensive picture of what an AI system does. That sentence is the argument for the artifact, and it came from the organization that shipped the prototype tool.
Meta’s worked example is image classification. Its image classification models are all designed to predict what is in a given image. Yet they may be used differently in an integrity system that flags harmful content. The same models also sit in a recommender system used to show people posts they might be interested in.
The same weights behave differently depending on what wraps them. Documentation pinned to the weights alone cannot describe the behavior a user actually meets. Readers who want the underlying mechanics can start with how large language models work.
Red Hat frames the same gap in application terms. An AI system card contains information about how a particular AI system is built: its architecture and components. That includes the models used by the system and the data used to train and augment those models.
Who Publishes System Cards?
Meta published the prototype AI System Card tool. OpenAI maintains a deployment-safety index carrying dated system cards for recent releases. Anthropic published a system card describing its evaluations of Claude Opus 4.5. Red Hat introduced an AI system card for its Ask Red Hat conversational chatbot.
Frontier Labs and product vendors both use the format, and neither group answers to a shared template.
Why Does a System Card Matter?
OpenAI documents GPT-5 as a unified system: a smart and fast model that answers most questions, plus a deeper reasoning model for harder problems. A real-time router decides which model to use based on conversation type, complexity, tool needs, and explicit intent. A document scoped to one model cannot describe that object at all.
The router is continuously trained on real signals, including when users switch models, preference rates for responses, and measured correctness. Behaviour the user meets therefore depends on a component sitting outside any single set of weights.
Our AI benchmark coverage keeps showing capability rankings turn over every few release cycles, which is when a dated document beats a marketing page. Release dates and version history across the major labs sit in the AI model release tracker.
A card also states what the developer tested and the standard the release went out under. Informed by the testing described in its system card, Anthropic deployed Claude Opus 4.5 under the AI Safety Level 3 Standard. That records a decision the developer made. It does not make the system safe.
A system card is the only public artifact where the wrapper around the model gets written down. Its usefulness is set by what it says about the guardrails, the deployment context, and the release decision. What it claims about the weights matters less.
Pros, Cons, and Risks
Advantages
- A system card can help enable a better understanding of how these systems operate based on an individual’s history, preferences, settings, and more.
- One document carries the safeguard tests, the honesty and agentic-safety tests, and the comprehensive alignment assessment for a release.
- The card states the intent and scope of the system’s security and safety posture, with a link to the security and safety issues that have been fixed and when they occurred.
- A card can capture essential details about how the AI system has been built, including its core components and data sources.
Trade-offs and Risks
- The evidence is largely self-generated. Anthropic’s card describes Anthropic’s own evaluations of Claude Opus 4.5. Self-reported documentation carries that structural limit, which is a property of the artifact rather than an accusation against any lab.
- Red Hat introduces AI system cards by extending an analogy the company had drawn earlier. Nothing in the record names an agreed format, so two honest cards can still resist a direct comparison.
- The Ask Red Hat system card can be accessed by Red Hat subscribers. Publication and public readability are separate questions.
- A documented safeguard and a safeguard that holds under attack are different things. Measured bypass rates sit in our AI jailbreaking statistics.
What a system card does not do: The document records safeguards, testing and deployment context. It does not certify a system, remove risk, or prevent misuse. Red Hat’s own framing puts the reader in the position of someone reading a label before deciding to buy, subscribe, or even use the services of an AI system. The judgment stays with the reader.
Types of System Card
Four versions of the artifact circulate, answering to different audiences. Three of them describe a deployed system, while the remaining one uses the same phrase for a different object entirely.
An arXiv paper on system cards for AI-based decision-making in public policy proposes system cards as scorecards presenting the outcomes of formal audits. The benchmark behind them holds 56 criteria organized within a four-by-four matrix. Rows cover data, model, code, and system, and columns cover development, assessment, mitigation, and assurance.
| Type | What it documents | Who publishes it | Source document date |
|---|---|---|---|
| Product-system card | The AI models that comprise a deployed product system, such as Instagram feed ranking | Meta | February 2022 |
| Frontier-release card | Capabilities plus safeguard, honesty, agentic-safety and alignment evaluations for one release | Anthropic, OpenAI | August 2025, November 2025 |
| Enterprise application card | Architecture, components, training and augmentation data, and the security and safety posture | Red Hat | September 2025 |
| Audit scorecard | The outcome of a formal audit against a 56-criteria accountability benchmark | Public-policy research literature | March 2022 |
Sources: Meta, Anthropic, OpenAI, Red Hat, arXiv
Two neighboring documents get mixed in constantly. A model card documents a single trained model, and a datasheet documents a dataset. Both are scope distinctions rather than competing formats.
Real-World Applications
Three publishers show the artifact in its working forms:
- A consumer product surface, where the card covers a ranking system.
- A frontier release, where each model ships with its own dated card.
- A shipped enterprise service, where the card doubles as a security disclosure.
Meta: the Instagram Feed Ranking Pilot
Meta’s pilot System Card, which it developed and continued to test, was for Instagram feed ranking. That process takes as-yet-unseen posts from accounts that a person follows. It then ranks them based on how likely that person is to be interested in them.
Ranking is the product, and the ranking is what the card is scoped to. A reader learns which models sit inside the surface instead of how one classifier scored on a benchmark.
OpenAI and Anthropic: One Card per Frontier Release
OpenAI’s GPT-5 system card labels the fast, high-throughput models as gpt-5-main and gpt-5-main-mini, and the thinking models as gpt-5-thinking and gpt-5-thinking-mini. Once usage limits are reached, a mini version of each model handles remaining queries. Naming every component is the part a model-scoped document has no room for.
Anthropic’s system card for Claude Opus 4.5 describes it as a frontier model with a range of powerful capabilities. Those show most prominently in areas such as software engineering and in tool and computer use. The same practice runs at OpenAI, whose deployment-safety index carries the GPT-5.5 System Card dated April 23, 2026.
Red Hat: a System Card for a Shipped Product
Red Hat introduced the AI system card for the recently released Ask Red Hat conversational chatbot. The card articulates the system’s intent and scope, offering stakeholders a concise view into its purpose, boundaries, and trust posture.
For a security reader, this is the useful shape. The document names the components and the data sources of a product already in customers’ hands. That turns it into a disclosure surface rather than a research nicety.
Across all three examples, the document is scoped to the deployed system and its safeguards instead of to a set of weights.
Is a System Card the Same as a Model Card?
No. The difference is scope. A model card documents a single trained model, while a system card covers the deployed system built around it, safeguards and deployment context included.
Meta gave the reason directly. Machine learning models do not always work in isolation to produce outcomes. Models may also interact differently depending on what systems they are a part of. Model cards therefore do not paint a comprehensive picture of what an AI system does.
The practical test is the object the document describes. A card that stops at the weights is doing the model card’s job under a larger name.
Are System Cards Required?
No captured source states that any jurisdiction requires one. The organizations named here publish voluntarily, on infrastructure they run themselves.
OpenAI maintains its own deployment-safety page and posts dated cards for recent releases. Meta framed its own card as one of the ways it was exploring increased explainability, through model and system documentation. Voluntary publication puts the burden on the reader to check whether a card exists and what it covers. Binding rules across jurisdictions are cataloged in the AI regulation tracker.
Conclusion
A system card outlines the AI models that comprise an AI system and gives a reader insight into that system’s underlying architecture. Why that scope matters shows up in OpenAI’s own description of GPT-5. The company documents it as a unified system with a fast model, a deeper reasoning model, and a real-time router between them. Everything the format picked up around that core hangs off one scoping choice.
The direction of travel runs from a single-model spec sheet toward a description of the whole deployed thing. The reader’s job does not change with it. Ask what object the document describes. If it stops at the weights, the cover is promising work the contents do not do.