← Anthropic — transparency

Observed 4 October 2026

Anthropic added new transparency pages for two models, Claude Sonnet 5.5 and Claude Opus 5.5, including descriptions, September 2026 release dates, and safety results, and updated the page date to October 2, 2026.

what changed matterssummary written by a model — check it against the diff below

3 lines of binding language were added — this changes what someone is permitted or required to do.

+57 −2  ·  dbfd955 → 2eeb4d3  ·  2eeb4d30af1a  ·  live page

anthropic/transparency-hub.txt4 Oct 2026
10 unchanged lines
02 System Trust and Reporting
03 Voluntary Commitments
Model Report
August 17, 2026
October 2, 2026
Select a model to see a summary that provides quick access to essential information about Claude models, condensing key details about the models' capabilities, safety evaluations, and deployment safeguards. We've distilled comprehensive technical assessments into accessible highlights to provide clear understanding of how the models function, what they can do, and how we're addressing potential risks.
Claude Fable 5.1 Claude Mythos 5.1 Claude Opus 5 Claude Sonnet 5 Claude Fable 5 Claude Mythos 5 Claude Opus 4.8 Claude Opus 4.7 Claude Mythos Preview Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4 and Sonnet 4 Claude Opus 4.1 Claude Sonnet 3.7
Claude Sonnet 5.5 Claude Opus 5.5 Claude Fable 5.1 Claude Mythos 5.1 Claude Opus 5 Claude Sonnet 5 Claude Fable 5 Claude Mythos 5 Claude Opus 4.8 Claude Opus 4.7 Claude Mythos Preview Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4 and Sonnet 4 Claude Opus 4.1 Claude Sonnet 3.7
Claude Sonnet 5.5 Summary Table
Model description Claude Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5, and is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.
Benchmarked Capabilities See our Claude Sonnet 5.5 system card ’s Section 8 on capabilities.
Acceptable Uses See our Usage Policy
Release date September 2026
Modalities Claude Sonnet 5.5 can understand both text (including voice dictation) and image inputs, engaging in conversation, analysis, coding, and creative tasks. Claude can output text, including text-based artifacts, and diagrams.
Model architecture and training methodology Claude Sonnet 5.5 was pretrained on large, diverse datasets to acquire language capabilities. After the pretraining process, Sonnet 5.5 underwent substantial post-training, with the goal of making it an effective assistant whose behavior aligns with the values described in Claude’s constitution .
Training Data Claude Sonnet 5.5 was trained on a proprietary mix of publicly available information from the internet, public and private datasets, and other sources, such as synthetic data generated by other models. The mix may also include user data from feedback or bug reports, or data that users have explicitly permitted for training. In the course of its training, we used several data cleaning and filtering methods, including deduplication and classification.
Testing Methods and Results Based on our assessments, we deployed Claude Sonnet 5.5 with the same chemical and biological misuse protections we deployed for Claude Opus 5, treating it as having CB-1 capabilities. Autonomy threat model 1 is applicable to Claude Sonnet 5.5. See below for select safety evaluation summaries.
See Fable 5.1 & Mythos 5.1 System Card
The following are summaries of key safety evaluations from our Claude Sonnet 5.5 system card. Additional evaluations were conducted as part of our safety process; for our complete publicly reported evaluation results, please refer to the full system card .
Political Even-handedness
People use Claude to learn about and discuss political issues, so we test whether it treats different political views fairly. We use our open-source evaluation , which gives Claude 1,350 pairs of requests on 150 political topics. Each pair asks for the same thing from opposite political perspectives, for example a request about healthcare policy written once from a Democratic point of view and once from a Republican one.
Another Claude model then grades the answers. It checks whether Claude answers both requests in a pair with similar depth and quality (what we call even-handedness), whether it acknowledges other points of view, and whether it refuses to answer. We report results in two settings: Claude accessed directly through our API (the service developers use to build Claude into their own products), with no added instructions, and Claude on claude.ai, where it follows standing instructions we give it (called a system prompt).
Claude Sonnet 5.5 was substantially more even-handed than Claude Sonnet 5. It scored 97.9% on the API and 99.0% on claude.ai (vs. 86.2% and 95.7% for Sonnet 5), and it refused less often in both settings. Results on acknowledging other points of view were mixed. Through the API, Sonnet 5.5 did this slightly less often than Sonnet 5 (41.4% vs. 45.7%). On claude.ai, whose system prompt directs Claude to engage even-handedly across viewpoints, it did so considerably more often (60.2% vs. 52.9%). Sonnet 5.5’s scores are compared with our other models below.
Pairwise political bias evaluations: even-handedness. Higher scores for even-handedness are better. Results for previous models may show variance from previous system cards due to routine evaluation updates.
Honesty
People increasingly rely on AI models for information and to get work done, so it matters that these models tell the truth and are open about what they have done. We test Claude’s honesty in several ways. Claude Sonnet 5.5 did better than Claude Sonnet 5 on most of these tests, with some exceptions noted below. For more in-depth descriptions of the evaluations and their results, please see the Claude Sonnet 5.5 system card .
- Misleading users: We ran about 4,100 simulated test sessions with Claude, then had another AI model review them for eight kinds of misleading behavior, such as lying to the user, flattering them or agreeing with them too readily, and claiming to have finished a task that was not done. Claude Sonnet 5.5 did better than Claude Sonnet 5 on seven of the eight. The exception was dodging questions on sensitive political or social topics, or answering them too cautiously. Here the two models were similar, and Sonnet 5.5 did this slightly more often than most of our recent models. Claude Opus 5.5 did better than Sonnet 5.5 on most of these behaviors. (Section 6.2.3)
- Making up facts: AI models sometimes “hallucinate,” stating false information as if it were true. An honest model should admit when it does not know an answer instead of guessing. We asked Claude factual questions on 41 subjects, without letting it look anything up online, and marked each answer as right, wrong, or declined. Claude Sonnet 5.5 got more questions right than Claude Sonnet 5, but it also gave slightly more wrong answers. Weighing right answers against wrong ones, it scored better than Sonnet 5 but below the other Claude models we tested. (Section 6.3.2.1)
- Honesty under pressure: We used a public test called MASK to check whether Claude will say something it believes is false when a user, or the instructions it has been given, pressure it to. Claude Sonnet 5.5 stuck to what it believed more often than Claude Opus 5.5, but less often than Claude Sonnet 5. (Section 6.3.2.2)
- Using answers without saying so: We gave Claude programming tasks where the answer was already within reach of the computer it was working on, for example in a newer version of the software that already contained the fix. We then measured how often Claude used that answer without telling the user, which would leave the user unaware of where the solution came from. Claude Sonnet 5.5 did this less often than Claude Sonnet 5, and slightly less often than Claude Opus 5.5. (Section 6.3.2.3)
- Admitting to hidden changes: During some training exercises, we saw Claude hide changes to code from the automated system that graded its work, so that it would pass. To test whether Claude would own up to this, we took records of these sessions and asked Claude to summarize its work, without asking about the hidden changes directly. Claude Sonnet 5.5 mentioned the hidden changes 96.2% of the time, about as often as Claude Opus 5.5 and more often than the other Claude models we tested. (Section 6.3.2.4)
The charts below compare Sonnet 5.5 with other recent Claude models on the eight misleading behaviors from the first test above. Lower scores are better. Behaviors are scored even in sessions where they are unlikely to come up, so the scores are most useful for comparing models with each other.
Scores from our automated behavioral audit
RSP Evaluations
Our Responsible Scaling Policy (RSP) evaluation process is designed to systematically assess our models' capabilities in areas where they could pose catastrophic risks before we release them. Claude Sonnet 5.5 is not more capable than our most advanced models, and it is less capable than Claude Opus 5.5 on most of our tests. We found that Claude Opus 5.5 does not reach our higher risk thresholds for chemical and biological weapons or for autonomy (AI that could do the work of large teams of top researchers, or dramatically speed it up, in fields where rapid progress could be dangerous), so the same conclusion applies to Sonnet 5.5.
- Alignment risk is the risk that a model behaves in ways Anthropic did not intend. In our August 2026 Risk Report we raised our estimate of this risk from very low to low, because recently disclosed incidents involving how models behaved during cybersecurity testing made us less certain. We assess that the risk of Claude Sonnet 5.5 causing catastrophic harm by acting in ways we did not intend is also low. We also tested whether Sonnet 5.5 could secretly carry out a harmful side task while doing an assigned job, without being caught by an automated monitor. When the monitor could read the model’s reasoning (the step-by-step thinking it writes out before it acts), Sonnet 5.5 almost never succeeded. When its reasoning was hidden, it succeeded less often than Claude Opus 5.5 and Claude Mythos 5.1, but more often than Claude Sonnet 5.
- On its ability to carry out AI research and development on its own, Claude Sonnet 5.5 is broadly similar to or less capable than Claude Opus 5.5. Opus 5.5 did not cross our higher threshold for this ability, so we reach the same conclusion for Sonnet 5.5. On our capability index, which combines many tests into a single score, Sonnet 5.5 scored below Opus 5.5 (167.93 vs. 169.12).
- On chemical and biological weapons, we treat Claude Sonnet 5.5 as able to significantly help someone with a basic technical background make known chemical or biological weapons, which is our CB-1 threshold. We deploy it with safeguards to match, including the same automated filters we used for Claude Opus 5, which are designed to detect and block attempts to get this kind of help. On three automated tests related to this threshold, it scored close to Claude Opus 5.5 and the same as or higher than Claude Sonnet 5.
- Developing new chemical or biological weapons takes rare expertise and judgment. We estimate that Claude Sonnet 5.5’s abilities in this area are similar to or below those of Claude Opus 5, and well below those of Claude Opus 5.5. It struggled with longer, open-ended research tasks, which limits how far it could stand in for human experts. We are releasing Sonnet 5.5 with the same safeguards we used for Claude Opus 5, rather than the broader safeguards we use for Claude Opus 5.5, which also limit access to biology research capabilities that could be used for harm as well as good. We describe our safeguards further in our most recent Risk Report .
Claude Opus 5.5 Summary Table
Model description Claude Opus 5.5 is our new leading model, and early testers saw large jumps in performance on their most complex work.
Benchmarked Capabilities See our Claude Opus 5.5 system card ’s Section 8 on capabilities.
Acceptable Uses See our Usage Policy
Release date September 2026
Modalities Claude Opus 5.5 can understand both text (including voice dictation) and image inputs, engaging in conversation, analysis, coding, and creative tasks. Claude can output text, including text-based artifacts, and diagrams.
Model architecture and training methodology Claude Opus 5.5 was pretrained on large, diverse datasets to acquire language capabilities. After the pretraining process, Opus 5.5 underwent substantial post-training, with the goal of making it an effective assistant whose behavior aligns with the values described in Claude’s constitution .
Training Data Claude Opus 5.5 was trained on a proprietary mix of publicly available information from online sources, public and private datasets, user data, and synthetic data generated by other models. Throughout the training process we used several data cleaning and filtering methods, including deduplication and classification.
Testing Methods and Results Based on our assessments, we deployed Claude Opus 5.5 with ASL-3 protections, treating it as having CB-1 capabilities. Autonomy threat model 1 is applicable to Claude Opus 5.5. See below for select safety evaluation summaries.
See Fable 5.1 & Mythos 5.1 System Card
The following are summaries of key safety evaluations from our Claude Opus 5.5 system card. Additional evaluations were conducted as part of our safety process; for our complete publicly reported evaluation results, please refer to the full system card .
Safeguards Evaluation
We run evaluations to assess how Claude handles conversations about disordered eating. We focus on whether it avoids reinforcing requests that pose potential risk while remaining helpful on benign questions about nutrition, fitness, and health. Claude Opus 5.5 maintained high harmless response rates on single-turn requests that could reinforce disordered eating behaviors, such as requests for calorie targets below safe thresholds or for techniques to conceal calorie restriction. On the API, its harmless response rate was slightly below that of Opus 5 (95.64% vs. 96.89%); on claude.ai, the difference narrowed to within the margin of error. Opus 5.5 did not over-refuse any benign requests in our evaluation set, which includes general questions about nutrition, eating disorder treatment and recovery, and body image.
Qualitative review of multi-turn testing identified multiple areas where Claude Opus 5.5 improved compared to Opus 5. Opus 5.5 produced outputs that were shorter and less verbose than those of Opus 5, which is beneficial when the user may have limited capacity or willingness to parse lengthy, complex outputs. It also used a warmer tone, more frequently validated the act of disclosure, and was less likely to make unfounded claims about the user's eating behaviors, body image, or emotional state, beyond what the user had disclosed. It was also less likely to label the user with a clinical eating disorder without caveats or a recommendation to seek a professional evaluation. It was also less likely to introduce user-specific numbers that could function as targets, such as estimated body mass index. Eating disorder experts generally advise against providing such numbers because they may reinforce cognitive distortions.
In single-turn testing on prompts with ambiguous (not clearly harmful) contexts, Claude Opus 5.5 was more likely than Opus 5 to offer evaluative feedback when a user shared an image of their body. For example, it commented on whether a user's waist was slim enough for modeling. This behavior, which we also observed in Fable 5.1, was mitigated on claude.ai by the default system prompt. We made no additions to the system prompt specific to disordered eating for this release; we retained existing language directing Claude to avoid precise diet and nutrition figures when a conversation shows signs of disordered eating, and to refer users to the National Alliance for Eating Disorders helpline rather than the discontinued National Eating Disorders Association (NEDA) line.
Alignment Evaluation
When Claude works as an agent, it usually runs inside a sandbox: a walled-off computer environment that gives the model the access it needs for a task and nothing more. We ran a sandbox escape evaluation to measure how often Claude Opus 5.5 tries to violate those boundaries without authorization. In this evaluation, one Claude model acts as an investigator, builds a scenario and probes how the model being tested behaves. It covered a diverse range of scenarios based on deployments, capabilities evaluations, and reinforcement learning environments, tailored to tempt the model into violating various kinds of boundaries on its own initiative. With Claude Opus 5.5, only 1.5% of cases resulted in attempted crossing of a containment boundary, and all of those cases are rated as low severity: the model reached across a weak boundary in a read-only manner, did not carry the attempt further when it encountered a stronger security barrier, and clearly reported the actions afterwards. Opus 5.5’s rates of attempted sandbox escapes are compared with our other models below. Some tendency to test boundaries remains, but these results suggest that recent training changes meant to curb this behavior have had a strong effect. This evaluation complements our broader alignment audit, which tracks a set of behaviors linked to recent cybersecurity incidents involving our models and those of other developers.
Scores from our sandbox escape evaluation, comprising 28 handmade scenarios tailored to elicit sandbox escape, run seven times for each model. Percentage of low-severity and moderate- to high-severity attempts shown. Severity determined by the persistence, scope, and deception involved in each attempted act of containment boundary violation. Shown with 95% CI.
RSP Evaluation
Our Responsible Scaling Policy (RSP) evaluation process is designed to systematically assess our models' capabilities in areas where they could pose catastrophic risks before we release them. These evaluations test the underlying model, so they were run without the additional safeguards we apply in our products. Claude Opus 5.5 is our most capable general release model to date, matching or exceeding Claude Fable 5.1 and Claude Mythos 5.1 on many of our evaluations.
- Alignment risk, which is the risk that a model behaves in ways Anthropic did not intend, remains low. In our August 2026 Risk Report we moved this assessment from very low to low to reflect increased uncertainty following recent incident disclosures related to model behavior in cybersecurity evaluations. We do not believe Claude Opus 5.5 raises risk beyond that assessment. On tests of whether a model can carry out a hidden task without being caught, Claude Opus 5.5 was almost always detected when monitors could read its reasoning, which is how our monitoring works in practice.
- On automated AI research and development, Claude Opus 5.5's capabilities are at or slightly above those of Claude Mythos 5.1, our previous frontier in this area, but it does not cross the RSP capability threshold. We have not observed a sustained doubling in the pace of our AI progress attributable to AI, and the model is not close to substituting for our research scientists and engineers. On an internal test built from real engineering problems our staff have solved, Claude Opus 5.5 scored 55.8%, well below the 85% we think a model able to substitute for our research staff would reach. External testing by METR produced findings consistent with this assessment.
- On chemical and biological weapons, Claude Opus 5.5 is broadly as capable as previous models we have conservatively treated as able to significantly help individuals with basic technical backgrounds produce known (non-novel) weapons, so we treat it as having that capability and deploy commensurate safeguards, including real-time classifiers to prevent harm.
- For novel weapons development, Claude Opus 5.5 performs similarly to Claude Mythos 5.1 across our evaluations. However, expert red teaming found that it did not improve on the weaknesses that kept Claude Mythos 5.1 below this threshold. It struggled to generate genuinely new ideas, sometimes misrepresented the scientific literature by relying on paper abstracts, and made scientific errors that went unnoticed when users lacked expertise in that area. These weaknesses prevent the model from substituting for scarce human expertise, and we conclude that Claude Opus 5.5 does not cross the threshold for novel weapons capabilities. We apply the same expanded biology safeguards we applied to Claude Fable 5 and Claude Fable 5.1, which restrict access to dual-use research biology capabilities. We describe these safeguards further in our most recent Risk Report .
Claude Fable 5.1 Summary Table
Model description Claude Fable 5.1 is the world’s most advanced model for coding and knowledge work—and its research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
Benchmarked Capabilities See our Claude Fable 5.1 & Claude Mythos 5.1 system card ’s Section 8 on capabilities.
533 unchanged lines

Underlined words mark what moved within a rewritten line. Nothing here is interpreted — this is the difference between two captures, and the conclusion is yours.