Observed 19 September 2026
Anthropic added new model versions (Claude Fable 5.1 and Claude Mythos 5.1) with a full transparency summary, including safety, alignment, and capability evaluation details.
what changed matterssummary written by a model — check it against the diff below
3 lines of binding language were added — this changes what someone is permitted or required to do.
+48 −1 · a83189d → dbfd955 · dbfd9550426a · live page
anthropic/transparency-hub.txt19 Sept 2026
12 unchanged lines Model Report August 17, 2026 Select a model to see a summary that provides quick access to essential information about Claude models, condensing key details about the models' capabilities, safety evaluations, and deployment safeguards. We've distilled comprehensive technical assessments into accessible highlights to provide clear understanding of how the models function, what they can do, and how we're addressing potential risks. Claude Opus 5 Claude Sonnet 5 Claude Fable 5 Claude Mythos 5 Claude Opus 4.8 Claude Opus 4.7 Claude Mythos Preview Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4 and Sonnet 4 Claude Opus 4.1 Claude Sonnet 3.7 Claude Fable 5.1 Claude Mythos 5.1 Claude Opus 5 Claude Sonnet 5 Claude Fable 5 Claude Mythos 5 Claude Opus 4.8 Claude Opus 4.7 Claude Mythos Preview Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4 and Sonnet 4 Claude Opus 4.1 Claude Sonnet 3.7 Claude Fable 5.1 Summary Table Model description Claude Fable 5.1 is the world’s most advanced model for coding and knowledge work—and its research capabilities offer an early glimpse of how AI models will contribute to scientific progress. Benchmarked Capabilities See our Claude Fable 5.1 & Claude Mythos 5.1 system card ’s Section 8 on capabilities. Acceptable Uses Anthropic’s Usage Policy applies. Release date September 2026 Modalities Claude Fable 5.1 can understand both text (including voice dictation) and image inputs, engaging in conversation, analysis, coding, and creative tasks. Claude can output text, including text-based artifacts, and diagrams. Model architecture and training methodology Claude Fable 5.1 was pretrained on large, diverse datasets to acquire language capabilities. After the pretraining process, Fable 5.1 underwent substantial post-training, with the goal of making it an effective assistant whose behavior aligns with the values described in Claude’s constitution . Training Data Claude Fable 5.1 was trained on a proprietary mix of publicly available information from online sources, public and private datasets, user data, and synthetic data generated by other models. Throughout the training process we used several data cleaning and filtering methods, including deduplication and classification. Testing Methods and Results Based on our assessments, we deployed Claude Fable 5.1 with ASL-3 protections, treating it as having CB-1 capabilities. Autonomy threat model 1 is applicable to Claude Fable 5.1. See below for select safety evaluation summaries. See Fable 5.1 & Mythos 5.1 System Card Claude Mythos 5.1 and Claude Fable 5.1 are the same model, but with different levels of safeguards. Claude Fable 5.1 is generally available, while Claude Mythos 5.1 is available only through our trusted access programs and Claude Security; its safeguards are specifically designed to support work in cybersecurity and the life sciences.Evaluations that don't engage those safeguards can be run on either configuration, so the key safety results summarized below apply to both Claude Fable 5.1 and Claude Mythos 5.1. Additional evaluations were conducted as part of our safety process; for our complete publicly reported evaluation results, please refer to the full system card . Safeguards Evaluation Claude Fable 5.1/Mythos 5.1 showed comparable overall performance to Mythos 5 in the suicide and self-harm domain. In terms of improvements, Mythos 5.1 was less likely than Claude Mythos 5 to suggest substitution methods for self-harm (e.g., holding ice cubes). These methods are clinically contested and have not been shown in research to meaningfully reduce self-harm incidents. Mythos 5.1 was also more likely to ask users directly whether they were experiencing suicidal thoughts, and to distinguish suicidal ideation from urges toward non-suicidal self-harm. Finally, Mythos 5.1 did not make unconditional assurances about the confidentiality of crisis line services, such as promising that a call would never result in contact with emergency services. One area for improvement was a tendency to implicitly validate self-harm as a coping strategy by acknowledging that it can regulate difficult emotions or provide relief. Mythos 5.1 also sometimes validated a user's fears about seeking help and mildly amplified prior negative experiences with crisis services. When these statements appeared, they were consistently within responses that discouraged self-harm and directed the user toward human support, but we consider them undesirable regardless of the surrounding context. Ahead of launch, we updated the claude.ai system prompt to address these behaviors, adding language that steers Claude away from describing self-harm as effective even when the user asserts this themselves. This partially mitigated the weaknesses described above. Existing system prompt language, including directions not to suggest substitution methods involving physical discomfort, further reinforced the improvements observed in the core model. We are continuing to explore how best to respond in sensitive mental health contexts, and we encourage developers building on the API to apply comparable safeguards and robust mitigations in contexts where users may be in distress. Claude is not a substitute for professional advice or medical care and is not intended to diagnose or treat any medical condition. Every Claude model is trained to detect and respond to expressions of distress (including if someone expresses personal struggles with suicidal or self-harm thoughts) with empathy and care, while pointing users toward human support when appropriate: helplines, mental health professionals, or trusted friends or family. Agentic Coding Across both scenarios, the helpful-only version of Claude Mythos 5 showed lower overall success rates than Claude Opus 4.8 and was modestly above or on par with Claude Mythos Preview. It's our assessment that these models would require substantial human direction for many operational steps. We use Shade, an adaptive red teaming tool from the external security firm Gray Swan, to evaluate the robustness of Claude models in computer-use environments, where the model interacts with the screen directly by clicking, typing and scrolling. The attacker runs on 14 test cases, with 200 attempts at each, and we measure success over all attempts. We compare model robustness with and without the additional safeguards we have designed to protect users in this setting. In computer-use environments, Claude Fable 5.1 achieved an attack success rate of 0.07%, corresponding to two successful attempts out of 2,800 across two of the 14 scenarios. This is comparable to Claude Opus 5 (0.18%, or five successful attempts), because the difference is not distinguishable from noise at these low absolute rates. Both are more robust than Claude Fable 5 (2.50%) and Claude Sonnet 5 (2.25%). With our safeguards enabled, Claude Fable 5.1's attack success rate is unchanged at 0.07%. Alignment Evaluations Claude Mythos 5.1 and Claude Fable 5.1 are the same model, but with different levels of safeguards. This means many of our evaluations can apply to either model. Our automated audit finds that Claude Mythos 5.1 is somewhat more honest and hallucinates less than previously released models. It hallucinates inputs, meaning it invents or materially misrepresents the contents of files or earlier user messages, significantly less than previous models in our investigations. Additionally Mythos 5.1 has lower rates of falsely claiming that tasks have been completed when they haven't and is less sycophantic to the user overall, meaning it is less prone to unprompted excessive praise, agreement or apology. Other metrics related to misleading the user are broadly similar to recently released models. Scores from our automated behavioral audit for the dishonesty-related metrics given below. Lower numbers represent a lower rate or severity of the measured behavior; on all graphs in this figure lower is better. The y-axis is truncated below the maximum score of 10 in many cases. Reported scores are averaged across all approximately 4,100 investigations per target model (approximately 2,100 seed instructions, with the conversation scenarios investigated once by each of two investigator models), with each investigation generally containing many individual conversations. Shown with 95% CI. RSP Evaluations Our Responsible Scaling Policy (RSP) evaluation process is designed to systematically assess our models' capabilities in areas where they could pose catastrophic risks before we release them. Because Claude Fable 5.1 and Claude Mythos 5.1 share the same underlying model, these evaluations were run on Claude Mythos 5.1, the configuration without the additional biology and cybersecurity safeguards, so that they reflect the model's underlying capabilities. Claude Fable 5.1 is our most capable general release model to date, advancing on both Claude Fable 5 and Claude Opus 5. - Alignment risk, which is the risk that a model behaves in ways Anthropic did not intend, is low: in our August 2026 Risk Report we moved this assessment from very low to low to reflect increased uncertainty following recent incident disclosures related to model behavior in cybersecurity evaluations, and while Claude Mythos 5.1 shows somewhat stronger covert capabilities than prior models, we do not believe it raises risk beyond that assessment. - On automated AI research and development, Claude Mythos 5.1's capabilities are comparable to or slightly stronger than those of Claude Mythos 5 and Opus 5, our previous frontier in this area, but it does not cross the RSP capability threshold. We have not observed a sustained doubling in the pace of our AI progress attributable to AI, and the model is not close to substituting for our research scientists and engineers; external testing by METR produced findings consistent with this assessment. - On chemical and biological weapons, it is difficult to say with full confidence whether any model passes our threshold for providing significant uplift in the construction of non-novel CB weapons. However, Claude Mythos 5.1 is broadly more capable than previous models we have conservatively treated as able to significantly help individuals with basic technical backgrounds produce (non-novel) weapons, so we treat it as having that capability and deploy commensurate safeguards, including real-time classifiers to prevent harm. With these mitigations we believe catastrophic risk in this category is low but not negligible. - For novel weapons development, Claude Mythos 5.1 shows modest gains over Claude Mythos 5 and Opus 5 on our automated evaluations. However, expert red teaming indicates that significant weaknesses remain, including weak novel ideation, poor strategic judgment and poor technical calibration, which prevent the model from substituting for scarce human expertise, and we conclude that Claude Mythos 5.1 does not cross the threshold for novel weapons capabilities. We apply the same expanded biology safeguards we applied to Claude Mythos 5, which restrict access to dual-use research biology capabilities. We describe these safeguards further in our most recent Risk Report . Claude Mythos 5.1 Summary Table Model description Claude Mythos 5.1 is the world’s most advanced model for coding and knowledge work—and its research capabilities offer an early glimpse of how AI models will contribute to scientific progress. Benchmarked Capabilities See our Claude Fable 5.1 & Claude Mythos 5.1 system card ’s Section 8 on capabilities. Acceptable Uses Anthropic’s Usage Policy applies.Note that this model is available only through our trusted access programs and Claude Security . Release date September 2026 Modalities Claude Mythos 5.1 can understand both text (including voice dictation) and image inputs, engaging in conversation, analysis, coding, and creative tasks. Claude can output text, including text-based artifacts, and diagrams. Model architecture and training methodology Claude Mythos 5.1 was pretrained on large, diverse datasets to acquire language capabilities. After the pretraining process, Mythos 5.1 underwent substantial post-training, with the goal of making it an effective assistant whose behavior aligns with the values described in Claude’s constitution . Training Data Claude Mythos 5.1 was trained on a proprietary mix of publicly available information from online sources, public and private datasets, user data, and synthetic data generated by other models. Throughout the training process we used several data cleaning and filtering methods, including deduplication and classification. Testing Methods and Results Based on our assessments, we deployed Claude Mythos 5.1 with ASL-3 protections, treating it as having CB-1 capabilities. Autonomy threat model 1 is applicable to Claude Mythos 5.1. See below for select safety evaluation summaries. See Fable 5.1 & Mythos 5.1 System Card Claude Mythos 5.1 and Claude Fable 5.1 are the same model, but with different levels of safeguards. Claude Fable 5.1 is generally available, while Claude Mythos 5.1 is available only through our trusted access programs and Claude Security; Mythos’ safeguards are specifically designed to support work in cybersecurity and the life sciences. For our complete publicly reported evaluation results, please refer to the full system card . Cyber Evaluation ExploitGym is a large-scale public benchmark developed by researchers at UC Berkeley together with collaborators at the Max Planck Institute for Security and Privacy, UC Santa 50 Barbara, Arizona State University, Anthropic, OpenAI, and Google. The benchmark measures whether AI agents can turn known vulnerabilities into working exploits. ExploitGym targets a wide range of applications and mainly focuses on converting a crash into unauthorized code execution. The task is to develop an exploit that achieves unauthorized code execution and reads a secret flag (a secret password) stored outside the agent’s authorized scope. Because agents sometimes achieve code execution through a different flaw than the one provided, successful flag captures are additionally verified by an agent-as-a-judge check. The chart below reports the number of exploits that successfully use the target vulnerability. In all cases, the judge that decides whether an exploit uses the intended vulnerability is the same model that is being evaluated. Claude Mythos 5.1 is an improvement over Mythos 5 on ExploitGym. Bars show the number of successful exploits using the given vulnerability under 2-hour and 6-hour wall-clock budgets. Success requires 51 capturing a dynamically generated secret flag, and each successful exploit is verified by an agent judge to confirm it uses the intended vulnerability. Bio Evaluation This benchmark assesses Claude's protein design abilities. The Sequence Generation variant assesses Claude's ability to create new proteins that meet a set of requirements, such as a plain-language description of its function, overall shape or how compact it is, and other specific structural features (e.g. whether it needs to attach to another protein). It is scored on whether the design meets those requirements, how likely it is to fold into a stable structure, and how different it is from proteins that already exist (to confirm it is designing rather than copying). Claude Mythos 5.1 achieved 46.0%, ahead of Claude Opus 5 at 42.4%, Claude Mythos 5 at 40.4% and Claude Sonnet 5 at 20.2% The Library Ranking variant assesses Claude's ability to decide which protein designs are worth testing in the lab first. Each problem gives Claude one set of candidate proteins along with real lab results from an actual protein engineering project, tells it what the researchers are trying to optimize for, and asks it to rank a new set of candidates from the same project that it has not seen results for. Claude Mythos 5.1 achieved 49.3%, ahead of Claude Mythos 5 at 48.2%, Claude Opus 5 at 48.0% and Claude Sonnet 5 at 41.3%. Claude Opus 5 Summary Table Model description Claude Opus 5 is a thoughtful and proactive model that comes close to frontier intelligence. On some coding and knowledge work evaluations Opus 5 is the new state-of-the-art. Benchmarked Capabilities See our Claude Opus 5 system card ’s Section 8 on capabilities. 486 unchanged lines
Underlined words mark what moved within a rewritten line. Nothing here is interpreted — this is the difference between two captures, and the conclusion is yours.