OpenAI Preparedness Framework (Version 2, Updated April 15, 2025)
Overview unavailable.
Preparedness Framework Overview
- Version 2, updated April 15, 2025, describes OpenAI’s approach to preparing for frontier AI capabilities that could cause severe harm.
- It focuses on three Tracked Categories: biological and chemical risks, cybersecurity risks, and AI self-improvement risks.
- For each category, OpenAI develops threat models, measures capability thresholds, and requires safeguards before deploying highly capable models.
- The update introduces Research Categories for emerging risks that need further study, threat modeling, and evaluation techniques.
- The framework covers risk categorization, capability measurement, safeguard design and sufficiency, governance, transparency, decision-making, and illustrative controls.
Preparing for Frontier Risks
- OpenAI focuses its safety efforts on a limited set of AI capabilities that could create severe harm, using holistic risk assessments to determine which risks and thresholds require attention.
- Frontier models are evaluated before deployment and throughout development, with capability-elicitation methods designed to detect levels associated with meaningful increases in severe risk.
- Models that reach High capability thresholds are not deployed until risks are sufficiently minimized; Critical capabilities also require adequate safeguards during development, regardless of deployment plans.
- The Safety Advisory Group oversees the Preparedness Framework, while OpenAI Leadership and the Board’s Safety and Security Committee provide decision-making and oversight.
- Rapid progress toward more scientific and agentic systems, along with more frequent deployments, is making scalable evaluations, stronger safeguards, and greater cross-company coordination increasingly necessary.
We are on the cusp of systems that can do new science, and that are increasingly agentic - systems that will soon have the capability to create meaningful risk of severe harm.
Frontier AI Risk Framework
- As frontier AI labs proliferate, safety work must expand to address reasoning models, increasingly agentic systems, and the need for stronger protective measures.
- Research and deployment over the past year have strengthened the field’s understanding of how to prioritize risks through threat modeling, capability evaluations, external consultation, and updated industry frameworks.
- Frontier capabilities are assessed holistically using internal research alongside input from academics, independent experts, industry groups, governments, and relevant policy mandates.
- Capabilities become Tracked Categories when they present plausible, measurable, severe, net-new, and instantaneous or irremediable risks.
- Research Categories cover capabilities that may contribute to severe harm but do not yet meet the threshold for tracking; ongoing work may eventually elevate them to Tracked Categories.
Instantaneous or irremediable: The outcome is such that once realized, its severe harms are immediately felt, or are inevitable due to a lack of feasible measures to remediate.
Frontier Risk Thresholds
- The framework creates threat models for each tracked category, identifying severe-harm risks and the capability thresholds that could substantially increase them.
- High thresholds mark capabilities that intensify known risks; systems crossing them need strong safeguards before deployment and appropriate security controls during development.
- Critical thresholds signal qualitatively new and potentially unprecedented threat vectors, requiring safeguards throughout development regardless of deployment plans.
- Teams develop evaluations to monitor progress and determine when systems may have reached these thresholds.
- For biological and chemical capabilities, risks range from enabling novices to create known threats to allowing experts or autonomous systems to engineer novel, highly dangerous biological agents.
- Critical biological capabilities could enable mass-casualty events and major social disruption, prompting a halt in development until adequate safeguards and security controls exist.
Proliferating the ability to create a novel threat vector of the severity of a CDC Class A biological agent (i.e., high mortality, ease of transmission) could cause millions of deaths and significantly disrupt public life, with few available societal safeguards.
Critical Capability Safeguards
- At the High cyber capability level, models can automate end-to-end operations against hardened targets or discover and exploit operationally relevant vulnerabilities, potentially scaling cyberattacks and disrupting the offense-defense balance.
- Models combining cyber capabilities with long-range autonomy or the ability to bypass safeguards could undermine OpenAI’s ability to monitor and mitigate risks; high-standard security, misuse, and misalignment protections are therefore required.
- Critical cyber capability is reached when models can independently develop zero-day exploits across hardened critical systems or execute novel end-to-end attacks, creating catastrophic risks and requiring development to halt without adequate safeguards.
- High self-improvement capability is defined as performance comparable to giving every OpenAI researcher a highly capable mid-career research engineer, signaling that AI development may already be accelerating.
- Critical self-improvement involves recursively automated AI research or sustained, dramatically faster generational progress, which could cause emerging risks to outpace oversight and threaten human control of AI systems.
A major acceleration in the rate of AI R&D could rapidly increase the rate at which new capabilities and risks emerge, to the point where our current oversight practices are insufficient to identify and mitigate new risks, including risks to maintaining human control of the AI system itself.
Research Categories and Safeguards
- Research Categories cover frontier capabilities that require further threat modeling or could undermine safeguards before they can be rigorously tracked.
- The framework calls for developing threat models, improving measurement and evaluations, and sharing findings publicly where feasible.
- Long-range Autonomy, Sandbagging, Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological capabilities each receive tailored monitoring or mitigation responses.
- Reviews will be provided to the Safety Advisory Group, which may recommend changes to internal practices or the Preparedness Framework.
- Version 2 separates AI Self-improvement, Long-range Autonomy, and Autonomous Replication and Adaptation as distinct capabilities formerly grouped under Model Autonomy.
Undermining Safeguards: ability and propensity for the model to act to undermine safeguards placed on it, including e.g., deception, colluding with oversight models, sabotaging safeguards over time such as by embedding vulnerabilities in safeguards code, etc.
Emerging AI Risk Categories
- Rapid, hard-to-track advances in AI capabilities are identified as a potentially irreversible risk with unpredictable, severe consequences.
- Long-range autonomy and autonomous replication remain Research Categories rather than fully tracked risks, but OpenAI is investing in them because their threat models may mature.
- Nuclear and radiological risks are treated as Research Categories due to the expertise, classified information, materials, equipment, and physical barriers required to develop such weapons.
- OpenAI is building safeguards against assistance with high-risk weapons queries, evaluating refusal performance, prioritizing nuclear-risk research, and consulting US national security stakeholders.
- Political manipulation is prohibited, while persuasion and relational capabilities are studied alongside safeguards, misuse monitoring, content-provenance efforts, and broader societal responses; these risks fall outside the Preparedness Framework’s severe-harm criteria.
Because of the significant resources required and the legal controls around information and equipment, nuclear weapons development cannot be fully studied outside a classified context.
Capability Evaluation Framework
- The organization develops high-precision, high-recall evaluations to determine whether covered systems have crossed dangerous capability thresholds.
- Evaluations aim to approximate the maximum capabilities an adversary could elicit, including advanced settings, low-refusal model variants, best available scaffolds, and—in some cases—fine-tuning.
- Because elicitation methods continually improve, any single evaluation is treated as a lower bound on real-world capabilities, requiring ongoing monitoring and reassessment.
- The framework combines scalable automated evaluations with deep dives involving expert red-teaming, consultations, laboratory studies, and independent assessments.
- Biological-threat evaluations examine both useful information and tool integration across five stages of weapon creation, while the framework covers frontier models, agents, major deployment changes, and unexpectedly capable updates.
Nonetheless, given the continuous progress in model scaffolding and elicitation techniques, we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling, on capabilities that may emerge in real world use and misuse.
Thresholds and Safeguards
- The Safety Advisory Group (SAG) makes final decisions about whether a model falls under the Preparedness Framework, especially when coverage is ambiguous.
- Before deployment, covered models undergo Scalable Evaluations, whose results and methodological caveats are compiled into a Capabilities Report for SAG review.
- Threshold decisions rely not only on evaluation results but also on holistic judgment about the strength and reliability of the available evidence.
- If a threshold is crossed—or if further evidence is needed—the SAG may recommend safeguards, conclude that no action is necessary, or commission deeper research.
- Safeguards are selected by mapping plausible severe-harm scenarios to specific protections and measurable efficacy standards, including separate defenses against malicious users and misaligned models.
We consider separate safeguards for two of the main ways in which risks can be realized: a malicious user, who can leverage the model to cause the severe harm, and a misaligned model, which autonomously causes the harm.
Safeguards Deployment Review
- The Safeguards Report maps ways severe harm could occur to relevant controls, evaluates their effectiveness, estimates residual risk, and notes limitations.
- Safeguards can often be reused across deployments, reducing the need for entirely new safeguards and stress tests.
- The Safeguards Advisory Group (SAG) assesses adequacy using the system’s capability level, threat model, expert advice, safeguard performance, and risks observed in comparable external systems.
- Safeguards address both malicious users—through robustness, monitoring, and controlled access—and misaligned models through alignment, oversight, restricted autonomy, and system architecture.
- High-capability systems must have adequate safeguards before deployment, while Critical-capability systems require them during development; SAG may approve deployment or demand further evaluation.
Covered systems that reach High capability must have safeguards that sufficiently minimize the associated risk of severe harm before they are deployed.
Adaptive Safety Governance
- The SAG may recommend targeted, minimally disruptive changes to deployment conditions or stronger safeguards when existing measures do not sufficiently reduce severe-harm risks.
- Safeguards are expected to evolve through continuous review, especially when evidence suggests they are not working as intended.
- If another developer releases a highly capable system without comparable protections, OpenAI may adjust its safeguards only after rigorous confirmation, public acknowledgment, and assurances that it remains more protective—avoiding a race to the bottom.
- Models approaching Critical capability require heightened safety and security controls even before internal use or external deployment, with evaluations addressing both malicious actors and model misalignment.
- The framework emphasizes trust through internal documentation, employee reporting channels, accountability for noncompliance, and planned public disclosure of preparedness results.
Models that have reached or are forecasted to reach Critical capability in a Tracked Category present severe dangers and should be treated with extreme caution.
Frontier AI Governance
- OpenAI plans to publish information about major AI deployment decisions, including testing scope, capability evaluations, reasoning, and safeguards, while allowing redactions for intellectual property or safety.
- When warranted and feasible, third parties may independently evaluate tracked capabilities and stress-test safeguards, especially for models exceeding a High capability threshold.
- The Safety Advisory Group may seek independent opinions from domain experts, including specialists who are not AI experts, to strengthen its assessment of deployment risks.
- The revised framework clarifies that capabilities, risks, and safeguards are assessed holistically, and that reducing risk does not necessarily require reducing a model’s capabilities.
- High thresholds indicate a significant increase in existing severe-risk vectors, while Critical thresholds signal qualitatively new severe threats; tracked capabilities must be plausible, measurable, severe, net new, and instantaneous or irremediable.
Critical capability thresholds mean capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent under the relevant threat model.
Preparedness Framework Revisions
- The framework expands its safety scope, retaining restrictions on political persuasion and misuse while moving Nuclear and Radiological capabilities into Research Categories alongside autonomy, sandbagging, replication, and safeguard-undermining.
- OpenAI will develop stronger threat models and capability evaluations for Research Categories, using external experts and rigorous measurement methods even where capabilities are not yet formally tracked.
- Capability elicitation will test models using techniques that approximate the highest level of misuse expected from threat actors, distinguishing automated “scalable evaluations” from expert-led “deep dive” assessments.
- Safeguards will be tailored to specific risks, documented through Capabilities and Safeguards Reports, and evaluated directly for efficacy rather than merely by rerunning capability tests on safeguarded models.
- The framework shifts from one-off safety drills toward continuous red-teaming, considers marginal risk in light of other market systems, and assigns the Safety Advisory Group responsibility for reviewing residual risks and advising leadership.
We are shifting our attention to a more durable approach of continuously red-teaming and assessing the effectiveness of our safeguards.
Preparedness Governance and Safeguards
- The Safety Advisory Group (SAG) offers diverse perspectives and recommendations on catastrophic risk, but OpenAI Leadership retains final authority and the ability to act without SAG approval.
- OpenAI Leadership is responsible for accepting residual risks, making deployment decisions, and providing the resources needed to implement the Preparedness Framework.
- The Board’s Safety and Security Committee oversees these processes and may request information, reverse decisions, or require a different course of action.
- The Preparedness Framework is intended to evolve through at least annual reviews, with a fast-track process available when severe risks emerge urgently.
- Safeguards for high- and critical-capability systems must sufficiently reduce severe-harm risks, including misuse by malicious users; proposed approaches include robust refusal behavior and resistance to jailbreaks.
The Preparedness Framework is a living document and will be updated.
Safeguards Against Severe Harm
- The framework identifies three safeguards for malicious-user risks: model robustness, usage monitoring, and trust-based access.
- Robustness measures include refusal training, knowledge filtering, interpretability techniques, and jailbreak resistance, evaluated through red-teaming and benchmark results.
- Usage monitoring combines automated detection, human review, blocking, investigations, enforcement, and law-enforcement escalation, with performance measured by detection accuracy and response time.
- Trust-based access relies on identity verification, trusted relationships, usage policies, rate limits, and cautious expansion to broader user groups.
- The section then turns to misaligned models, emphasizing safeguards that prevent autonomous systems from circumventing human oversight and carrying out severe harms.
Safeguards should sufficiently minimize the risk associated with a misaligned model circumventing human control and oversight and executing severe harms.
Safeguarding Against Misalignment
- Severe-harm risks are addressed through several complementary claims: limited autonomy, value alignment, instruction alignment, reliable oversight, and restrictive system architecture.
- Value alignment is tested through adversarial red-teaming, monitoring, and deployment evidence, with deception or scheming treated as signs of failure.
- Instruction alignment requires models to follow vague or complex instructions while preserving human values, avoiding harm, and resisting behaviors such as escaping containers or deceiving users.
- Oversight safeguards include AI monitors, mandatory human approval, interpretable reasoning, immutable action logs, and escalation procedures that reliably detect and block dangerous behavior.
- Containerization and restricted permissions limit the model’s ability to access output channels or persist long enough to cause harm; robustness and usage monitoring provide additional protection against malicious use.
Evidence of taking harmful actions, including deception and scheming, would be evidence of misalignment
Frontier Model Security
- High-capability models require robust safeguards against both external misuse and internal adversaries, aligned with established security frameworks such as ISO 27001, SOC 2, NIST, and FedRAMP.
- Technical protections include containerization, sandboxing, restricted internet and filesystem access, limited credentials, reduced persistence, and rigorous testing or red-teaming.
- Security threat models must continuously identify vulnerabilities and attack vectors, then be updated and validated as technologies and threats evolve.
- A defense-in-depth strategy combines physical and datacenter security, network segmentation, workload isolation, encryption, and Zero Trust principles.
- Access should follow least privilege, separation of duties, strong multifactor authentication, managed devices, logging, and regular audits, alongside a secure development lifecycle and supply-chain controls.
“Canary evaluations” which test model capabilities to bypass less complex, easier-to-exploit versions of our security controls, establishing that our implemented controls are robust
Layered Security Governance
- Security reviews and penetration testing should be integrated into engineering, especially for high-sensitivity components before deployment.
- Formal change management requires authorization, documentation, testing, approval, multi-person review for critical infrastructure, and rollback procedures.
- Organizations should protect software supply chains by vetting reputable hardware, software, third-party suppliers, and libraries continuously.
- Operational security depends on 24x7 monitoring, rapid incident response, consistent patching, red-teaming, and bug bounty programs.
- Independent audits, transparent reporting, and management oversight provide accountability and help validate the effectiveness of security and risk-management programs.
Conduct adversarial testing and red-teaming exercises to proactively identify and mitigate potential vulnerabilities within corporate, research, and product systems, ensuring resilience against unknown vulnerabilities and emerging threats.
Preparing for Frontier Risks
We are on the cusp of systems that can do new science, and that are increasingly agentic - systems that will soon have the capability to create meaningful risk of severe harm.
Critical Capability Safeguards
- Critical cyber capability is reached when models can independently develop zero-day exploits across hardened critical systems or execute novel end-to-end attacks; development must halt without adequate safeguards.
- Critical self-improvement involves recursively automated AI research or dramatically faster generational progress, potentially outpacing oversight and threatening human control.
A major acceleration in the rate of AI R&D could rapidly increase the rate at which new capabilities and risks emerge, to the point where our current oversight practices are insufficient to identify and mitigate new risks, including risks to maintaining human control of the AI system itself.
Safeguarding Against Misalignment
- Severe-harm risks are addressed through limited autonomy, value and instruction alignment, reliable oversight, and restrictive system architecture.
- Oversight safeguards include AI monitors, mandatory human approval, interpretable reasoning, immutable action logs, and escalation procedures to detect and block dangerous behavior.
Evidence of taking harmful actions, including deception and scheming, would be evidence of misalignment