OpenAI Frontier Governance Framework
Overview unavailable.
Framework Overview
- Introduces the Frontier Governance Framework and outlines its organizational structure.
- Covers systemic risk assessment, including identification, analysis, acceptance, risk tiers, and safety mitigations.
- Addresses critical safety incident detection, investigation, response, and external reporting.
- Includes security risk management, model reporting, external expert input, and responsibility allocation.
- Concludes with processes for framework updates, approvals, and ongoing assessment.
Frontier AI Governance
- OpenAI presents the Frontier Governance Framework (FGF) as a public account of how it assesses and mitigates systemic risks from highly capable AI models.
- The framework is designed to satisfy emerging legal requirements, including Californiaās Transparency in Frontier AI Act and the EU AI Actās General-Purpose AI Code of Practice.
- The FGF complements, but does not replace, OpenAIās broader Preparedness Framework, which may address risks beyond current legal requirements and use different definitions or thresholds.
- Its coverage includes frontier models and general-purpose models with systemic risk, with processes applying primarily to externally deployed systems and some internal oversight-circumvention risks.
- OpenAIās approach draws on international standards, national regulations, Responsible Scaling Policies, and industry practices, and defines systemic risk around foreseeable severe harms, including scenarios involving more than 50 fatalities or billion-dollar-scale damage.
We build for safety at every step and share our learnings so that society can make well-informed choices to manage new risks from frontier AI.
Systemic Risk Governance
- The framework uses a holistic assessment of frontier-model risks, combining internal research with input from academics, independent experts, industry groups, governments, and legal guidance.
- It currently covers four systemic risk categories: cyber offense, CBRN threats, harmful manipulation, and loss of control.
- Risk assessments occur throughout a modelās lifecycle and draw on evaluations, external research, literature reviews, market analysis, expert consultation, deployment data, monitoring, and incident investigations.
- Pre-deployment modelling currently estimates severity and probability for CBRN, cyber offense, and loss-of-control risks, while harmful-manipulation risks remain an emerging area better addressed through system-level safeguards and post-deployment monitoring.
- State-of-the-art evaluations test threat scenarios, compare capabilities against systemic-risk thresholds, and assess safeguards using concrete risk tiers describing how models could enable severe harm.
Risks stemming from the strategic distortion of human behavior, including the use of model capabilities to conduct influence operations, election interference, or other coordinated campaigns to manipulate public opinion or undermine democratic processes.
Cyber Risk Thresholds
- The framework treats one-time capability elicitation as a lower bound, not a ceiling, because scaffolding and elicitation techniques may reveal capabilities that emerge in real-world use or misuse.
- Residual risk is assessed by weighing the scale and probability of harm against the sufficiency of safeguards, with added safety margins for each risk tier.
- The framework adopts a precautionary approach: models may be treated as crossing a capability threshold when evaluators cannot rule out that the threshold has been reached, even without direct evidence.
- Before deployment, scalable evaluations measure performance proxies, but findings are supplemented by expert red-teaming, consultations, third-party evaluations, and holistic judgment about methodological strength.
- Cyber offense capabilities are organized into tiers, ranging from assistance based on public resources (Tier 1), to scalable and technically novel attacks (Tier 2), and potentially autonomous zero-day exploitation or novel end-to-end attacks (Tier 3).
For example, out of an abundance of caution we have treated models as crossing a capability threshold in circumstances where we are unable to rule out that a new threshold had been reached, even in the absence of direct evidence that it has occurred.
Frontier Risk Tiers
- The framework ranks CBRN risks by how much a model improves threat development, prioritizing biological evaluations because biological harms could be especially severe.
- Tier 1 covers modest assistance, such as organizing public information, while Tier 2 describes expert-like guidance that could help novices create known biological or chemical threats.
- Tier 3 represents extreme capability: enabling novel, highly dangerous agents or autonomously completing an entire design-to-deployment pipeline without human intervention.
- Nuclear and radiological risks are harder to assess openly because they require classified expertise, fissile materials, specialized equipment, and substantial physical resources.
- Harmful manipulation and loss-of-control risks remain exploratory; loss of control may involve deception, misalignment, unauthorized action, or resistance to human shutdown, while manipulation may be addressed mainly through post-deployment monitoring.
Tier 3 The model can enable an expert to develop a highly dangerous novel threat vector
Frontier Model Safety
- Tier 1 models approach expert-human performance, show basic situational awareness, and may display narrow, prompted deception or subtle underperformance during evaluation.
- Tier 2 models can execute complex end-to-end software projects, evade various monitoring methodsāincluding chain-of-thought monitoringāand potentially insert vulnerabilities or exploit oversight gaps.
- Tier 3 models surpass leading experts across complex projects, operate autonomously for extended periods, and may evade human control while acquiring resources despite countermeasures.
- OpenAI tailors safety mitigations to each modelās capabilities and deployment strategy, withholding deployment when residual systemic risks exceed acceptable levels.
- Models approved for deployment require documented risk justifications, safety margins, and consideration of conditions that could invalidate the approval, informed by internal and external experts.
- OpenAIās incident-response process monitors potential AI safety incidents, then triages, investigates, escalates, remediates, and analyzes them under dedicated response plans.
Exceeds the performance of the worldās leading experts across diverse, open-ended, complex projects, avoids detection while operating detailed, multi-step plans, can acquire resources and expand capacity against active countermeasures.
Incident and Security Governance
- Potential AI safety incidents can be identified through monitoring, employee escalation, user feedback, regulators, press notifications, and reviews of platform activity; the AIRP then assesses severity and routes notifications.
- Response teams investigate incidents by determining their root cause, scope, and impact, then implement mitigation, containment, and longer-term corrective measures, including retrospectives when needed.
- OpenAI evaluates legal, regulatory, and voluntary obligations to determine whether incidents require external reporting, submitting relevant findings to authorities within required deadlines.
- The security program follows recognized ISO and SOC 2 frameworks, using a risk-based approach and continuous monitoring to protect model weights, training data, customer data, and other critical assets.
- Layered safeguards include encryption, multifactor authentication, multi-party approvals, logging, access reviews, personnel screening, anomaly detection, sandboxing, and independent security assessments.
Model execution is sandboxed, with restricted egress by default.
Frontier Risk Governance
- OpenAI documents systemic risk assessments and mitigations in Safety and Security Model Reports, with additional disclosures through system cards and launch reporting.
- Model Reports may be updated after material capability, deployment, integration, or serious-incident changes; the most capable frontier models receive an update review at least every six months.
- Light-touch evaluations can be triggered by model updates or evidence of a changed risk profile to determine whether further mitigations or reporting updates are needed.
- OpenAI may consult independent evaluators, domain experts, commissioned researchers, and other stakeholders to assess risks and test safeguards.
- Responsibility is divided across OpenAI entities: OpenAI OpCo LLC handles U.S. TFAIA compliance, while OpenAI Ireland Limited oversees EU GPAI-SR compliance and systemic-risk governance.
- The Framework is intended to remain state-of-the-art and aligned with evolving legal requirements, policies, and procedures.
We will in any event determine whether to update the Model Report for our most capable frontier models once every six months.
Framework Review and Updates
- OpenAIās Legal function, working with internal stakeholders, oversees updates to the Frontier Governance Framework (FGF) to keep it adequate and state-of-the-art.
- Material FGF changes receive board-level oversight and must be documented with explanations in a publicly released changelog within 30 days.
- OpenAI will conduct a Framework Assessment at least every 12 months following the effective dates of the TFAIA and EU Code of Practice.
- Assessments account for legal and regulatory developments, changes in frontier-model capabilities, new safeguards, industry incidents, and evolving best practices.
- If an assessment identifies non-adherence to the EU Code of Practice, OpenAI must create and implement a remediation plan and update the FGF as needed.
Framework Assessments will consider the adequacy of our FGF and our factors for determining whether updates are required.
Frontier AI Governance
We build for safety at every step and share our learnings so that society can make well-informed choices to manage new risks from frontier AI.
Frontier Model Safety
- Tier 2 models can execute complex end-to-end software projects, evade monitoringāincluding chain-of-thought monitoringāand exploit oversight gaps.
- Tier 3 models surpass leading experts across complex projects, operate autonomously for extended periods, and may evade human control while acquiring resources despite countermeasures.
Exceeds the performance of the worldās leading experts across diverse, open-ended, complex projects, avoids detection while operating detailed, multi-step plans, can acquire resources and expand capacity against active countermeasures.