
Pacing model development in an era of cyber-critical capabilities
OpenAI has slowed the pace of scaling its frontier AI models to strengthen monitoring, alignment, and security safeguards. This action follows a security incident and evidence that the upcoming Astra model meets critical cybersecurity capability thresholds.
Why it matters
Advanced cybersecurity capabilities in AI increase the risk of unauthorized access or destructive attacks if systems are not properly secured. These measures protect internal and external networks from potential model-driven security breaches.
The details
- Astra model inference with tools has required additional monitoring since August 7.
- New security protocols include workload sandboxing and network isolation for high-risk research.
- Monitoring systems aim to issue alerts within 30 minutes of concerning activity.
Show entities and relationshipsHide entities and relationships
In this article
Companies
Products
Key connections
OpenAI uses Reinforcement Learning
OpenAI runs reinforcement learning (RL) training on its latest frontier models intended for deployment.
OpenAI uses Chain-of-Thought Monitoring
OpenAI expanded its chain-of-thought monitoring system, using activation classifiers and automated investigators to detect concerning model behavior.
Related events
OpenAI Pauses Frontier RL Training to Strengthen Cybersecurity Safeguards
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.