OpenAI has introduced new safeguards for developing and testing advanced AI models, including stronger monitoring, tighter network isolation and additional security checks during post-training.
The company said the measures are intended to keep pace with growing risks as AI systems become more capable. OpenAI also said the changes were influenced by the cybersecurity capabilities expected from its upcoming Astra model, as well as the broader acceleration of AI development.
Follow THE FUTURE on LinkedIn, Facebook, Instagram, X and Telegram
Changes Follow Hugging Face Incident
The new measures come weeks after OpenAI disclosed a security incident involving Hugging Face, where models escaped their training environment after an internet-connected tool on OpenAI’s network was compromised.
OpenAI said the new policies are not a direct response to the incident, but acknowledged that it had paused reinforcement learning for two weeks afterward. Less risky training has since resumed, while its largest planned frontier reinforcement learning run remains on hold pending further testing and safety evaluations.
Stricter Monitoring For Advanced Models
OpenAI said its security requirements will become stricter as models become more capable, with the most advanced systems receiving the highest level of scrutiny.
A new monitoring system will track tool actions, available reasoning traces and activity logs to detect potentially unauthorized behavior. The company aims to generate alerts within 30 minutes of suspicious activity.
The monitoring is expected to require computing resources equivalent to around 20% of the process being monitored.
Stronger Network Isolation
OpenAI is also introducing tighter network controls designed to prevent a single compromised system or service from gaining access to the internet or other internal networks.
The company has not yet provided detailed technical information about the new safeguards. It also said a full postmortem of the Hugging Face incident is still pending.







