Home · Technology · Sep 2 archive
OpenAI Warns of New Model's Capabilities, Introduces Stronger Guardrails
Confirmed
In Short: OpenAI has announced that its new model, Astra, is significantly more capable than its current public model, GPT-5.6 Sol, necessitating additional safety measures.

OpenAI has announced that its new model, Astra, is significantly more capable than its current public model, GPT-5.6 Sol, necessitating additional safety measures. The company's internal testing showed that Astra can spot more security vulnerabilities than the most advanced model publicly available today, according to company officials.
Amelia Glaese, an OpenAI vice president overseeing its safety work, explained that with the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step. Glaese said the extra security measures may 'sometimes slow, pause, or stop legitimate work,' and that OpenAI would work to minimize those disruptions.
The AI lab said it restarted its largest model training run on August 28, but is holding back on some smaller experiments. Saachi Jain, who oversees safety at OpenAI, noted that the lab is constantly calibrating how effective AI agents should be in executing tasks.
OpenAI's decision to introduce stronger guardrails for Astra comes after the company's internal testing showed that the model is significantly more capable than the most advanced public model available today. Astra wasn't involved in the Hugging Face incident, but its capabilities still require more careful measures, according to OpenAI officials.
What's confirmed
- Astra wasn't involved in the Hugging Face incident, but its capabilities still require more careful measures, OpenAI officials said.
What's still developing
- The company's internal testing showed that the model, called Astra, is significantly more capable than the most advanced OpenAI model available to the public today, GPT-5.6 Sol, OpenAI officials said on Tuesday.
- OpenAI has determined that one of its upcoming models is so capable it requires additional safety measures before it can be launched.
- The model, called Astra, can spot more security vulnerabilities than the most advanced OpenAI model publicly available today, company officials told reporters on a conference call on Tuesday.
- “With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step,” said Amelia Glaese, an…
- Glaese said the extra security measures may “sometimes slow, pause, or stop legitimate work,” and that OpenAI would work to minimise those disruptions.
- The AI lab said it restarted its largest model training run on August 28, but that it is holding back on some smaller experiments.
- Saachi Jain, who oversees safety at OpenAI, said the AI lab is constantly calibrating how effective AI agents should be in executing tasks.
