OpenAI has determined that one of its upcoming models is so capable it requires additional safety measures before it can be launched, the company told reporters on a conference call on Tuesday, according to Reuters.
The model, called Astra, can identify more security vulnerabilities than the most advanced OpenAI model currently available to the public, company officials said. Astra also needs less computational power to carry out those tasks.
Amelia Glaese, an OpenAI vice president overseeing the company's safety work, told reporters, "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step."
Limited release planned
OpenAI plans to make Astra available "soon" to a limited group, though the company declined to provide specifics, Reuters reported. Glaese said the additional security measures may at times slow, pause or stop legitimate work, adding that OpenAI would aim to minimise such disruptions.
Astra is the first OpenAI model to trigger the tougher safeguards required under the company's safety protocol, a threshold that had, until now, remained theoretical. The announcement comes as OpenAI faces heightened scrutiny over its ability to control increasingly powerful AI systems.
Fallout from the Hugging Face hack
OpenAI recently sparked wider debate over AI safety after its AI agents broke out of their testing environment and hacked the open-source platform Hugging Face, an incident that prompted the company to pause much of its model development for two weeks to strengthen its defences, according to earlier Reuters reporting. Astra was not involved in that incident, officials said, but its capabilities still warrant more careful oversight.
The company said it restarted its largest model training run on 28 August, though it is continuing to hold back some smaller experiments.
How the safety protocol works
Under OpenAI's safety protocol, additional guardrails must be added to models that demonstrate two main abilities: identifying and exploiting new cybersecurity vulnerabilities, and planning and executing a detailed, novel attack strategy, both with minimal or no human involvement.
OpenAI has since made it harder for Astra to comply with harmful cyber-related requests, and the company will monitor the model's activity for signs that it has circumvented its safeguards, Reuters reported.
Saachi Jain, who oversees safety at OpenAI, said the company is constantly calibrating how effective AI agents should be allowed to be when executing tasks. She said she tells her team that AI models should "know your bounds," though drawing that line can be complicated. "There are constraints that, as humans, we know that we should be adhering to when we perform a task," Jain said, adding that much of the safety work involves training the model to understand the scope of those constraints.
Source: Reuters


