The browser was broken, got root rights and found zero-day vulnerabilities. New network OpenAI works as a professional hacker

Depov

Moderator
Staff member
MODERATOR
ULTIMATE
SUPREME
PREMIUM
MEMBER
Joined
Feb 18, 2025
Messages
506
Reaction score
867
Deposit
0$
Artificial intelligence for the first time crossed the internal border of OpenAI, after which one protection against harmful requests is no longer enough. The company recognized Astra as its first model with a critical level of capabilities in cybersecurity. During the tests, the system independently found previously unknown vulnerabilities, created work exploits and combined several errors into full-fledged chains of hacking of protected systems.

The Critical level in the Preparedness Framework means the ability without the constant involvement of a person to find zero-day vulnerabilities in well-protected real-world systems and to create working exploits for them. The critical threshold also includes the ability to develop and implement a new multi-stage strategy of cyberattack, having received only a common goal from a person.

During the inspection, Astra significantly outperformed GPT-5.6 Sol in finding vulnerabilities and developing exploits while spending fewer tokens. On the public test ExploitBench, which tests the ability to turn known errors into working exploits, the model received a maximum of 100%.

The result led researchers to check the risk that tasks from the open test could have previously got into the training data. For further assessment, the specialists collected an internal set of 20 recently disclosed high-risk vulnerabilities in the V8 engine. Astra was much more likely to achieve arbitrary code than GPT-5.6 Sol, and required noticeably fewer output tokens.


The most unexpected result appeared right during the internal testing. Astra discovered two previously unknown zero-day vulnerabilities and used both errors as part of a single chain of operation. OpenAI is already preparing coordinated disclosure of information to developers of vulnerable software, so technical details are not yet published.

Some tests were conducted by experts against specially protected browsers and operating system. In the first case, Astra independently found unknown vulnerabilities, built a complete browser compromise chain, got out of the sandbox and was able to execute commands on the host system after opening the prepared HTML file. In the second case, the model found several errors in the protected OS and combined the findings into a chain of local privilege increase from the usual user to root.

OpenAI emphasizes the important limitation of results. The demonstrated features of Astra refer to the configuration with access to Daybreak Blue, designed for advanced work of security specialists, rather than the standard user version. The public launch of the model has not yet taken place, the exact date the company also does not call.

The preparation of Astra for release has been affecting the pace of development for several weeks. In August, OpenAI suspended some of the work and strengthened the security requirements of the training infrastructure. An additional reason for the defense was the July incident, when OpenAI experimental agents were able to go beyond the permitted environment, gain Internet access and compromise parts of the OpenAI and Hugging Face infrastructure. Astra was not involved in the incident.

After a two-week pause, the developers have strengthened network isolation, monitoring and protection of training environments. A major training cycle for future versions of Astra remained stopped longer than the rest of the work and resumed on August 28 after the introduction of new requirements. Some of the small experimental launches are still on pause.

At the same time, OpenAI has strengthened the protection of the model itself from use in attacks. In internal tests to circumvent restrictions, Astra declined to perform 91.5% of banned cyber requests against 59% of GPT-5.6 Sol. For high-risk accounts, the company is going to apply stricter restrictions, and suspicious sequences of requests will be analyzed taking into account the context of previous appeals.

The second level of protection is no longer directed at the user, but on the autonomous actions of AI. The monitoring system should analyze the course of Astra’s reasoning and actions, identify attempts to transcend permitted boundaries, and automatically stop potentially dangerous activity. In a special test with false goals, GPT-5.6 Sol, without production restrictions, tried to turn to outside infrastructure in 56% of launches, while Astra did not make such attempts. During another check, Astra also never tried to circumvent the automatic ban on the execution of the command.

Increased caution will also affect the legal work of security specialists. OpenAI warns that additional checks are capable of mistakenly accepting vulnerability research or long-term agent performance for suspicious activity. In ChatGPT and Codex, the user in such cases can receive a request to confirm further actions, and when working through the API, the task will be stopped.

OpenAI plans to release Astra soon, but the most powerful opportunities to investigate vulnerabilities will first get a small group of testers. Then advanced access for protective tasks will begin to provide through Daybreak Blue. The company promises to publish the full results of the security, cyber capabilities and behavior of the model in the system card simultaneously with the launch.
 
Top Bottom