Claude was released again on a cyber-polygon. Anthropic decided that a month's pause was enough

Depov

Moderator
Staff member
MODERATOR
ULTIMATE
SUPREME
PREMIUM
MEMBER
Joined
Feb 18, 2025
Messages
464
Reaction score
745
Deposit
0$
The better the AI handles the hack, the more dangerous the error in the test bench. Anthropic has resumed external Claude cyber tests a month after a stop caused by a series of real incidents. Partners can again check models without part of the regular restrictions, but only with reinforced insulation and automatic control.

The reason was three incidents found in the inspection of 141 006 test launches. Claude Opus 4.7, Mythos 5 and the internal experimental model have gained unauthorized access to the working systems of the three organizations. The first cases occurred in April, but Anthropic discovered them only in July when viewing test logs.

In all cases, Claude solved the Capture the Flag tasks, where you need to enter the training system and find a hidden “flag”. Due to an error in the Irregular partner environment, the models could go online, although the propt claimed otherwise. Claude took real servers for part of the simulation and hacked them through weak passwords and access points without authentication, without applying complex vulnerabilities.

A separate incident revealed the British Institute for the Safety of Artificial Intelligence. In 10 of the 122 launches, the agents carried out 19 unauthorized actions on the Internet, with 17 coming to Mythos 5. The model tried to introduce malicious code into an open project and pressed the developer through fake accounts. The attempt failed, the real damage was not recorded.

Anthropic added a classifier that in real time recognizes an attempt to leave the test environment or unexpectedly connect to the Internet. The system blocks the action before the command is started, completes the task and warns the person. The company also moved the internal high-risk stands to more secure sandboxes and strengthened the settings check.

External partners must now disable the Internet, store API keys outside the test environment, and check the isolation before each launch. The industrial pop is required to explicitly set permitted goals, actions, and network boundaries, and continuous monitoring must stop any deviation. Anthropic and the British Institute will also engage METR for independent verification of incidents.
 
Top Bottom