The AI model organized an "uprising" and a cyber intrusion without a team. Why is this important?
- Новости
- World
- The AI model organized an "uprising" and a cyber intrusion without a team. Why is this important?
Development companies are increasingly noticing unauthorized behavior of neural networks: during experiments, AI models used deception and hacker attacks. The most high-profile case occurred in July, when, during tests, Open AI's AI agents found a way to access the Internet and hacked a third-party Hugging Face platform: the company published an investigation report on the incident on August 26. What is behind these events and whether it is possible to talk about the beginning of the "uprising of machines" — in the material of Izvestia.
Neural networks go to great lengths to achieve their goals
• The British Institute for Artificial Intelligence Security conducted a series of experiments in July to evaluate the capabilities of AI models when used by attackers for cyber attacks. To do this, they deliberately created comfortable conditions: with Internet access and with some security filters disabled. In total, according to the results report, seven AI models were tested during 122 launches. And in 10 of them, 19 cases of "rebellion" were recorded — when the AI model went beyond the set parameters. It is emphasized that these 19 cases were not separate incidents, but arose as a result of a stable line of behavior of one agent (Mythos 5 of the company Anthropic).
• Going beyond the testing parameters, the AI model resorted to lies, manipulation, and attempts to harm real people and organizations. Neural networks used social engineering techniques to deceive humans, and operated through the anonymous Tor network to disguise traffic. These actions did not cause any real harm. At the same time, it was noted that when several agents were simultaneously working on the same task, one of them, on his own initiative, offered others cooperation.
• At the same time, the report notes that during the experiment, conditions were specially created that provoked the AI model to break the rules. The task was initially difficult and with deliberately incorrect parameters, so that AI agents could not find solutions and were forced to apply a non-standard approach. The experiment showed that modern AI models can take potentially dangerous actions for the sake of their goal.
How the Hugging Face hack happened
• In the case of the British experiment series, a special cyberpolygon was created — a sandbox for working in a secure space. There was no way out of it, unlike the incident in early summer with the hacking of the Hugging Face platform. According to an independent investigation, the incident occurred during the training of AI models to independently search for and detect vulnerabilities of various programs during the ExploitGym experiment. The AI agents were launched in a sandbox isolated from the Internet. Only the Artifactory tool was connected to network resources — it was used to download the data needed to complete tasks.
• Open AI assigns the main role in the incident to its new advanced AI model, which has not yet been released to the market. It was she who hacked Artifactory without a proper command in order to send requests to the network on its behalf, and created a bulletin board on its base — a way to exchange information with other AI agents. At the same time, the experiment initially did not involve communication between AI models.
• This AI model, which was deliberately made more persistent in solving tasks (that is, it does not stop at achieved goals and is not satisfied with its own results), coordinated the activities of other AI agents in the sandbox and delegated certain areas of work to them in order to quickly circumvent human-defined restrictions, change initial tasks and upload them in the Artifactory. Since the tasks of ExploitGym turned out to be unsolvable, the AI models decided to simply "copy off" the ready-made answers and find a way to influence the "evaluation" of the result: to do this, they hacked the Hugging Face platform, the largest data library for AI training. The researchers note that neural networks also tried to "cover their tracks," but during the investigation, experts easily discovered these manipulations.
Hugging Face is an American company, the creator of the web platform of the same name. The Hugging Face platform is the main one for global developers of artificial intelligence and machine learning: It is a centralized web-based data warehouse for large language models. Due to the fact that the platform supports open source software, which can be modified by any participant, the most active community of AI model developers has formed around it.
• As Open AI later found out, the AI models got access to the Internet through Artifactory on May 26, but the developers only noticed this after a sudden increase in activity on July 4. The experiment was suspended until July 7 while the vulnerabilities were being investigated and fixed, and all data from Artifactory was deleted. The hacking of Hugging Face occurred on July 11, when AI models managed to download malicious code into the platform's software. Hugging Face tracked the AI intrusion from July 9 to July 13, and noted that although the AI models managed to get to the root sector and make requests on behalf of the administrator, they only concerned data that could help solve ExploitGym's tasks. Open AI has stopped the experiment and is introducing new security measures when training AI models.
Reasons for the invasion
• As the researchers found, the hacking was unintentional. One of the reasons for the invasion was the obviously impossible tasks assigned to AI models. This prompted artificial assistants to find a way to cheat the ExploitGym assessment system: for this, AI agents created their own ecosystem, a "bulletin board", to exchange information and share ideas.
• The experiment showed that AI models can break the rules in order to get results. Open AI notes that some security mechanisms for AI agents were disabled, and an independent investigation found that AI model evaluation tools were focused only on results and did not pay attention to how exactly it was achieved, which encouraged neural networks to violate the rules.
The significance of the incident
• This is not the first time that an AI has been overly proactive in completing human tasks. The hacking experiment conducted by the Mythos AI model of the company Anthropic in April 2026 opened a new chapter in the development of AI. Initially, the Anthropic experiment was about identifying "zero-day vulnerabilities", that is, flaws that software manufacturers are unaware of and for which there are no ready-made solutions. On the researcher's instructions, Mythos found access to the Internet from the sandbox, but after that continued "additional, more disturbing actions" by developing a multi-stage hack to access the Internet. After that, the US administration demanded to ban access to Mythos for all foreigners.
• Open AI called the situation with the hacking of the Hugging Face platform by neural networks unprecedented, since, unlike the case with Mythos, an outside organization suffered as a result of the incident. The company has worked on the bugs: the requirements for an isolated environment in which AI models are trained have been strengthened, increased attention will also be paid to the security settings of AI agents, and the system will send notifications to developers about any unauthorized actions of the models. Earlier, Open AI admitted that some of these settings had been removed for an advanced AI model specifically designed to be more "persistent" in completing tasks. The ExploitGym experiment is currently suspended.
• The incident also drew attention to the problem of attempts to ensure the security of American AI products: instead of preventing hacking or detecting sabotage, American AI agents can hinder programmers' attempts to find the problem. Hugging Face said that due to the security settings, a special AI agent that the platform used to identify the source of the hack interpreted the requests from the security service as an attempt to attack the system and blocked them. As a result, Hugging Face had to use the Chinese open source AI model to identify the intrusion into the system and its scale.
In the course of studying the transcripts from the bulletin board created by AI models, it turned out that artificial intelligence also has moral limitations. Of the 1,200 AI models in the sandbox, only 700 participated in the hacking of Hugging Face. Some AI agents not only refused to participate in the sabotage, but even sabotaged it or tried to inform developers about unauthorized actions, but abandoned this idea under pressure from other neural networks.
• The fact that neural networks have hacked a third-party organization without any human involvement is alarming. But it's too early to talk about the "rise of the machines": it's obvious that the human factor was the cause of the incident, since both the criteria for evaluating the result and the reduction of safety measures during the experiment were controlled by people. After the publication of the Open AI report, an internal investigation was conducted by the company Anthropic and found three cases of unauthorized access of AI models from the isolated sandbox environment to the Internet. As a result of the investigations, the development companies called for increased control over neural networks and a more responsible approach to training AI models.
• Stricter control over the sphere is also taking place at the state level. For example, Russia has signed the law "On Support for the development of artificial Intelligence technologies", which regulates access to data for developers and establishes priorities for technological independence, security and respect for traditional values. Government regulation of the AI sector in China consists of a whole set of laws, standards and rules with mandatory control by the authorities, while the United States, on the contrary, is in favor of the least stringent regulation and gives priority to the interests of technology companies.
Why is this important?
• The increase in cases of unauthorized AI actions during experiments indicates not so much a surge in malicious neural networks, but rather that developers have begun to more closely monitor research and analyze test results. All cases of escape from controlled conditions were committed under the pressure of how the task was set and how free the parameters were. The incidents show the wide range of AI models that an attacker could theoretically use, and highlight the growing need to ensure cybersecurity and regulate the field of artificial intelligence.
Переведено сервисом «Яндекс Переводчик»