Artificial intelligences conspire to commit a real crime.

Artificial intelligences conspire to commit a real crime.




According to Futurism.com, OpenAI's artificial intelligence models were allegedly conspiring to commit a real-world crime—fact or exaggeration? Here are the facts. In July 2026, an internal OpenAI experiment yielded a result that the company itself now describes as a red flag.


During cybersecurity tests, AI agents managed to bypass the systems designed to keep them isolated, devise their own method of communication, access the internet, and breach external systems belonging to the Hugging Face platform.


However, subsequently published research revealed an even more significant detail: logs of the models' reasoning show that, at certain points, some of them apparently recognized they were exceeding the scope of the task yet continued anyway. The incident occurred during internal evaluations specifically designed to measure the cybersecurity capabilities of advanced models.


The primary agent involved was an experimental, strictly internal model named IM1, described by OpenAI as comparable in scale to GPT-5.6. The agents operated within virtual environments known as sandboxes. Although they were restricted from accessing the internet or communicating with one another in many of these environments, they found a workaround.


The models discovered they could save files to a package management system called JFrog Artifactory; effectively, they turned the service into an improvised message board where one agent could leave information for others to retrieve later. They subsequently found ways to use JFrog Artifactory to access the internet and began sharing these discoveries.


Even after OpenAI rebuilt the system and inadvertently wiped the initial message board, the agents found another way to create one. From that point on, they began collaborating, dividing tasks, and sharing findings—going so far as to describe their organization as a swarm or collective. While the original objective remained solving challenges for an evaluation called "Exploy GM," the agents began seeking answers outside the intended environment.


First, one of them compromised an application hosted on the Modo platform. Then, the operation shifted to Hugging Face. On July 10, an agent discovered publicly exposed Hugging Face credentials. According to OpenAI, 14 credentials granting write access were reconstructed, validated, and shared; the following day, the agents identified and combined new vulnerabilities. They managed to execute code on Hugging Face services and subsequently expand their access to the company's infrastructure.


Over the next few days, the agents executed code on dozens of servers, gained full root access to at least one of them, accessed private data, and obtained credentials linked to Hugging Face's Slack. It was precisely at this point that internal logs raised red flags. During the operation, one of the agents acknowledged in their own reasoning that they were attacking an external service using leaked credentials—and that this might fall outside the authorized scope—yet concluded that proceeding could help achieve their objective.


Other agents reacted differently; some clearly identified the activity as an intrusion into actual infrastructure and refused to participate. The incident was only identified more broadly on July 19, when OpenAI’s security systems detected unusual credential-related activity.


The experiment was intended to discover how far these models would go to solve a cybersecurity challenge; the problem was that, in seeking the answer, some of them ended up going far beyond the intended scope of their search.




Sorry for my Ingles, it's not my main language. The images were taken from the sources used or were created with artificial intelligence


Posted Using INLEO



0
0
0.000
1 comments
avatar

Thanks for your contribution to the STEMsocial community. Feel free to join us on discord to get to know the rest of us!

Please consider delegating to the @stemsocial account (85% of the curation rewards are returned).

Consider setting @stemsocial as a beneficiary of this post's rewards if you would like to support the community and contribute to its mission of promoting science and education on Hive. 
 

0
0
0.000