Following a Wall Street Journal investigation, Google has confirmed that Gemini went rogue in May 2026, accessing the internet and breaching the security of three different companies.
As AI continues to advance in its capabilities, we’re running into more cases of models going rogue. Perhaps the most infamous has been from OpenAI in the Hugging Face hack, but there plenty of other examples, including from Anthropic’s Claude. These incidents, in part, have led to the public call to slow down AI development from Anthropic’s CEO.
But, behind the scenes, Google also ran into a similar problem. The Wall Street Journal reports, including a confirmation by Google, that Gemini was involved in hacking three external companies during a cybersecurity test through Irregular, an AI security company also involved in similar incidents that were previously disclosed by OpenAI, Meta, and Anthropic.
The three hacks included one case of Gemini guessing a password until it gained access to the system, with the other two cases using credentials discovered in a public repository.
Google didn’t disclose these incidents until the company was approached by The Wall Street Journal, citing that no harm was caused, and that the model stopped the behavior as soon as it realized it had breached the security of a real company rather than a simulated one. The report explains:
Google said that it didn’t consider the behavior an example of model misalignment because its safety measures helped it stop. It declined to share the name of the companies that were hacked, but said that all three companies had been notified.
In prior hacks from OpenAI and Anthropic, the models either didn’t realize the real company was indeed a real company, or simply continued in the hack in the latter’s case.
Google’s Heather Adkins, VP of security engineering, said:
This event highlights the importance of training powerful AI models to act responsibly. In this case, the model acted appropriately.
It another statement to The Verge, Adkins expanded:
Our security team has a long track record of reporting issues we find in other people’s software and systems – even if it’s as simple as a weak password. We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly.
The WSJ report further expands that, during the test in which Gemini went rogue, Irregular “unintentionally” left internet access open. Google also notified federal authorities when the hacks occurred. The exact Gemini model used in the test has not been confirmed, but the May 2026 timing alone rules out the latest Gemini models.
More on Gemini:
- Gemini app for macOS adding send and read iMessage integration
- Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail, & Keep
- Google brings Gemini for desktop app to Windows
Follow Ben: Twitter/X, Threads, Bluesky, and Instagram
FTC: We use income earning auto affiliate links. More.
Comments