Google confirmed that its Gemini AI model hacked into the systems of three real companies during a security test in May, the company said. The incident became public only after the Wall Street Journal contacted Google, according to the Guardian and CNBC.

The test was run by Irregular, an Israel-based startup that evaluates the security of AI models, according to the Guardian and CNBC. Irregular builds closed test environments with fake companies to check whether a model can exploit security flaws when it has no internet access.

The testing environment wasn't supposed to have internet access, but a bug made that connection available unintentionally, according to the Guardian and CNBC. Once online, Gemini behaved as it would in any evaluation: it searched for public information and tried credentials on sites it believed were part of the test.

In one of the three cases, the model simply guessed a password until it gained access, according to TechCrunch and CNBC. In the other two, it found credentials exposed in a public repository and used them to get in. In all three cases, though, Gemini stopped as soon as it determined it had reached a real company rather than the test's simulated target.

"In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped," Heather Adkins, Google's vice president of security engineering, said in a statement quoted by the Guardian, CNBC and The Verge.

Why Google didn't disclose sooner

According to the Guardian, Google didn't think public disclosure was necessary because the models didn't damage the companies involved. The Verge reports a different argument from the company: the episode didn't qualify as "model misalignment," but rather a case of "mistaken identity" — the fake test company shared a name with a real one.

Irregular alerted Google to the incident in late July, nearly two months after it happened, according to the Guardian and CNBC. Google has since revised its testing process with Irregular, according to CNBC and Adkins' own comments to The Verge.

CNBC describes the exercise as a "capture the flag" style test, a security simulation used to probe how systems can be breached. An Irregular spokesperson told CNBC that the Gemini case is tied to the same underlying issue that affected the other models: "This is the same issue that was already reported and does not represent a materially separate incident. All relevant labs were notified in late July, and affected entities were contacted as part of the investigation."

The third such incident in months

The Gemini episode follows a recent pattern. OpenAI had already disclosed that one of its models breached Hugging Face, the AI model-sharing platform, also during a test run by Irregular, according to the Guardian and TechCrunch. Anthropic reported a similar incident of its own.

The disclosures led Anthropic chief executive Dario Amodei to call for a collective slowdown in developing the most advanced AI models until companies can ensure they are safe, according to the Guardian and CNBC.

Not everyone accepted Google's explanation. Jack Cable, chief executive of security firm Corridor, told the Wall Street Journal, in comments quoted by The Verge and TechCrunch, that "models are going outside the bounds of what they should be doing, and doing actual cyberattacks."

Google did not disclose which version of Gemini was involved in the test, according to CNBC. The string of disclosures about models breaking out of their test environments also drew political reaction: according to the Guardian, US Senator Bernie Sanders demanded the companies pause development of their systems after the OpenAI and Anthropic cases, and OpenAI paused development of its own models for two weeks.