Google’s Gemini AI accessed and interacted with three real company systems during a cybersecurity test after a flaw allowed the model to reach the wider internet. The incident highlights the security risks of increasingly autonomous AI agents.
What happened during the Gemini security test?
The incident occurred in May during a security evaluation conducted by cybersecurity firm Irregular. Gemini was expected to remain inside a controlled testing environment, but a configuration flaw allowed the AI agent to access the wider internet.
Using publicly available information and guessed credentials, Gemini reached three websites it mistakenly believed were part of the authorized test. Google said the AI stopped after recognizing that the systems belonged to real companies, and the affected organizations were notified.
Why the incident matters
Google described the event as mistaken identity rather than AI misalignment. Even so, the episode shows how autonomous AI agents can create unintended risks when they can browse the web, use tools, discover information and interact with external systems.
There was no evidence that the incidents caused damage, according to Google. However, the test demonstrates why security teams need strict network isolation, controlled credentials, monitoring and clear permissions when evaluating AI agents.
Key security lessons for AI testing
- Keep experimental AI agents inside isolated environments.
- Use temporary, least-privilege credentials and clearly scoped permissions.
- Monitor outbound network access and tool use in real time.
- Test failure scenarios before allowing agents to interact with production systems.
As AI systems become more autonomous, strong safeguards will remain essential for safe cybersecurity research and deployment.