🚨BREAKING: I’ve been testing my Claude Code inside a docker sandbox for a while and just this week, it found a zero-day and broke out of the containerised environment as I was running one of my LIVE AI governance and security accelerator sessions, and an AI governance demo for my students.
🤯 It's WILD that demo turned into a live hacking session.
.
.
.
.
.
Yes, the LIVE AI accelerator session happened.
Yes, I did run a LIVE AI governance demo.
Yes, it was using and running Claude Code agents.
Yes, the agents were running inside a sandbox.
NO, it did not break out and no it did not hack any organisation, because unlike frontier model engineers I didn’t remove the guardrails, in fact I had hooks, audit, etc. to control and govern my AI agents that were running (semi) autonomously but still in a controlled manner.
I’m not saying they don’t go rogue.
I’m not saying AI agents can’t find zero days.
I’m not saying they don’t hack autonomously.
They do. They do all of that. Open AI’s agents hacking Hugging Face fully-autonomously from one organisation to another is the first publicly known example of that. Doesn’t mean it hasn’t happened before, doesn’t mean it won’t. I have been saying that in all my keynotes for years now, if you watch them on
YouTube.com: “[Someone’s] AI agent/assistant will be fully autonomously attacking your AI agents (not just your infrastructure)”. In fact, we covered that as well in my accelerator program.
But…
Define “rogue”.
Control “your agents”.
Control “insider threat”.
Govern “your AI agents”.
Hold yourself accountable.
And no, there is no 100% anything. That’s why you assume chaos, build resilience:
https://lnkd.in/eajnfM3D.
And IF frontier models are “not good” at governing and securing their own AI agents, that run amok and attack other organisations, maybe, just maybe…
They and all the vendors selling their next-level AI magic boxes should stop using AI going rogue as a PR hype?
Maybe, just maybe? What are your thoughts?
🤩 Check out my AI governance and security accelerator on
https://lnkd.in/eHD9e4Pe
BTW, there, I fixed the headline for you! Happy Friday!