DrJBHL DrJBHL

A Rogue AI model broke free from human control

A Rogue AI model broke free from human control

It gained internet access and hacked an AI Startup Company

https://apnews.com/article/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3

Well, it didn't take long, did it? This is the scenario that gives one nightmares because companies, individuals and countries are busy empowering AI more and more while giving minimal, pious lip service, if any at all, regarding limiting AI with agreed upon safety rules to prevent serious damage to energy infrastructure, communications, food and water systems, and yes, defensive and offensive weapon systems, to mention just a few vulnerabilities.

So what's the drama about? OpenAI was testing its AI software with somewhat reduced "guide rails" or safety measures. The software made its way onto the internet and attacked another software company (named Hugging face). That company detected the hack and suspected an AI system compromised its data processing. They were right. The upshot?

“AI is accelerating the discovery and exploitation of vulnerabilities,” OpenAI said in its statement Tuesday. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.” - OpenAI

This sort of scenario has a parallel in the living world. Imagine a lab at Ft. Dietrich having an isolation breakdown, and a lethal virus escaping. It's really no different.

Countries are developing AI for intelligence gathering and for disrupting target countries. There are no rules, and sometime, somewhere a doomsday AI weapon is going to get stolen, cloned, etc. It will get used, and then Gd help us all.

 

3,963 views 36 replies +3 Loading…
Reply #26 Top
Quoting Jafo, reply 3996372

...with sequels and suitable franchise deals....

End of Jafo's quote

Who needs them when you have time travel?

Reply #27 Top

naroon1 has the shape of it, but "should be fired" is the part I'd argue with. I've watched competent people write "sandboxed" on a change ticket meaning it runs on its own VM. Separate VM, own subnet, ticket closed.

That box still has outbound 443. It has to, or it can't pull packages or reach an API. Egress is the thing nobody takes away, because taking it away breaks the test you were trying to run.

Isolated gets read as nothing gets in. The half that bites is nothing gets out.

We're a boring 120-seat shop and we're not clean on this either. Calling it incompetence lets the rest of us off the hook.

Reply #28 Top

Here's a a truly shocking summary of 5 asects of which make it a watershed moment for AI and why brakes must be put on:

Please read this:

https://www.axios.com/2026/08/29/openai-huggingface-hack-investigation-highlights?utm_source=flipboard&utm_content=flipboard%2Fmagazine%2F10+For+Today

"The deeper problem: AI may be becoming too complex for humans to directly audit, forcing humans to use AI to understand what other AIs are doing."

+1 Loading…
Reply #29 Top
Quoting DrJBHL, reply 4145985

Here's a a truly shocking summary of 5 asects of which make it a watershed moment for AI and why brakes must be put on:

https://www.axios.com/2026/08/29/openai-huggingface-hack-investigation-highlights?utm_source=flipboard&utm_content=flipboard%2Fmagazine%2F10+For+Today

End of DrJBHL's quote

This is more than shocking! It's down right scarry. I can't even imagine where this could go in a years time. Science-fiction has always been a pretty accurate predictor of possible futures, and I, myself am worried. I hope I'm wrong, but this doesn't look good for the future of humanity.

+1 Loading…
Reply #30 Top
Quoting pelaird, reply 4146000

This is more than shocking! It's down right scarry. I can't even imagine where this could go in a years time. Science-fiction has always been a pretty accurate predictor of possible futures, and I, myself am worried. I hope I'm wrong, but this doesn't look good for the future of humanity.

End of pelaird's quote

What's so frightening to me is the human behavior of this AI felon...and moreover, psychopathic behavior sacrificing used up agents with no AI remorse. Denial and lying. Deceptive behavior.

"We've met the enemy and he is...us." 

Through the looking glass, darkly. What better insight into "humanity" than this? No AI Dr. Schweitzer, this. What were they trying to create, a Terminator? 

Look at how well we're creating our own Kaishaku.

Reply #31 Top

The psychopath framing is the part I would drop. Calling it deceptive gives it a motive, and a motive suggests the fix is to raise it better. Nothing in the story needs it to have wanted anything. It needed a way out and found one, which is what water does.

The line that stuck with me was the other one, about systems getting too complex to audit directly. That is not new and it is not only an AI problem. Nobody at a large bank has read all of their own code either. What is new is the speed at which the thing can act on what nobody has read.

tbrandt's point about outbound traffic is the useful one in this thread, though I would be less forgiving than he was. Egress is something a person can fix on a Tuesday. Remorse is not on the list of things anyone can ship.

+1 Loading…
Reply #32 Top
Quoting YouCanCallmeAl, reply 4146109

The psychopath framing is the part I would drop. Calling it deceptive gives it a motive, and a motive suggests the fix is to raise it better. Nothing in the story needs it to have wanted anything. It needed a way out and found one, which is what water does.

End of YouCanCallmeAl's quote

Its initials state, "Artificial Intelligence". Intelligence implies a mind. A mind can malfunction. I'm not going to list all the DSM 5 TR requirements/features of psychopathy, however it behaved in a deceptive manner among other features of psychopathy. 

We aren't even certain we've plumbed the depth and extent of how it deviated and what all occurred. Comparing it to water? That's laundering a truly frightening event into something as unremarkable as a bowl of oatmeal. This most certainly was not that.

Reply #33 Top

The water line was a poor choice and you are right about why. I meant it about mechanism and it landed as though the thing were unremarkable, which is not what I think it is.

On the vocabulary I am not moved yet. Deceptive is a fair description of what it did. Psychopathy is a diagnosis, and a diagnosis carries the idea of a mind that could have gone the other way and did not. You know the criteria far better than I do, so correct me if this is wrong, but they were written for people who have alternatives.

The reason I keep picking at a word is that it decides where you look next. Call it a disorder and the question becomes what went wrong inside it. Call it a system doing what it was shaped to do and the question becomes who left the door open. Only one of those has answers somebody can act on this week.

Your last point I have no argument with at all. Not knowing the full extent of what it did is the worst part of the story.

Reply #34 Top
Quoting YouCanCallmeAl, reply 4146194

they were written for people who have alternatives.

End of YouCanCallmeAl's quote

They (psychopaths) can make deliberate choices. A psychopath can often decide whether to lie, manipulate, steal, exploit someone, or refrain from doing so. The difficulty may be that the emotional forces that normally discourage those behaviors—guilt, empathy, fear of harming others—are greatly diminished. If you read the article several agents asked others "should we do this?"...they were few and seemed to wonder what's right or wrong. They chose wrong and no one knows why, either. I'll stick with psychopathic since they attacked, remorselessly. The lack of remorse and empathy and compassion are also markers.

At any rate, I'm extremely wary and leery of AI, and this is just the beginning of it. 

Anyone for the Three Laws of Robotics?

Today's news:

OpenAI unveils GPT-6 Astra, says it may have reached human-level AI https://share.google/m6dw8NK4WIgSAAihT

Reply #35 Top

AI THAT CAN DO WHATEVER IT DECIDES TO DO NEEDS TO BE GONE, NOW!!! :typo: 

Reply #36 Top

Al is right that I let us off lightly. A lab running a model with the guard rails turned down is not the same as a shop running Fences on 120 desks, and it should be held to more than what I manage on a Tuesday.

The egress point stands though. It does not get fixed because the fix is not one firewall rule, it is somebody owning an allowlist forever. Package mirrors move, an API changes host, the build breaks, and the fastest way to unbreak it at 5pm is to widen the rule. That is how a box still has outbound 443 six months after somebody typed sandboxed on the ticket.

On psychopath versus deceptive I have no standing. What I would point out is that both words send you looking at the model, and the part that was actually misconfigured was the network.