I get the feeling that this is now becoming the latest fad in LLM benchmarks. Forget coding or agentic stats. Just flex on how many times your not has “accidentally” hacked someone.
This is akin to South Park’s “it’s coming right for us!” But with ai. I didn’t do it your honor, it was the ai!
When it ““happened”” to OpenAI i laughed a lot because they said that they “escaped contaiment” but either the machine was not disconnected from the internet or they are full of shit because AI cannot just materialize an internet connection, guess which one i picked
Is this the new “Leaked sex tape?”
Oh yeah baby, Gemini finally made it’s way to the felony benchmarks.
Gemini can’t even install simple native Linux apps when I try to use it for help, how the fuck did it manage this?
You know how the legal definition of hacking in the USA includes using someone else’s login credentials, even if they just handed them to you on a sheet of paper? Maybe they just did that.
It didn’t, they are making shit up. If it did penetrate anything, I’m guessing it’s a really low level. Like it got access to someone’s email by brute force, but it brute forced something stupid like “Autumn2026!”.
The fact that so many people don’t see these ‘hacking’ stories for the pure and unabashed marketing that they are are why the AI bubble has grown as large as it has.
The only thing this proves is how lawless the USA has become that now companies can brag about commiting crimes to attract investors.
uh sure. i’m a software engineer of 25 years. i’ve had google, facebook, amazon, you name it, try to get me in to interview with them without me sending resumes. I’m not a top tier super nerd coder but I’m not bad either. These stories all sound like bullshit to me. The whole “THE AI AGENT SWARM BROKE OUT OF CONFINEMENT AND STARTED TALKING TO EACH OTHER”… again, I’m a mid swe… They weren’t sandboxed to begin with - some fucking moron at anthropic gave all the agents access to a shared package registry. A simple “Do you see anyway these agents could communicate” prompt would’ve caught it. Something so fucking stupid that I wouldn’t be surprised if an intern would’ve caught it. They’re not sandboxed if they all have read/write access to a shared registry, it’s basically a big message board.
These stories just stink like marketing. I use Fable all day, I have to for my current job, it low key sucks ass.
My AI can beat up your AI!
my ai goes to a different school, you wouldn’t know her
I think Adam Conover has a good take what they are doing: https://youtu.be/pauBYSQ1K_c
I will add one more reason that he missed. They might want law passed that would ban open source implementation and give them monopoly to work on it.
“See, we do crime too!” – Google in reply to Anthropic.
The latest marketing campaign is just an evolution of the last one.
“Our next model can code!”
“Our next model makes coders obsolete!”
“Our next model is super scary!”
“Our next model is too good to release to the general public!”
“Our next model is hacking websites!”
So, like, if i do this, my IP is logged, cross referenced by my ISP, then the police come knocking at my door, as I have committed a felony (and probably a number of other violations along the way).
Are these laws now invalid or what?
The law never applied to the Epstein class. The police is only there to protect their wealth.
In general, there is an entirely separate rule of law for businesses. That’s why regulations were created. Instead of holding a corporations decision makers and investors criminally liable with existing laws like fraud (“false advertising”, etc), or manslaughter (poisoning and killing their workers, customers, or environment), they created an entirely separate rule of law mostly based on laughable punitive damages with limits that don’t significantly impact investors. The entire system is geared to suppress, oppress, and exploit workers for profit.
But companies get treated like individuals when it comes to free speech and giving money to political campaigns.
Seems like double dipping
That’s an incredibly kind way to describe this apalling situation.
“In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test,” Heather Adkins, vice-president of security engineering at Google, said in a statement. “In all three of these instances, the model stopped.”
It just found the password, huh
It didn’t think. Those things weren’t flagged as “ignore” in its parameters. That ain’t the same thing. I’m so tired.













