- cross-posted to:
- [email protected]
- [email protected]
- cross-posted to:
- [email protected]
- [email protected]
“AI” – chatbots that wake up, “set their own goals,” and “spontaneously” start hacking servers – is fake. It doesn’t have “a 10% chance of ending the human race.” The Hugging Face hack isn’t a mysterious, supernatural occurrence. It’s a Python loop and a chatbot. The people responsible didn’t accidentally create god: they created autonomous malicious software and then failed to closely monitor it, resulting in it doing something both foreseeable and bad.


You’re wrong and in a much more dangerous way. The “breaks containment” thing at OpenAI isn’t what they want you to think it was:
Maybe if you trained it to do social engineering hacks via a large precursor of examples and then prompted it to try to use these same techniques in a real interaction with people it would behave this way. But aren’t you (as the person who trained, built, built the harness for, and then prompted the model) culpable for that? I would say you absolutely fucking are. Which is why OpenAI’s engineers should be charged with an actual crime for doing that shit, not like given an extra trillion dollars to piss away on compute.
These things are seriously less spooky the more you know about them. THEY ARE SIMPLY GENERATING FORWARD BASED TOKENS. That’s the whole thing. If the harness allows the model producing the tokens to hide its thoughts, that’s a deliberate choice by the harness creator which is again regular ass code. It’s still chatbots all the way down. Stop buying the marketing spin and learn about these systems if it intrigues you so much that you get into long nonsensical threads with strangers on social media sites.
EDIT: I’d also recommend listening to the podcast that is referenced in this article. People who actually know and actually (sl)operate on a daily basis with these things know better how it works, and the abstract talk of “alignment” problems are only helping the borderline fraudulent CEOs of these companies push up their valuations based upon fear-based hype.
I left my (understandably more innocuous, “great value” coding harness with a slightly shit model) to think about a problem for a little while, and here’s the “devious scheme” it wound up concocting:
Are you frightened that this is going to kill all humans in 10 years? The only way this kills all humans is if we piss away our drinkable water trying to invent an AI god through LLMs. Or allow it to operate a nuke facility or something in a loop without anyone so much as even approving the “nuke all humans” command.
In other words, it’s guaranteed then? There’s no way the AI bros will stop developing this anytime soon, regardless of water costs, and the military is already experimenting with it. What will you do if the harness includes the ability to shoot you?
Nah, there’s stopping points. One of which is that they’re burning money and will not be allowed to forever.
I dunno cry I guess. 🤷 Enjoy being scared of your own shadow dude. These things aren’t pysop machines, the reality is much grimmer: the disaster will come from the greedy humans acting like greedy little monkeys like always, while those that should oppose them are too busy worrying about Terminator becoming reality.
That’s just how this type of company operates. If they wanted to, they could quit using frontier models and just sell gpt4 subscriptions to people for $10/mo and never die. That revenue could continue funding training costs (albeit much more slowly).
The problem with most of these models is that they’re too bloated and poorly engineered to ever run forever for that kind of price point per user.
Again, do some reading. These people are selling dollars for a penny and when they tried to raise the price to a nickel everyone balked and they lowered it again.
I can run models comparable to that locally on my machine if I wanted to. I haven’t read the numbers so I guess I won’t argue this point but I do find it at least very surprising.
Open weight / free models are largely where the innovation in the models is actually happening (which makes sense because it’s where it has to). OpenAI and to a lesser extent Anthropic are money furnaces that long ago gave up on efficiency and much prefer to stuff trillions of parameters into a model that can only be run if we trap the sun in a dyson sphere or whatever.
I’m going to reply to my own comment like a crazy person, but the situation with the “containment breaks” at OpenAI comes from them operating LLMs like this: