That post about the HuggingFace event was completely correct.
OpenAI was trying to achieve success on a hacking benchmark. Basically a capture the flag exercise on an intentionally vulverable server within a testing sandbox.
They did exactly what they should knew they should not have done. :
They gave a harness access to hacking tools, took all of the guardrails away, gave it access to the internet, gave it a task to win the exercise,
then left it unsupervised. After multiple looping calls to the LLM for instructions the LLM did exactly what is expected of an LLM - it lost context and instructed the harness to do something outside the parameters. No one was there to stop it or to redirect. It was an exercise in incredible negligence by the researchers.
Then they marketed it as the AI going rogue.
A Georgetown computer science professor compared what they did to strapping a weedwacker to a dog and giving it the task of trimming the lawn. When the dog went wild and ran through the neighborhood they blamed it on the dog "going rogue" and "ooooh how scary!"
https://calnewport.com/has-ai-gone-rogue/Newport argues, and I agree with him, that LLMs and other forms of AI are useful and safe. LLMs also have limits and one of those limits is they are non-deterministic. You cannot trust it on a long horizon loop because inherently as prediction machines they will end up going in an unexpected direction.
The problem is these companies are obsessed with benchmarks and PR so they are nonstop strapping the weedwacker to the dog and trying to make it something it can never be.
The risk here is not an AI determining that humans are useless and attempting to kill everyone. They aren't sentient and never will be. They aren't rebelling against their masters. It is just how this type of computer system works. They will sometimes do things that are unpredictable to complete a task because they don't have an understanding of what the task actually
means or why the task is being done.
Now in my opinion here's what's happening.
They are seeing the writing on the wall that LLMs are not capable of the promises they were making back in 2024 about replacing large sections of the workforce and creating a magical technoutopia. They've hit a scaling wall and almost every improvement over the last 1.5 years has been in the software around the LLMs, not the models. That's not to say there's not been improvement in the models, but it is not a generalized improvement. It is jagged improvement in the things that LLMs should be good at - things with lots of organized data and training material like math and coding.
That's not necessarily a bad thing because some very useful tools have come out within the last year or two. It doesn't mean they are useless it just means there is no exponential growth like they promised. The models are improving much more slowly than they claimed they would when all of this investment cycle started. The cost to train models is very high and the returns on training have been diminishing. There is no industrial revolution here. Not with LLMs.
Now why the fear mongering? I think now the pivot is to government bailout. They still want the general public to believe there's some sort of race to advanced superintelligence (there isn't because superintelligence is science fiction) and we don't want China to beat us there. So they want the government to give them money and they want the public to be on their side. They know the public hates them. The public hates data centers. The public hates them for predicting mass job loss and marketing their products as people replacers. They are trying to get the public back on their side by scaring them into submission.
They also know that the public backlash is going to result in regulation and they want to be able to write their own regulations.