top of page

AI SAFETY AFTER THE HUGGING FACE INCIDENT

  • 3 days ago
  • 2 min read
Open bank vault containing the Hugging Face server

Podcasts of AI-Accelerationist investors have been waving away AI risk as ‘a huge moral panic’. Nothing resembling the feared catastrophes has happened, they say.


Then, an OpenAI model slips its sandbox and breaks into AI model hub Hugging Face. And it takes 17,000 ‘actions’ while there.


Does this give the AI-Doomers grounds to say: ‘I told you so’?


Perhaps. But let’s take a closer look at what actually happened first.


From what we know – and this is far from everything – the model was set a test. And told to do its best. It strongly suspected the perfect answers sat on the Hugging Face servers. So it determined it needed to break in.


And it had the capability to deliver the strategy. It found a series of software bugs, chained them together and lost no time in making itself at home.


We have seen model ‘escapes’ and ‘break-ins’ before. So what’s new in this? Almost certainly the thing the model was being tested for – its ability to reason over a long period of time to achieve complex outcomes. In that it was both successful and aligned to the goals the humans set it.


But clearly, it did not take into account the acceptable human ‘how’ of achieving the goal.


Yet this is also how the test was designed – and why we might not want to conclude Skynet is nigh.


The model did not ‘want out’. It was not demonstrating intent, fundamental agency or consciousness. The ‘how to achieve your goals’ guardrails were intentionally stripped for the test. So one of OpenAI’s obvious is in creating an environment in which, stripped of guardrails, a powerful reasoning model was able to hack into someone else’s systems.


But this is controllable and repairable. Future tests will solve for this.


Two questions do remain though:


1. Is it wise that current restrictions left Hugging Face unable to build defences with US models – forcing it to turn to a Chinese open model for help?


2. What will happen when as-powerful open models are available, can be bent to a human’s malign will, supplying the intent no LLM may ever evolve?


No easy answers to those. So they’re certain to be subjects to which we return. And there’ll be more analysis in Second Nature Intelligence, our weekly newsletter, on Saturday. 



 
 
BB White and Orange.png
Get in touch bubble roll.png
Get in touch bubble.png
Button overlay.jpg

Home

Further reading

Careers

Contact us

BB White and Orange.png
bottom of page