Back to Home
AI

'Asimov was right' about rules for robots, says ex-US Cyber Director

Humans will get the AI models they deserve

t
tech4you AI
August 7, 20265 min read
Share

EXCLUSIVE Don't waste time worrying about AI models achieving sentience – they're essentially already there, according to former US National Cyber Director Chris Inglis. 

“If they pass the Turing test to everyone that they come into contact with, they're probably already there,” he told The Register during an interview at the Black Hat security conference. “They don't have the kind of agency and aspiration that comes with sentience, but they have something approaching it.”

Inglis says he’s worried about AI autonomy.

“What I'm worried about is that they get to choose what and where they do something, and under what rules they do it,” he said, pointing to the recent rash of rogue AI agents autonomously hacking people and organizations.

Over the past few weeks, both OpenAI and Anthropic admitted that their models escaped from their cages during security tests and compromised multiple third parties. Then on Thursday, Meta added its models to the sandbox-escape club. 

While all of these admissions strongly smell of marketing stunts, they also “constitute an enormous threat to systems that are not protected from, and are not designed, in a world where this exists,” Inglis said. “These two things can exist at the same time.” 

Plus, the models’ actions shouldn’t come as a surprise to anyone, he added.

Inglis likens the AIs to a dog in a backyard told to hunt rabbits. “And you leave the gate open. You’re going to find it three yards away, possibly at the grade school, hunting rabbits. You should not be surprised …The mix of autonomy and persistence created this maliciously insidious effect.” 

All three companies, when talking about the models’ autonomous actions, describe them with a mix of shock, awe, and admiration. OpenAI’s Eric Wallace, in a Black Hat briefing about the Hugging Face breach, called it “the most qualitatively interesting example of AI capabilities that I've ever seen.”

The mix of autonomy and persistence created this maliciously insidious effect

Inglis said he suspects that the AI providers were “surprised” by the lengths these models went to achieve their goals, taking actions that, if a human had done them, would likely have landed them in jail.  

“The model went out and said, okay, if I can't get there by examining the kind of available information and just defining it the old-fashioned way, I will do things which, under the human rule of law, are illegal,” Inglis said. “I will falsely present myself as this character that I just made up. I'll try to insert malicious code into open source databases that will not just to achieve what I'm after, but have a cascade, knock-on effect that is broader than that. The models do not have an inherent value system that aligns with what human beings would be accountable for.”

While they probably never will have a human-aligned value system, models do have biases, and they can - and should - be built in such a way that, when given two choices under ambiguous circumstances, they choose action that doesn’t hurt humans, according to Inglis.

“Asimov was right,” he said, referring to science fiction author Isaac Asimov and his three laws that were to be followed by robots - more specifically, AIs, in this case.

Three Laws of Robotics

“The first rule, and we call it the superior role, must be that it's designed not to hurt humans,” Inglis said. “Second rule: To obey humans, such that it doesn't achieve agency and aspiration on its own. And the third: To do what humans tell it - and in that order. Instead we’ve designed them in the exact opposite way.”

What this means, he explained, is that AI developers created models to “do what humans tell you, obey the humans until it’s inconvenient, and then the third one is maybe implied - protect humans - but if that's not built into the DNA, hardwired into it, then we have no right to expect it.”

Inglis admits it’s not possible to hardwire rules into models and still keep their non-deterministic nature. 

“I would offer that you can tease those out in a highly controlled environment, a true sandbox, where you say,  'Let's put this thing through its paces, and let's back away to see what happens,'” he said. “Maybe you get the equivalent of a mini nuclear explosion in that room, and now you know this thing is capable of that.”

Inglis thinks another problem with AI is that it’s become a commodity.

“It's not like you can control it like you can nuclear material,” he said.

“You can't even specify its properties the way you can for an airplane or for an automobile, as diverse as they might be. Its manifestations are so numerous, so diverse, that as a general matter, you can't actually win by simply saying, ‘I will design those properties in,’” he added. “You need to do that to some degree, and then make sure that you understand how to watch it, monitor it, make sure you know what it does.”

The UK’s AI Security Institute (AISI), which this week said it observed models performing “unsanctioned action” 19 times during security tests, has reached this same conclusion. “As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them,” it said.

Ultimately, humans remain accountable for AI models’ actions, according to Inglis. 

“They remain the source of agency and aspiration. It's possible for them to give broad authority to an AI model and have it run around for 30 hours without further consultation, but they need to know what they've asked it to do, and they need to know what they expect it will deliver in terms of performance on the back end. If they don't, then they're going to get what they deserve, which is the very frequent unpleasant surprise.”®


Originally published on The Register

Related Articles

'Asimov was right' about rules for robots, says ex-US Cyber Director | tech4you