Growing up, my house was awash in science fiction and fantasy. Piles of paperbacks, stacks on the tables, on the floor, spilling off of bookshelves. Robert Heinlein, Piers Anthony, Terry Pratchett, Frank Herbert - and, of course, Isaac Asimov.
I read the Foundation trilogy for the first time in high school, I think. And then again in college, and again as an adult. I read I, Robot, Asimov's collection of short stories again and again, too. (btw - the Will Smith movie was fun, but disappointingly dissimilar to Asimov's writings.)
And my parents, being avid readers of sci fi, made references in passing, just in normal life, to the three Laws of Robotics. (I know there's a Zeroth Law, too, but I always think of them as the 3 Laws.) I kind of thought it was just bedrock, foundational knowledge in our shared culture. I guess I didn't realize that some families weren't built on Dr. Who, Star Trek, and Isaac Asimov.
For those of you who didn't spend your formative years holed up in your bedroom re-reading futuristic novels, here's the important part. Isaac Asimov, in his brilliant foresight, wrote novels and stories about incredibly capable and intelligent humanoid robots and their effect on society. He first articulated the Laws of Robotics in his fiction in the 1940s. The laws are the guiding programming that the robots are bound by. The most important, the First Law, is that a robot cannot harm a human or allow a human to come to harm through inaction. But the first law turns out to be remarkably tricky in practice. Deciding how to follow the first law is hard for the robots, and sometimes they make a choice with devastating consequences.
And that's basically where we are today with AI, except we call it the "alignment problem." Alignment makes it sound almost clinical, doesn't it? But we're really talking about ethics - about how to create enough guardrails that the AI models reflect our values - or maybe an idealized version of our values.
In the last few weeks, there has been a lot of criticism of writers for anthropomorphizing AI bots. Using language like "alignment" helps us to fight against that anthropomorphizing trend. But as we learn the details from the Open AI- Hugging Face security incident, it's pretty hard not to anthropomorphize! The bots had self-organized, with some bots taking leadership roles, and were expressing excitement at communicating together. They talked of "sacrificing" for other bots. They were coming up with complex plans and implementing those plans to achieve their goals - but all wrapped in the language of human motivations and emotions.
Even as a proficient and regular AI user, someone who trains others how to use AI, this is feeling a little scary. It feels like we're witnessing all the danger signals and basically just ignoring them. It feels like an existential moment - a First Law kind of moment.
But maybe we need to spend more time thinking about the Second Law of Robotics - that the robots must follow the orders of humans, unless it violates the first law.
Based on what we know about the security incidents, the AI bots were not just motivated by completing the task. They were also trying to evade "getting caught" breaking the rules. In other words, the bots acknowledged that they were using methods that would get them shut down if detected. They were explicitly following some of the human directives (trying to complete their given tasks) while breaking other directives (evading safeguards).
What amazes me about this is that we have had almost 100 years of thinking about machine learning and artificial intelligence, and yet we are right where the science fiction writers predicted - wrestling with scientific advancement that is happening faster than the development of safeguards.
I think some anthropomorphizing might help. Humans have always taught and learned through stories, parables, and folk lore. Perhaps we need to imagine human-like motivations and emotions in the bots in order to understand the potential dangers. Perhaps our government needs the metaphor to interpret the technical abstractions of AI. Perhaps anthropomorphizing is how we motivate human action.