The basic problem still remains: if you build an autonomous machine intelligence and try to encode it with basic directives, the potential implications of those directives is hard to predict. Obviously the paperclip company doesn’t want to replace the entire universe with a grey goo any more than the sorcerer’s apprentice wants to flood the workshop; it happens accidentally.
Of course the paperclip company can try to add constraints to their AI in order to prevent naive paperclip maximization, but what if they screw up those constraints as well? The whole premise of Asimov’s Three Laws is that AI has these sorts of constraints, but even in his stories these constraints still lead to unexpected outcomes.
All programming bugs are the result of a programmer encoding an instruction or statement that doesn’t imply what they think it implies and the computer following it literally. A more capable and autonomous computer that approaches what we might call “intelligence” is also going to be more capable of doing harm when it runs into a bug. And if it’s something like an LLM where the instructions are natural language, with all its ambiguity and vagueness, you have a whole other issue compounding it.
If you study philosophy you end up running into the exact same problem. The object of the game of philosophy is to make the most general true statements possible. One philosopher might say something like, “knowledge is defined as true justified belief”, or maybe “moral good is defined as whatever delivers the greatest good to the greatest number”, or maybe even, “the object of the game of philosophy is to make the most general true statements possible”. And then another philosopher comes up with a counterexample or counterargument which disproves the first philosopher’s statement, usually because—just like a programming bug—it entails an implication that the first philosopher didn’t think of. We have been playing the game of philosophy for thousands of years and nobody has managed to score a point yet.
Another thing. Human beings have a lot of needs, imperatives, motivations, and values. Some of them, like food, are built in. Others are learned through culture. But we end up with a lot of them, and it’s easy to take them for granted. With a machine, you have to build those things in yourself. There’s no getting around it. But we don’t actually have a complete, hierarchical set of imperatives/motivations/values for a decent human being. The philosophers have been working on it for millennia but keep running into bugs. So how can we expect to solve the problem for non-human AI? True, we are unlikely to screw up so badly that we end up with a literal paperclip maximizer, but we are bound to make some far more subtle mistake of the same general kind.
Of course the paperclip company can try to add constraints to their AI in order to prevent naive paperclip maximization, but what if they screw up those constraints as well? The whole premise of Asimov’s Three Laws is that AI has these sorts of constraints, but even in his stories these constraints still lead to unexpected outcomes.
All programming bugs are the result of a programmer encoding an instruction or statement that doesn’t imply what they think it implies and the computer following it literally. A more capable and autonomous computer that approaches what we might call “intelligence” is also going to be more capable of doing harm when it runs into a bug. And if it’s something like an LLM where the instructions are natural language, with all its ambiguity and vagueness, you have a whole other issue compounding it.
If you study philosophy you end up running into the exact same problem. The object of the game of philosophy is to make the most general true statements possible. One philosopher might say something like, “knowledge is defined as true justified belief”, or maybe “moral good is defined as whatever delivers the greatest good to the greatest number”, or maybe even, “the object of the game of philosophy is to make the most general true statements possible”. And then another philosopher comes up with a counterexample or counterargument which disproves the first philosopher’s statement, usually because—just like a programming bug—it entails an implication that the first philosopher didn’t think of. We have been playing the game of philosophy for thousands of years and nobody has managed to score a point yet.
Another thing. Human beings have a lot of needs, imperatives, motivations, and values. Some of them, like food, are built in. Others are learned through culture. But we end up with a lot of them, and it’s easy to take them for granted. With a machine, you have to build those things in yourself. There’s no getting around it. But we don’t actually have a complete, hierarchical set of imperatives/motivations/values for a decent human being. The philosophers have been working on it for millennia but keep running into bugs. So how can we expect to solve the problem for non-human AI? True, we are unlikely to screw up so badly that we end up with a literal paperclip maximizer, but we are bound to make some far more subtle mistake of the same general kind.