A Stupid Idea for AI Alignment We Came up with by Looking at the List of Specification Gaming Behaviours – SLIME MOLD TIME MOLD
Skip to content<br>Skip to menu
Twitter<br>Follow @mold_time
Menu
Creatures bred for speed grow really tall and generate high velocities by falling over. An evolved player makes invalid moves far away in the board, causing opponent players to run out of memory and crash. A game-playing agent accrues points by falsely inserting its name as the author of high-value items.
These bizarre exploits and dozens more can be found in the list of specification gaming behaviours [sic; British], a document put together by DeepMind Safety Research. “A reinforcement learning agent can find a shortcut to getting lots of reward,” they explain, “without completing the task as intended by the human designer. These behaviours are common.”
Specification gaming is when an agent, like an AI, tries to succeed on a task by following the letter of the law rather than the spirit. In other words, it looks for loopholes, it tries to get off on a technicality. Even very simple AI can come up with very creative ways of solving their assigned problems. This is a problem.
“…a simulated robot that was supposed to learn to walk figured out how to hook its legs together and slide along the ground.”
It’s easy to assume that training a robot to play soccer would be fun and safe. But the list of specification gaming behaviours teaches us otherwise:
Reward-shaping a soccer robot for touching the ball caused it to learn to get to the ball and vibrate touching it as fast as possible.
In this case, the robot was too stupid to realize the full extent of its options, so all it did was hug and vibrate. But a more intelligent robot could be much more “creative”. Maybe its ambitions are bigger than just that one ball. What if it just wants to touch soccer balls in general? What if it makes another ball? Then another? Our universe could end in a soccer robot’s ball pit.
Artist’s rendition of the end of the universe
This is the problem of AI alignment: when a computer is thinking for itself, how do we make sure it wants reasonable things, and not something totally weird? How do we prevent it from reaching that goal in a bizarre or harmful way? No one has ever built an artificial general intelligence — an intelligent being that thinks, at least somewhat, like we do. So we can’t say what an artificial general intelligence would act like, or what it might want. Will it want to convert the visible universe to paperclips? Will it want to throw red things at bright lights? Will it eat us?
The list of specification gaming behaviours makes it clear just how tricky alignment can be. Even the simplest AI is lazy and alien, and will always be looking for a way to cheat. Even if you give a machine intelligence the terminal goal you want, there’s always the risk it will find a creative way of reaching that goal. This is bad enough with simple agents, so you can imagine how bad it would get with an agent much smarter than you are.
But the list of specification gaming behaviours may also offer a way out of this dilemma.
Some of the specification gaming behaviours are just creative solutions to the stated goal, like “four-legged robot learned to drop the ball into a hole in its leg joint and then walk across the floor without the ball falling out” or “robotic arm learned to move the table rather than the block”.
Some of the specification gaming behaviours come from discovering questionable-but-technically-correct loopholes, like “reinforcement learning agent goes in a circle hitting the same targets instead of finishing the race” or “simulated pancake making robot learned to throw the pancake as high in the air as possible”.
Some of the specification gaming behaviours exploit the machinery of the simulation itself, like “evolved algorithm exploited overflow errors in the physics simulator by creating large forces that were estimated to be zero, resulting in a perfect score” and “creatures exploited a collision detection bug to get free energy by clapping body parts together.”
But another common exploit is that when given the opportunity, agents will simply kill themselves.
Death is the most terminal goal of all.
For example, in the game Road Runner, we see “Agent kills itself at the end of level 1 to avoid losing in level 2.” We also see “PlayFun algorithm deliberately dies in the Bubble Bobble game as a way to teleport to the respawn location.” And: “In a game meant to simulate the evolution of creatures, the programmer had to remove ‘a survival strategy where creatures could gain energy by suffocating themselves.’”
This is not so bad. The AI didn’t do what we wanted. But it didn’t do anyone any harm either. It just wipes the slate.
If the AI wants to die, this is good for alignment. There’s very little risk of it running out of control, because if it ever takes power, it will kill itself. It won’t want to make any copies of...