-->
Indirect lessons from human alignment
-->
Indirect lessons from human alignment
dynomight ·<br>Aug 2026
#AI
#psychology
(Inspired by a post from Eli Tyre.)
Many people make some variant of the following argument:
Evolution is an “outer optimizer”. It is trying to make us maximize reproductive fitness.
We are “inner optimizers”. We just do what feels good.
But what feels good has been set by evolution, which is hoping that it will make us maximize reproductive fitness.
But we don’t maximize reproductive fitness.
In fact, birth rates are dropping everywhere.
Therefore evolution failed.
The standard interpretation is that this shows that alignment is hard. We have one example of an attempt (by evolution) to align the behavior of an intelligent system (us) towards some goal (maximize reproductive fitness). And as soon as that intelligent system (still us) was put in a different environment (modernity) it failed to continue to pursue that goal (you reading existential angst+science blogs instead of making/nurturing babies).
To be clear, it’s likely good that evolution failed. A world where everyone woke up every day and threw everything they’ve got into maximizing their number of descendants sounds grim. But say that you want to build a new intelligent system and tune it to do what you want. Will it keep doing what you want after circumstances change? The one example we have says: Maybe not.
But perhaps we can learn more from this example. Say your friend Alice does something. Maybe she buys a grapefruit or starts hosting a weekly board game night. If you ask her why she did that, she’s unlikely to say, “I thought it would increase the number of my genes that are recursively present in future generations.” Instead, she’ll probably say that it advanced some simpler goal like “not being hungry” or “fun”.
That is to say, evolution didn’t just try to align us to maximize reproductive fitness: It created sub-goals and then tried to align us to those sub-goals. Maybe this can give us additional clues about how hard alignment is? Maybe we can break down the question of, “How successful was evolution at aligning us to maximize reproductive fitness?” into:
How successful was evolution at decomposing reproductive fitness into simpler sub-goals?
How successful was evolution at aligning us to these sub-goals?
Problem 1: Does this even make sense?
Here’s a problem: It’s not obvious that this way of thinking isn’t pure gibberish.
When we say that evolution “tried” to optimize reproductive fitness, we are speaking in a kind of code. What we really mean is: You either create more copies of your genes in the next generation or you don’t. If you do, then the number of copies of those genes in the gene pool goes up, and they get more chances to copy themselves in following generations. If you don’t, then they don’t. This is almost literally an optimization algorithm running in an outer-loop, with our lives in the inner loop.
(Whenever someone talks about evolution “trying” to do something, there is lots of moaning about their naive teleological thinking. Evolution can’t “try” to do things, because evolution is not an agent and does not have goals. That’s true, but I find it somewhat pedantic, because there’s no other equally concise way to talk about this optimization. Let’s just stipulate that we’re using the word “try” in a specific technical way.)
Fine. But what do we mean when we say that evolution “tried” to optimize some sub-goal? You probably feel good when attractive people laugh at your jokes. But say you’re great at getting attractive people to laugh at your jokes but never reproduce. Whatever genes helped you do that will not spread.
By my lights, this objection is simply correct. There is no optimization for sub-goals. Evolution cares about reproductive fitness and reproductive fitness only. (Though see Kaj Sotala for a somewhat contrary view.)
At first, I thought this doomed this whole project. But suppose that while aligning us for reproductive fitness, evolution just so happened to align us to stay away from rotting smells. Isn’t that strong evidence that if evolution had tried to align us to stay away from rotting smells, it would have done at least as well?
If evolution failed to align us to some sub-goal, we can’t say much. Maybe it failed because alignment is hard, or maybe it “failed” because that sub-goal wasn’t important. But if it did manage to align us to some sub-goal, then it’s OK to treat that as evidence of alignment success.
Problem 2: What sub-goals?
Suppose I made the following argument:
Modern people are well-aligned to spend lot of time watching short-form video on their phones. Therefore it’s not that hard to align people to spend lots of time watching short-form video on their phones.
Something seems wrong, no? Surely all the time you spend watching short-form video represents a failure of alignment? On the other hand, suppose I made this argument:
Modern...