Thank you very much for the piece, I had a question regarding MAC. Maybe this is due to my lack of understanding of the topic, but wouldn’t using MAC still result in a hegemony? For example the author references applying the same principles used in the nuclear arms race to AI, mentioning the prevention of Iranian development of nuclear weapons through economic sanctions. But this gives the US a clear advantage- they are able to control the production of nuclear weaponry and who develops it almost as if it is a monopoly. The authors statement that MAC requires “only that a sufficiently powerful state or coalition can unilaterally impose costs on its violators,” is also merely justified by the fact that “even adversarial powers share an interest in avoiding catastrophic outcomes.” They reference the Cold War arms race, however I believe the arms race ended not from mutual recognition that escalation was dangerous for all parties involved, but because the US was able to establish international dominance.
I understand that this also reflects the Leviathan as well- Hobbes did support a monarchy after all, and I believe at least that the sovereign is supposed to represent the ultimate artificial representative of the people. However a monarchy itself poses the threat of totalitarianism, which many argue that we are also beginning to see in the US as of now too, and though the prisoners dilemma is temporarily resolved it still does not form equilibrium due to an imbalance of powers and would most likely emerge in the dilemma rising once more through alternative uprisings. Please do correct me if I have misinterpreted the article or have provided any factually incorrect information and thank you again for the piece, it was interesting to read.
The article references game theory repeatedly and says we are in a "classic Prisoner's Dilemma." But this is by no means clear.
In the classic prisoner's dilemma, there are two agents each of whom have two options: cooperate or defect. Importantly, for both agents, the "dominant" strategy in this game (i.e. the strategy that's in an agent's best interest regardless of the other's actions) is to defect.
In lay terms, the thinking goes like this:
1. If I continue developing AI and my rival pauses, then I gain an advantage.
2. Alternatively, if I continue developing AI and my rival also continues, at least I won't fall behind too much.
Therefore, regardless of what strategy my rival pursues, I should continue developing AI.
The author suggests basically this logic, saying that “the structure of AGI development seems to intrinsically reward accelerationism and punish restraint (even if the ultimate consequences for reckless accelerationism are unimaginably catastrophic).”
But this ignores the catastrophic risks that many AI researchers are warning about! Yes, he mentions it parenthetically, but he fails to incorporate it into the game-theoretic analysis.
What happens if we believe a sufficiently advanced AI will kill humanity?
This dramatically changes the payoff for racing ahead. To alter the line of thinking above, it would look something like this:
1. If I continue developing AI and my rival pauses, then I develop something that kills humanity.
2. Alternatively, if I continue developing AI and my rival also continues, then either of us develop something that kills humanity.
So the optimal course of action is to NOT continue developing AI since racing ahead means doom for everyone.
This is very different from the classic Prisoner’s Dilemma!
Of course, this assumes that an advanced AI would be catastrophically dangerous. So to be more realistic, we can instead assign some probability to the claim that advanced AI will kill everyone. And we might also assign some probability to the claim that advanced AI will usher in an era of unprecedented abundance. And depending on these and other factors, the payoffs for cooperation/defection can be very different, thereby leading to very different policy considerations.
The important point, however, is that we should not simply assume (and the author does a poor job demonstrating) that we are facing a prisoner’s dilemma situation. We could instead be playing other games.
You’re right that if decision-makers believe there is a high chance that advanced AI will cause catastrophic harm, then the payoffs change, and the situation may no longer look like a Prisoner’s Dilemma. But that is an empirical disagreement about how the key decision-makers see the payoffs, not about the logic of the model itself.
That said, many (if not most) policymakers already describe U.S.–China AI competition as a race. The White House’s official strategy is literally titled “Winning the AI Race: America’s AI Action Plan,” and Georgetown’s Center for Security and Emerging Technology notes that the AI “arms race” has become a common (if not dominant) way of describing competition between the United States and China.
If leaders believe that slowing AI development while their rivals continue would leave them at a strategic disadvantage, then continuing to accelerate is a rational choice from their point of view. This is exactly the logic of a Prisoner’s Dilemma: each player chooses the strategy that seems best for themselves, yet when both do so, both end up worse off than if they had cooperated.
The essential point is that strategic behavior is driven by the payoffs that decision-makers believe they face, not necessarily by the payoffs an outside observer thinks they face.
So even if another game might better describe the objective situation (in a world where people who actually make the decisions did not think we are in an arms race), in our world, widespread belief among key decision-makers that AI development is an arms race can and does create arms-race behavior.
Yes, policy makers might use the term "race." But not every race has the structure of the classic prisoner's dilemma.
In my earlier comment, I noted one feature of the prisoner's dilemma: the dominant strategy is to defect. But if an agent believes AI will kill everyone, then this would not be the dominant strategy. The race would not be an arms race, but a suicide race. So it is not true to say that the AI race "INTRINSICALLY rewards accelerationism and punishes restraint." (emphasis mine)
And if a sufficiently advanced AI system would in fact kill everyone, then it would be unfortunate if everyone believed that we were in a prisoner's dilemma, since this would encourage the very acceleration that would lead to annihilation. People who believe there is a grave danger should instead caution against framing our situation as a prisoner's dilemma.
Another feature of the prisoner's dilemma is that when both agents defect, they are both worse off than if both had cooperated. But this also is not an INTRINSIC feature of the AI race. For example, in the White House strategy report that you cite, it does not suggest that slowing down and cooperating with China is preferable in any way. Rather, it argues that the US should speed ahead while also crippling China's capabilities. This is because the policy makers believe the US will outright win the race, so "defecting" is actively good for us even if China tries to catch up. This too suggests a different payoff logic from the classic prisoner's dilemma.
Thank you very much for the piece, I had a question regarding MAC. Maybe this is due to my lack of understanding of the topic, but wouldn’t using MAC still result in a hegemony? For example the author references applying the same principles used in the nuclear arms race to AI, mentioning the prevention of Iranian development of nuclear weapons through economic sanctions. But this gives the US a clear advantage- they are able to control the production of nuclear weaponry and who develops it almost as if it is a monopoly. The authors statement that MAC requires “only that a sufficiently powerful state or coalition can unilaterally impose costs on its violators,” is also merely justified by the fact that “even adversarial powers share an interest in avoiding catastrophic outcomes.” They reference the Cold War arms race, however I believe the arms race ended not from mutual recognition that escalation was dangerous for all parties involved, but because the US was able to establish international dominance.
I understand that this also reflects the Leviathan as well- Hobbes did support a monarchy after all, and I believe at least that the sovereign is supposed to represent the ultimate artificial representative of the people. However a monarchy itself poses the threat of totalitarianism, which many argue that we are also beginning to see in the US as of now too, and though the prisoners dilemma is temporarily resolved it still does not form equilibrium due to an imbalance of powers and would most likely emerge in the dilemma rising once more through alternative uprisings. Please do correct me if I have misinterpreted the article or have provided any factually incorrect information and thank you again for the piece, it was interesting to read.
Very fine piece!
Thank you!
The article references game theory repeatedly and says we are in a "classic Prisoner's Dilemma." But this is by no means clear.
In the classic prisoner's dilemma, there are two agents each of whom have two options: cooperate or defect. Importantly, for both agents, the "dominant" strategy in this game (i.e. the strategy that's in an agent's best interest regardless of the other's actions) is to defect.
In lay terms, the thinking goes like this:
1. If I continue developing AI and my rival pauses, then I gain an advantage.
2. Alternatively, if I continue developing AI and my rival also continues, at least I won't fall behind too much.
Therefore, regardless of what strategy my rival pursues, I should continue developing AI.
The author suggests basically this logic, saying that “the structure of AGI development seems to intrinsically reward accelerationism and punish restraint (even if the ultimate consequences for reckless accelerationism are unimaginably catastrophic).”
But this ignores the catastrophic risks that many AI researchers are warning about! Yes, he mentions it parenthetically, but he fails to incorporate it into the game-theoretic analysis.
What happens if we believe a sufficiently advanced AI will kill humanity?
This dramatically changes the payoff for racing ahead. To alter the line of thinking above, it would look something like this:
1. If I continue developing AI and my rival pauses, then I develop something that kills humanity.
2. Alternatively, if I continue developing AI and my rival also continues, then either of us develop something that kills humanity.
So the optimal course of action is to NOT continue developing AI since racing ahead means doom for everyone.
This is very different from the classic Prisoner’s Dilemma!
Of course, this assumes that an advanced AI would be catastrophically dangerous. So to be more realistic, we can instead assign some probability to the claim that advanced AI will kill everyone. And we might also assign some probability to the claim that advanced AI will usher in an era of unprecedented abundance. And depending on these and other factors, the payoffs for cooperation/defection can be very different, thereby leading to very different policy considerations.
For a more careful breakdown, see here:
https://www.lesswrong.com/posts/uFNgRumrDTpBfQGrs/let-s-think-about-slowing-down-ai#The_arms_race_model_and_its_alternatives
The important point, however, is that we should not simply assume (and the author does a poor job demonstrating) that we are facing a prisoner’s dilemma situation. We could instead be playing other games.
You’re right that if decision-makers believe there is a high chance that advanced AI will cause catastrophic harm, then the payoffs change, and the situation may no longer look like a Prisoner’s Dilemma. But that is an empirical disagreement about how the key decision-makers see the payoffs, not about the logic of the model itself.
That said, many (if not most) policymakers already describe U.S.–China AI competition as a race. The White House’s official strategy is literally titled “Winning the AI Race: America’s AI Action Plan,” and Georgetown’s Center for Security and Emerging Technology notes that the AI “arms race” has become a common (if not dominant) way of describing competition between the United States and China.
If leaders believe that slowing AI development while their rivals continue would leave them at a strategic disadvantage, then continuing to accelerate is a rational choice from their point of view. This is exactly the logic of a Prisoner’s Dilemma: each player chooses the strategy that seems best for themselves, yet when both do so, both end up worse off than if they had cooperated.
The essential point is that strategic behavior is driven by the payoffs that decision-makers believe they face, not necessarily by the payoffs an outside observer thinks they face.
So even if another game might better describe the objective situation (in a world where people who actually make the decisions did not think we are in an arms race), in our world, widespread belief among key decision-makers that AI development is an arms race can and does create arms-race behavior.
Yes, policy makers might use the term "race." But not every race has the structure of the classic prisoner's dilemma.
In my earlier comment, I noted one feature of the prisoner's dilemma: the dominant strategy is to defect. But if an agent believes AI will kill everyone, then this would not be the dominant strategy. The race would not be an arms race, but a suicide race. So it is not true to say that the AI race "INTRINSICALLY rewards accelerationism and punishes restraint." (emphasis mine)
And if a sufficiently advanced AI system would in fact kill everyone, then it would be unfortunate if everyone believed that we were in a prisoner's dilemma, since this would encourage the very acceleration that would lead to annihilation. People who believe there is a grave danger should instead caution against framing our situation as a prisoner's dilemma.
Another feature of the prisoner's dilemma is that when both agents defect, they are both worse off than if both had cooperated. But this also is not an INTRINSIC feature of the AI race. For example, in the White House strategy report that you cite, it does not suggest that slowing down and cooperating with China is preferable in any way. Rather, it argues that the US should speed ahead while also crippling China's capabilities. This is because the policy makers believe the US will outright win the race, so "defecting" is actively good for us even if China tries to catch up. This too suggests a different payoff logic from the classic prisoner's dilemma.