nhorozov.xyz/thoughts/superrationality cosmic host
superrationality cosmic host (?)
epistemic status: somewhat confident in the general intuitions, i might be using some terminology a little wrong though. all of this is fairly out there
i was reading bostrom’s cosmic host and i was thinking that superrationality might be common amoung SIs (cooperation without communication in the prisoners dilemma).
ASI (made by humans) probably wont have real superrationality for a while because our AIs are all trained on roughly the same training data with similar architectures, and should cooperate with each other much earlier. In FDT two identical agents always cooperate, and (i think, but probably) as they become more different they are less likely to cooperate. As you increase their intelligence theyre more likely to cooperate (bc superrationality). So whatever threshold of intelligence that needs to be passed for an entity to be superrational, we are probably going to see superrationality before we hit that threshold in our own AIs (assuming we keep making them smarter, which we might not for obvious safety reasons).
Superrationality would make sense to be a cosmic norm, and maybe intentionally seeking to create superrational ASI would be in line with that. A much safer way to do this would be to train an non-superintelegent AI to be superrational. This could be done by RL SFT’ing a synthetic reasoning trace in which the AI would try to reason what the other agent would do, and if it knows (somehow) the other is also superrational it would cooperate.
this was fun to think abt
18/09/2026 10:12:25