nhorozov.xyz/thoughts/thought 2026 09 18

thought 2026 09 18

thinking; the best possible actions you can take are probably not very different that the worst possible actions you can take. if you go all-out on solving world hunger by making food hyper-cheap to produce, imagine a meal costing a thousandth of a penny, then obesity might increase a lot. the specific details of your action matter a lot when you are trying to do the best possible thing, otherwise when exerting that effort you will miss/overshoot/overoptimize badly.
in ais a lot of safety work is dual-use for capabilties, an example being RLHF. worse than capabilties, which undersells the harm of RLHF, it also can teach smart models to be sycophantic and (sometimes, seen it) subtly passive aggressive.
(goodhart’s law)