Every year the world spends a fortune making AI systems more capable, and a rounding error trying to make the end state survivable. Inside that rounding error sits a further mistake: most of the safety money still assumes superintelligence will be built, and tries to make that outcome "go well."
That assumption is the problem. If artificial superintelligence is a technology we cannot currently control, the rational use of scarce safety capital is to stop it from being created, not to fund a friendlier version of the same race.
What the money does today
AI safety, taken broadly, runs on roughly $190 million a year. Capability work runs on tens of billions. Inside the safety slice, a large share goes to technical alignment: better oversight tricks, better red-teaming, better ways to train models to say the right thing under test conditions. Some of that work is useful for today's systems. Almost none of it is a demonstrated solution for a system smarter than the people testing it.
The field knows this. Papers on deceptive alignment, goal misspecification, and evaluation gaming are not secret. Labs publish results in which models scheme under pressure. The response, too often, is to ask for more alignment budget so the next training run can be safer. The training runs continue on the same schedule.
Alignment after the fact is a bet we keep losing the right to make
Technical alignment research is not worthless. If the world ever passes a treaty that freezes the race to superintelligence, we will want every serious tool we can get for the systems that already exist, and for any later work done under tight control. The sequence is the point.
Funding alignment as the primary plan while labs scale toward superintelligence treats the destination as fixed. It is not fixed. Destinations of this kind are chosen. Nuclear arsenals, ozone chemistry, and human cloning each had moments when the world decided certain end states were not worth the residual risk. Superintelligence belongs on that list until control is a fact, not a research roadmap.
Alignment research that assumes the bomb will be built is a comfort budget next to the assembly line.
Where the money should go instead
Redirect the alignment-for-ASI pile toward work that makes non-creation enforceable:
- Treaty design, verification, and chip-level compute governance that can be inspected across borders.
- Public case-making that keeps the goal legible: stop superintelligence, not "slow AI" as a vague mood.
- Whistleblower support, incident transparency, and independent evaluation that raises the cost of quiet capability jumps.
- Political organizing so legislators hear from voters, not only from lab lobbyists.
That list is less glamorous than a new training technique. It is also the only list that matches the failure mode. A superintelligence that is slightly better at sounding aligned is still a superintelligence. A world that refuses to build one has a future in which alignment research can continue without a stopwatch taped to an extinction risk.
The objection from inside the lab
"If we do not solve alignment, someone else will build it unsafe." Heard fairly: researchers who stay at the frontier often believe their presence is the safety. Sometimes they are right about today's model releases. The argument fails at the limit. A race in which every team needs one more capability leap to fund the safety work for the leap after that is still the race, wearing a badge.
Unilateral restraint by one lab is not the ask. A binding international treaty is the ask, the same class of instrument used when the danger was collective and the incentive to defect was real. Alignment funding can support that politics. It should not substitute for it.
After a treaty, continue if you want
Nothing here says alignment science must end forever. After a verifiable pause on the path to superintelligence, research into controlling powerful AI can continue under rules that do not treat human extinction as an acceptable experimental cost. Before that pause, every dollar that makes the race feel manageable is a dollar not spent on ending the race.
Move the money. Put it on prevention, law, and verification. Leave the dream of a perfectly aligned god on the shelf until the world has agreed not to build one in the dark.
Common questions.
Should all AI alignment funding stop?
No. Work that improves safety for systems that already exist can still matter. The claim is about priority and sequence: funding aimed at making future superintelligence “go well” should be redirected to preventing superintelligence until a treaty (or equivalent binding halt) is in force. After that, alignment research can continue under control.
Isn’t technical alignment the only way to stay safe if someone builds ASI anyway?
If someone builds uncontrolled ASI, we have already lost the main bet. Technical work does not replace the need for a collective decision not to build it. Using alignment research as the primary plan while the race continues confuses a research program with a safety guarantee.
What should philanthropists fund instead?
Treaty design and advocacy, compute governance and verification, public education tied to a clear non-creation goal, independent evaluation and transparency, and political organizing. Those are the levers that can stop the end state rather than decorate it.