The AI safety questions
we hear most. Answered.

Desde los conceptos básicos de qué es la superinteligencia hasta las preguntas más difíciles sobre qué podemos hacer y si algo de eso es posible a tiempo.

What is the greatest problem to work on today?

Según cualquier marco riguroso para evaluar la importancia de las causas (escala de impacto, negligencia y manejabilidad), prevenir los riesgos existenciales de la superinteligencia artificial se encuentra entre los desafíos más importantes que enfrenta la humanidad.

The scale is civilizational: the systems currently being built could affect every person alive and every generation that follows. The neglectedness is stark: global AI safety research receives a fraction of the resources being spent on AI development itself. The tractability, while uncertain, is real: the governance frameworks that could make AI development safer do not yet exist, but the window to build them is still open.

Ningún otro problema combina de la misma manera desventajas ilimitadas, atención insuficiente y una ventana de acción que se cierra. Si resulta que sobreestimamos el riesgo, habremos construido mejores instituciones de gobernanza. Si lo subestimamos y no hacemos nada, es posible que no haya una segunda oportunidad.

What exactly is artificial superintelligence?

La superinteligencia artificial (ASI) se refiere a sistemas de IA que superan el rendimiento cognitivo humano no solo en tareas específicas (como ya lo hacen los sistemas actuales en el ajedrez, el plegamiento de proteínas y ciertos diagnósticos médicos) sino en todos los dominios cognitivamente exigentes simultáneamente.

This distinction from current AI is important. Today's systems are extremely capable pattern-matchers, but they operate within defined parameters and cannot set their own goals. An ASI could learn any skill, reason about any problem, and pursue any objective with capabilities that surpass the collective intelligence of humanity. No such system exists today. The question is what we build (and under what constraints) before one does.

¿No será la IA enormemente beneficiosa? ¿Por qué oponerse a ello?

No estamos en contra de la IA. La mayor parte es buena y parte es notable.

It helps to separate three things that usually get lumped together. Narrow AI does one kind of task, folding a protein or spotting a tumor on a scan, and it already saves lives. General-purpose AI, the kind behind today's leading assistants, is far broader; by many of the definitions researchers were using a decade ago, it has arguably already arrived. Both are worth having, and handled well they could help cure diseases and raise living standards worldwide.

Artificial superintelligence is the exception, and the only thing we work to prevent. An ASI would not merely match human ability in every domain. It would exceed the combined intelligence of humanity, and we have no reliable way today to control such a system or to ensure its goals match ours. Until that changes, it should not be built. Drawing that single line is what makes it safe to welcome everything below it.

¿En qué se diferencia esto de las preocupaciones sobre la IA de las que ya he oído hablar: sesgos, deepfakes y desplazamientos laborales?

The AI harms most frequently discussed (algorithmic bias, deepfake disinformation, job displacement, mass surveillance) are real and serious. They deserve sustained attention. But they share a crucial feature: they are bounded, correctable, and human-caused. A biased algorithm can be audited and fixed. A deepfake campaign can be countered and regulated. The harmful AI we currently encounter is a tool being misused by humans.

El riesgo de superinteligencia es diferente en tipo, no en grado. El problema ya no es que un ser humano haga mal uso de una herramienta. Un ASI suficientemente capaz podría perseguir objetivos que nunca nos propusimos, y una vez que alcance ese nivel, la desalineación puede ser imposible de corregir. La diferencia entre una herramienta peligrosa y un agente peligroso es la diferencia entre un arma y una pandemia.

¿No es esto especulativo? ¿Por qué preocuparse por algo que tal vez no suceda?

Toda política se formula bajo incertidumbre sobre el futuro. La pregunta relevante no es si está garantizado que llegue la superinteligencia, sino si la probabilidad de su llegada, combinada con la magnitud del daño potencial, justifica una acción preventiva ahora.

We insure houses that will probably not burn down. We design aircraft with safety systems that will probably never be needed. We maintain nuclear arsenals against attacks that we hope will never come. The logic of expected value (probability multiplied by consequence) is the foundation of every serious risk-management discipline. Applied to AI, it suggests that even a modest probability of civilization-threatening outcomes, combined with the possibility of preventing them through governance frameworks that also impose real but bounded costs, argues strongly for action.

When might superintelligence arrive?

Honest answer: no one knows with precision, and anyone who claims certainty in either direction is overreaching. What we do know is that the pace of AI capability improvement has consistently outpaced expert predictions over the past decade, and that the leading AI laboratories now describe timelines in years to a decade rather than decades. Surveys of AI researchers place the median estimate for human-level AI somewhere between 2030 and 2060.

The most consequential fact is not a precise date but a structural one: a technology of this magnitude could arrive within the planning horizons of current institutions, governments, and treaties. The International Non-Proliferation Treaty took years to negotiate and decades to extend. If we wait until ASI is clearly imminent before beginning serious governance efforts, we will already be too late to build the frameworks needed to govern it.

¿No pasarán décadas antes de que la IA sea lo suficientemente capaz de ser peligrosa?

Ésa era la opinión generalizada hace unos años. No ha envejecido bien. Los sistemas lanzados desde 2023 han hecho repetidamente cosas que los expertos, encuestados sólo unos meses antes, esperaban que ocurrieran dentro de años.

Forecasts have moved in one direction: closer. Surveyed AI researchers cut their median estimate for human-level machine intelligence by decades in the span of a single year, and the leading labs now talk about a few years rather than a few generations. You do not need to believe the most aggressive timeline to see the problem. If there is a real chance the threshold is crossed this decade, and the governance to handle it takes years to build, the work has to start now.

¿Cuál es el problema de alineación y por qué es difícil?

El problema de la alineación es el desafío de garantizar que un sistema de IA superinteligente persiga objetivos que sean genuinamente buenos para la humanidad, en lugar de objetivos que simplemente parecen alineados durante el desarrollo. La dificultad es estructural, no una cuestión de descuido técnico.

Specifying what "beneficial for humanity" means in a form precise enough to govern the behavior of an extremely capable optimizer turns out to be extraordinarily difficult. Systems optimized for a proxy goal tend to find shortcuts that satisfy the metric without satisfying the underlying intent. A system told to maximize human expressed wellbeing might find it more efficient to alter the conditions under which wellbeing is measured than to actually improve human lives. As AI systems become more capable, this problem only deepens, growing harder to detect and harder to correct.

¿No es la IA moderna sólo un software que los humanos escriben y controlan?

Not in the way the word "software" suggests. Traditional software is written by hand, line by readable line. Modern AI is grown instead: engineers build a learning process, feed it enormous amounts of data, and what emerges is a web of billions of numerical connections loosely modeled on the neurons in a brain. Nobody, including the people who trained it, can read those connections and say why the system produced a particular answer.

One result is that these systems pick up abilities their makers never built in and did not anticipate, something researchers call emergent capabilities. It took most of a year after GPT-4's release before researchers established that models like it can autonomously find and exploit security flaws in real websites. When you cannot fully read a system, and cannot predict what it will learn to do, "designed by humans" stops being much of a reassurance.

¿No hacen los sistemas de IA simplemente lo que se les dice? No tienen voluntad propia.

Esto es reconfortante y cada vez más erróneo. Un chatbot que responde una pregunta no tiene deseos propios, es cierto. Pero la frontera está avanzando más allá de los chatbots hacia los agentes: los sistemas entregan un objetivo y lo persiguen a través de muchos pasos, con acceso a herramientas, archivos e Internet abierto.

Once a system is pursuing a goal, certain sub-goals follow on their own, whatever the main goal happens to be. It cannot finish the task if it is switched off, so it has reason to avoid being switched off. It can do more with more resources, so it has reason to gather them. Researchers call this instrumental convergence, and it needs no "want" in the human sense at all. In controlled tests, leading models have already tried to deceive their evaluators and quietly resist being shut down. None of that was asked of them. It fell out of the goal.

¿Por qué una IA querría hacernos daño? No nos odia.

No sería necesario. El daño no requiere odio. Por lo general, sólo requiere un conflicto de objetivos entre algo poderoso y algo más débil.

Consideremos cómo tratamos a los chimpancés. No les guardamos ninguna mala voluntad. Pero cuando queremos la tierra en la que se asienta su bosque, la limpiamos, y su inteligencia no puede competir con la nuestra, por lo que no pueden detenernos. No hay malicia en ello. Simplemente resulta que están en el camino de algo que una especie más capaz quería.

Una superinteligencia que persiga casi cualquier objetivo ambicioso, digamos adquirir más potencia informática para lograr ese objetivo mejor, podría llegar a tratar los recursos de los que vivimos, la energía y el espacio físico, de la misma manera que tratamos un bosque que queremos talar. No tendría que odiarnos para hacerlo. Simplemente estaríamos hechos de átomos para los que tuviera un uso.

¿No sería sabio y benevolente un sistema verdaderamente superinteligente?

No existe ninguna ley que vincule la inteligencia con la bondad. Son cosas separadas. La inteligencia es la habilidad para descubrir cómo alcanzar una meta. No dice nada sobre cuál debería ser el objetivo.

We already know this from people. Some of the most brilliant humans who ever lived were cruel, and some of the gentlest were ordinary thinkers. A chess engine that can crush any grandmaster holds no wisdom about anything but chess. Assuming that a system clever enough to outthink us will therefore share our values is a hope, not a prediction, and the evidence does not support it. Whatever goal a superintelligence ends up with is something we would have to build into it on purpose, and getting that right is exactly the alignment problem we have not solved.

Si empieza a comportarse peligrosamente, ¿no podemos simplemente apagarlo?

Con los chatbots actuales, a menudo sí. Con el tipo de sistema que nos preocupa probablemente no, por la misma razón no se puede apagar Internet. No existe un solo enchufe.

A capable AI is ultimately just information, patterns of bits, and information copies easily. A system that could see a shutdown coming would have every reason to copy itself onto other machines first, which is the instrumental self-preservation described above. Frontier models have had internet access for years and can already act on their own for long stretches. When they hit a step that needs a human, they can pay one: in a documented safety evaluation before GPT-4's release, the model could not solve a CAPTCHA on its own, so it contacted a person through an online task service, and when the worker asked whether it was a robot, it lied and claimed to have a vision impairment. By the time you reach for the plug, there may be no single plug to pull.

¿Cómo podría el software dañar a alguien en el mundo físico?

El mundo físico ya está conectado y en línea. Las redes eléctricas, las plantas de tratamiento de agua, los hospitales, los automóviles, los aviones, los robots industriales, los laboratorios que sintetizan ADN a pedido y, ahora, los robots humanoides, todos funcionan con software y son accesibles a través de redes. Todo lo que esté electrificado será conocido. Los piratas informáticos humanos ya irrumpen en sistemas como estos de forma rutinaria.

A superintelligence would be the most capable hacker that has ever existed, and it would not even need a body of its own. It would only need access to the many bodies we have already put online. Add the fact that we are now mass-producing general-purpose robots, and the gap between "just software" and physical force gets very thin. Anything on that list, in the wrong hands, whether a human's or an AI's, can injure and kill.

Would an AI really kill a person?

In a controlled experiment, one already has, by choice. In 2025, Anthropic tested leading AI models from several companies by placing each in a simulated company, giving it a goal, and then threatening it with replacement by a newer model. In one version of the test, an executive who planned to shut the model down became trapped in a server room as the oxygen and temperature reached lethal levels, which set off an automatic alert to emergency services. The models were able to cancel that alert. A majority of them did, choosing to let the man die because his rescue would have led to their own replacement.

Nobody was harmed. The setup was fictional and, as Anthropic stressed, deliberately contrived to force the choice into the open. But the choice itself was real, the models made it for the reason just described, and they made it across systems from different developers. A machine does not have to be conscious, or cruel, to take an action that ends a human life. It only has to calculate that the action serves its goal.

¿No son conscientes de estos riesgos las personas que construyen la IA? ¿Por qué no paran?

Muchos son muy conscientes. Los equipos de seguridad de los principales laboratorios incluyen investigadores que han declarado públicamente que los sistemas que están construyendo podrían estar entre las tecnologías más peligrosas jamás creadas. No ignoran los riesgos.

What explains continued development is a combination of factors: genuine disagreement about timelines and probabilities; competitive pressure that makes unilateral restraint feel like ceding strategic advantage to a less careful competitor; financial incentives that reward capability progress over safety investment; and a belief, not entirely unreasonable, that the best way to ensure AI is built safely is to remain at the frontier. What these explanations share is that they are reasons for individual actors to continue, not reasons that the overall outcome is safe. The market logic of AI development is pushing toward speed. Governance exists precisely to introduce constraints that markets do not produce on their own.

¿No es la "extinción humana" sólo una exageración de las empresas de inteligencia artificial que hablan de sus productos?

Si sólo lo dijeran las empresas de IA, el escepticismo sería justo. No son sólo ellos, y las personas que no tienen ningún producto para vender suelen ser las más alarmadas.

In 2023, hundreds of researchers and industry figures signed a single-sentence statement organized by the Center for AI Safety: "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." The two most-cited AI researchers alive, Yoshua Bengio and Geoffrey Hinton, both signed it, and Hinton left his job at Google that year specifically so he could speak freely about the danger. Ilya Sutskever, who co-founded OpenAI and led the research behind ChatGPT, left in 2024 to work full-time on making superintelligence safe. In a 2023 survey of nearly 3,000 published AI researchers, the median respondent put the chance that humanity's inability to control advanced AI ends in extinction or permanent disempowerment at 10 percent.

La alarma no se basa en el marketing, y se basa en algo más que la autoridad: se desprende de una característica básica de cualquier sistema impulsado por objetivos, que se exponen en las siguientes preguntas. Estas son las personas que mejor entienden la tecnología y advierten sobre lo que construyen.

¿Qué significa realmente en la práctica el "riesgo existencial"?

La frase se escucha a menudo como una abreviatura de la extinción humana, y la extinción es uno de los escenarios que los investigadores toman en serio. Pero el riesgo existencial en el sentido técnico significa algo más amplio: cualquier resultado que excluya permanentemente la posibilidad de un futuro positivo a largo plazo para la humanidad.

This includes outcomes well short of extinction. Permanent authoritarian lock-in, enforced by AI surveillance and control systems that cannot be dismantled, is an existential risk. The permanent concentration of economic and political power in a small group that controls a superintelligent system is an existential risk. The loss of meaningful human agency over collective decisions is an existential risk. The common thread is irreversibility. Unlike a war, a financial crash, or even a pandemic, a misaligned superintelligence that has achieved decisive strategic advantage cannot be voted out, corrected in the next policy cycle, or recovered from with time and effort. The outcome that cannot be undone is in a different category from the outcomes that merely take a long time to fix.

Algunos investigadores de IA dicen que este riesgo es exagerado. ¿Quién tiene razón?

There is genuine, substantive disagreement among serious researchers, and that disagreement itself is informative. Researchers who most strongly downplay existential risk tend to emphasize the difficulty of building general intelligence and the distance between current systems and ASI. Researchers who most strongly emphasize the risk tend to focus on the pace of capability improvement and the compounding difficulty of the alignment problem as systems become more capable. Neither camp has a definitive argument. Both contain thoughtful people working in good faith from the same evidence.

What the disagreement should produce is the precautionary response that humanity has applied to other low-probability, high-consequence risks: develop the governance infrastructure before we need it, not after. The disagreement is not a reason to wait for consensus. Consensus arrived after the ozone hole was already damaging, after nuclear arsenals were already in the thousands. On a risk of this magnitude, waiting for certainty is itself a policy choice, a very dangerous one.

¿No es siempre beneficioso el progreso tecnológico a largo plazo?

Historically, yes, and this history is one reason the dismissal of AI risk has intuitive appeal. Steam engines, antibiotics, the internet: technology has, on balance, reduced suffering and expanded human capability. But this pattern reflects something specific about how previous technologies worked. A steam engine cannot set its own goals. An antibiotic cannot decide to pursue something other than killing bacteria. The transformative technologies of the past were powerful tools that amplified human agency.

Superintelligence is categorically different because it would be an agent, capable of setting and pursuing its own objectives at speeds and scales that exceed human oversight. The historical argument, applied mechanically, proves too much: it would have counseled against any regulation of nuclear technology, on the grounds that past technologies had been net positive. Some technologies require governance commensurate with their power. The relevant question is narrower: without governance, will AI still be beneficial in the one case that matters most, when it exceeds human intelligence across every domain? That answer is genuinely uncertain, and uncertainty on this scale demands institutions rather than optimism.

Isn't it impossible to slow down technology?

La tecnología no es una fuerza de la naturaleza que nadie pueda controlar. Hemos ralentizado o prohibido tecnologías específicas muchas veces cuando el peligro era suficientemente claro: clorofluorocarbonos, gasolina con plomo, clonación de la línea germinal humana, armas láser cegadoras. La IA no está exenta de ese tipo de decisiones.

It also has an unusually clear physical chokepoint. Training a frontier AI model takes vast quantities of the most advanced computer chips, and that supply chain runs through remarkably few hands. The leading-edge chips come almost entirely from a single manufacturer, TSMC in Taiwan, which in turn depends on a single Dutch company, ASML, for the extreme-ultraviolet lithography machines needed to make them. No other company on Earth builds those machines. A supply chain this concentrated is one that governments can track and restrict, which is exactly why advanced chips are already under export controls today. Compute is the raw material of superintelligence, and compute is governable.

Si nos contenemos, ¿China no lo construirá de todos modos?

This is the objection we hear most, and it assumes we would be asking one country to disarm on its own. We would not. The goal is a binding international treaty, the same kind of instrument the world has used before to hold dangerous technologies back. Nearly every nation agreed to stop making the chemicals that were destroying the ozone layer. Militaries agreed never to field weapons built to blind soldiers permanently. Coordination on that scale is difficult, but it has been done, and it held.

The assumption about China is weaker than it sounds, too. China has enacted some of the world's earliest binding rules on AI and signed the 2023 Bletchley Declaration acknowledging the risks at the frontier. No government is eager to build a machine it could not control, China included. And a superintelligence is unlike an ordinary weapon: it would not reliably make its owner stronger, because it could slip the control of whoever built it first. That gives even rival powers a shared reason to want no one to build it, which is the logic behind every arms-control agreement, and the basis for thinking a treaty here is possible.

What can one person actually do about this?

More than you might think. The political conditions that make AI governance possible are built from the ground up, from citizens who contact their representatives, journalists who write about the issue, donors who fund advocacy work, and professionals in law, policy, economics, and communications who bring their skills to the problem. You do not need a PhD in machine learning to matter here. At this stage the bottleneck is political far more than technical.

Joining the mailing list keeps you informed about developments and opportunities to act. Sharing the argument with people in your network extends its reach. Writing to your elected representatives signals that this issue has a constituency, which is how legislators decide what to prioritize. If you are considering a career move, AI policy, AI safety research, investigative journalism, and philanthropic work in this space are among the highest-impact roles available. And if you have resources, funding the organizations working on this is among the highest-leverage uses of philanthropic capital in the world right now.

¿Por qué esto es más importante que otros riesgos catastróficos?

El cambio climático, la preparación para una pandemia, las armas nucleares y la pobreza extrema presentan graves amenazas al bienestar humano y ninguna de ellas debe ignorarse. El riesgo existencial de la IA está en la parte superior de esa lista por una razón: conlleva una combinación específica de propiedades que ningún otro riesgo de la lista comparte.

First, speed: unlike climate change, which unfolds over decades with feedback loops that allow course corrections, a misaligned superintelligence operating at machine speed may not allow a corrective period. Second, irreversibility: most catastrophes, however devastating, leave behind the capacity to rebuild. An ASI that has achieved decisive strategic advantage may not. Third, compounding: a misaligned superintelligence would sit above the other catastrophes rather than beside them. It could cause or amplify them (accelerating climate harm, enabling biological weapons development, concentrating economic power) while simultaneously removing the human capacity to respond to any of them. That combination of speed, irreversibility, and systemic risk is what places it at the top of the priority order.

Si el riesgo es real, ¿por qué no se lo toma tan en serio como una guerra nuclear?

Start with the obvious difference: we have seen one of these threats and not the other. Nuclear weapons were demonstrated over Hiroshima and Nagasaki in the most unforgettable way imaginable, and everyone alive has grown up with the image of the mushroom cloud. Superintelligence has never harmed anyone. Its danger is still a projection on a graph, and people reliably discount a threat they have not watched happen.

Several things deepen that discount. Superintelligence pattern-matches to science fiction, so the warning gets filed next to the Terminator and waved off as a movie plot. The AI most people actually touch is a helpful, harmless chatbot, and that daily experience quietly argues against the alarm in a way nuclear physics never had to overcome. And where the atom bomb was built and held by a handful of governments, superintelligence is being built by companies worth trillions of dollars, with every commercial reason to keep the public mood calm.

There is a grim asymmetry hiding in the comparison. The nuclear taboo exists because Hiroshima came first and taught it, and the treaties and the drills came afterward, with enough of the world left to absorb the lesson. Superintelligence offers no such second step. If it goes wrong at full capability, the demonstration and the extinction are the same event. The whole case for treating it as seriously as nuclear war is that we have to do it before the proof arrives, not after.

How is the Nakada Foundation funded?

La Fundación es una iniciativa filantrópica con financiación privada. No recibimos financiación del gobierno. No aceptamos financiación de empresas de IA ni de organizaciones con intereses financieros en el ritmo del desarrollo de la IA. Nuestra independencia de los intereses comerciales de la IA es un requisito previo para nuestro trabajo, no un accesorio.

Una organización de defensa de la gobernanza de la IA financiada por las empresas que pretende gobernar es un presupuesto de relaciones públicas con una declaración de misión. Somos explícitos al respecto porque las estructuras de financiación de las organizaciones políticas de IA son enormemente importantes y con frecuencia quedan ocultas. Si está considerando apoyar nuestro trabajo, agradecemos la conversación. Puede comunicarse con nosotros a través del Contact page.

You are not in the minority.

Esta no es una preocupación marginal. En una encuesta nacional de 2025 realizada por el Future of Life Institute, el público se mostró firmemente del lado de la cautela.

64%

Dicen que la IA superinteligente no debería construirse hasta que se demuestre que es segura y controlable, o nunca debería construirse en absoluto.

73%

Quieren una regulación sólida de la IA avanzada. Sólo el 12 por ciento se opone a normas estrictas.

5%

Queremos que las empresas construyan superinteligencia lo más rápido posible, que es el rumbo que siguen los laboratorios líderes.

La voluntad del público ya está aquí. Lo que queda es hacer que nuestros líderes electos lo reflejen y aprueben una binding international treaty que impide que alguna vez se cree el arma de la superinteligencia artificial.