Arms control treaties begin with definitions. The Chemical Weapons Convention defines a chemical weapon as any toxic chemical or its precursor used as a method of warfare. The Nuclear Non-Proliferation Treaty defines nuclear weapons states as those that manufactured and exploded a nuclear weapon before January 1, 1967. The Strategic Arms Limitation Talks defined ballistic missiles by range and payload. These definitions were not politically neutral: each one reflected strategic choices about who would be covered, what activities would be constrained, and how the interests of the parties would be balanced.
AI governance faces a definitional problem with several features that have no close precedent in arms control. Understanding these features is the prerequisite for designing definitions that can actually govern what they are intended to govern.
Three approaches to defining dangerous AI
Current governance frameworks have used three distinct approaches to define which AI systems fall within their scope. Each approach has strengths and limitations that a treaty-level framework would need to address.
Compute thresholds define dangerous AI by the computing resources used to train a model, measured in floating-point operations, or FLOPs. The EU AI Act defines frontier AI models as those trained with more than 10 to the 25th power FLOPs. The US Executive Order on AI issued in October 2023 required reporting to the federal government for models trained with more than 10 to the 26th power FLOPs. The Frontier AI Safety Commitments signed at Seoul apply to models at or approaching frontier capability, with compute as a proxy for frontier status.
Compute thresholds have two significant advantages for treaty design. First, compute is observable. The hardware required to train large AI models, specialized chips produced by a small number of manufacturers, consumes distinctive patterns of energy and generates measurable heat. The supply chains for these chips are narrow and pass through jurisdictions amenable to monitoring. Second, compute is harder to fake than capability claims, because producing the hardware and running the training runs generates observable signatures that are difficult to conceal at scale.
The limitation of compute thresholds is that they become obsolete as algorithms improve. A model trained today with 10 to the 26th power FLOPs has substantially different capabilities than a model trained five years ago with the same compute, because algorithmic improvements produce more capable models per unit of compute over time. A compute threshold set today will either be too restrictive as efficient algorithms produce powerful models below the threshold, or too permissive as the frontier advances above it.
Capability-based definitions define dangerous AI by what the system can do rather than how it was trained. A system that can autonomously identify and exploit cybersecurity vulnerabilities above a defined severity threshold is dangerous under a capability-based definition, regardless of how much compute was used to produce it. A system that can provide meaningful assistance to bioweapons synthesis above a defined level of specificity is dangerous, regardless of its training compute.
Capability-based definitions are more precisely targeted at the harms they seek to prevent, but they create a significant measurement problem. Capability evaluations are not standardized. Different evaluation methodologies produce different results. Frontier AI developers currently conduct capability evaluations using proprietary methodologies that are not publicly audited. A capability-based definition in a governance treaty would require standardized, independently verifiable capability evaluation methods, which do not yet exist in forms that all major AI-developing states would accept.
The Chemical Weapons Convention combined a general-purpose criterion with scheduled chemical lists to cover both known and unknown dangerous agents. The general-purpose criterion defines chemical weapons as any toxic chemical used as a method of warfare, regardless of whether it appears on a list. The scheduled lists then establish graduated verification requirements for specific chemicals. This layered approach allowed the treaty to cover agents not yet known when it was negotiated. AI governance could use an analogous structure: a general-purpose criterion defining dangerous AI by the character of its effects, combined with scheduled capability lists for known dangerous applications.
Domain-based definitions define dangerous AI by the sector in which it is deployed rather than by its training compute or general capabilities. A system deployed for autonomous lethal decision-making in military contexts is dangerous. A system deployed in critical infrastructure control is dangerous. A system deployed for large-scale persuasion targeting democratic elections is dangerous.
Domain-based definitions have the advantage of clearly linking the governance obligation to the harm it prevents. They are also more politically tractable than general capability definitions, because states can accept constraints on AI in specific high-risk domains without conceding the principle that their AI programs generally are subject to international oversight. The significant limitation is that domain-based definitions allow a highly capable system to operate without governance obligations if it is deployed outside the defined domains, even if the same system could easily be redirected to dangerous applications.
The moving frontier problem
Arms control definitions are typically written against a stable technology. The range and yield of a ballistic missile, the toxicity of a scheduled chemical, the fissile mass required for a nuclear weapon. These do not change significantly over the life of the treaty. AI capabilities advance rapidly, meaning that a definition adequate today may be obsolete within years.
This creates what might be called the moving frontier problem: how to write treaty definitions that remain relevant as the technology develops without requiring constant renegotiation of the treaty itself. Three mechanisms have been used in arms control to address analogous problems.
Delegated definitional authority: the treaty establishes a governing body with authority to update technical definitions within defined parameters, subject to approval by the Conference of the Parties. The Organisation for the Prohibition of Chemical Weapons uses this approach for its scheduled chemical lists, which can be amended through the Conference of the States Parties without reopening the underlying treaty. This is faster than full treaty renegotiation but requires that parties trust the governing body with definitional authority.
Technology-neutral definitions: the treaty defines dangerous AI by the character of its effects rather than by specific technical parameters. A definition that covers any AI system capable of causing mass casualties above a defined threshold, or any system capable of autonomous action consequential enough to threaten critical national infrastructure, would remain relevant as the underlying technology changes because it describes outcomes rather than inputs. The limitation is that outcome-based definitions are harder to verify, because the relevant test is what a system could do, not what it demonstrably does.
Sunset and renewal provisions: treaty definitions are written with a defined review cycle at which they are reassessed and updated. The treaty enters into force with initial definitions, continues to operate during the review period, and is updated at each review conference. This is the approach the NPT uses for its periodic review conferences, though the NPT's definitions of nuclear weapons states have proved impossible to update through this mechanism because of the interests of the P5 in maintaining their privileged status.
The strategic interest dimension
Definitional negotiations are not purely technical exercises. Each definition choice has strategic implications for which states bear the greatest compliance cost and which retain the greatest freedom of action. These implications are visible to negotiating governments and shape their positions on definitional questions.
States with more advanced AI programs prefer broad definitions that impose equal constraint on competitors at lower levels of development. A compute threshold set above the current frontier of the most advanced programs imposes no constraint on the leading states while requiring all others to stay below the threshold indefinitely. States with less advanced programs prefer narrow definitions that allow them to reach current frontier levels before governance obligations apply, and object to definitions that would permanently disadvantage them relative to states that developed frontier AI first.
"The definitional negotiation is where most of the real work of an AI governance treaty happens. A definition that is politically achievable but technically inadequate produces a treaty that governs something other than the actual risk. A definition that is technically adequate but politically unachievable produces no treaty at all. The only path through this problem runs through shared technical evidence, which is why the AI Safety Institutes matter more than most people realize."
This strategic dimension is why the empirical work being done by the UK's AI Security Institute and the US Center for AI Standards and Innovation (both established as AI Safety Institutes) is foundational for treaty negotiations, even though those bodies were not designed with treaty preparation in mind. By producing shared evidence about what AI systems at various capability levels can and cannot do, and about which capabilities correlate with which risks, they create a factual basis for definitional negotiations that is harder for strategic interests to simply override. Agreement on facts does not automatically produce agreement on definitions, but it changes the character of the disagreement from a purely strategic one to a partly technical one, where shared evidence disciplines the negotiating positions of all parties.
The definitional work needs to begin before negotiations do. When arms control negotiations begin with no pre-existing technical consensus on definitions, negotiators tend to produce definitions that reflect the lowest common denominator of political agreement rather than the technical requirements of the governance goal. The more technical work that is done in advance, the stronger the foundation on which negotiators can build definitions that actually govern what they are meant to govern.
Common questions.
What is the compute threshold approach to defining dangerous AI?
The compute threshold approach defines dangerous AI by the computing resources used to train a model, measured in floating-point operations (FLOPs). The EU AI Act uses 10^25 FLOPs as the threshold for frontier AI systems; the US Executive Order on AI used 10^26 FLOPs as a reporting trigger. Compute thresholds have two advantages for governance: the hardware required is observable through supply chains and energy consumption, and it is harder to fake than capability claims. The main disadvantage is that algorithmic improvements reduce the compute needed for given capability levels over time, requiring regular threshold updates to remain relevant.
How did negotiators define chemical weapons in the CWC?
The Chemical Weapons Convention combines a general-purpose criterion with scheduled chemical lists. The general-purpose criterion defines chemical weapons as any toxic chemical or precursor used as a method of warfare, covering agents not yet identified when the treaty was negotiated. Scheduled lists then establish graduated verification requirements: Schedule 1 chemicals have few legitimate uses outside weapons; Schedule 2 have some legitimate uses but significant weapons potential; Schedule 3 have substantial legitimate uses requiring lighter verification. This layered approach could be adapted for AI governance, combining a general-purpose criterion defined by effects with scheduled capability categories for known dangerous applications.
Why do states disagree about what counts as dangerous AI?
Three underlying tensions drive disagreement. The dual-use problem: capabilities that enable dangerous AI are often the same capabilities that drive commercial value, so definitions that cover danger tend to constrain commerce, and countries with large AI industries resist broad definitions. The moving frontier problem: definitions written today become obsolete as the technology develops. And strategic interest: countries with more advanced AI programs prefer broad definitions that equally constrain competitors; countries with less advanced programs prefer narrow definitions that allow them to catch up before obligations apply.
What role do AI Safety Institute evaluations play in resolving definitional problems?
The UK's AI Security Institute and the US Center for AI Standards and Innovation (both established as AI Safety Institutes) are building an empirical record of what AI systems at various capability levels can and cannot do in specific domains. This record changes the character of definitional negotiations: rather than arguing purely from strategic interest about where thresholds should be set, negotiators can point to shared evidence about what capabilities at given levels actually enable. Agreement on facts does not automatically produce agreement on definitions, but it disciplines the negotiating positions of all parties in ways that pure strategic argument cannot. This is why the technical evaluation work is foundational for treaty negotiations, even though the Safety Institutes were not designed with that purpose.