Skip to main content
Chapter 60 · Risks and Choices

The Great Risks: Misuse, Accidents, and Existential Threats

The Technologies That Could End Humanity

Every chapter in this book has described transformative potential. This one describes what the same capabilities look like pointed the other way.

The symmetry is not rhetorical. A system that can propose novel protein structures for a drug target can propose them for a toxin. A model trained to find vulnerabilities so they can be patched finds vulnerabilities. Automated laboratories that compress a year of experiments into a week compress dangerous experiments at the same rate. In nearly every case the dual use is not a separate application built by different people—it is the same capability, and often the same model, asked a different question.

That property is what makes this chapter necessary and what makes the governance problem hard. Restricting the harmful use means restricting something inseparable from the beneficial one.

The purpose here is not alarm. It is calibration: which risks are severe, which are merely frightening, and which of the proposed responses would actually work.


2026 Snapshot — The Risk Landscape

What Is Already Happening

Some risks in this chapter are speculative. These are not.

AI-enabled fraud is operational and scaling. Voice cloning requires roughly a minute of audio and is being used routinely for impersonation scams against families and against corporate finance functions. This is not a warning; it is a current line item on fraud-loss reports.

Synthetic media has entered politics. Deepfaked audio of candidates has appeared in real elections across multiple countries. The most consequential effect so far has not been successful deception but the liar's dividend: the existence of convincing fakes gives anyone caught on genuine recording a plausible denial.

AI-assisted intrusion is a live capability. Models materially accelerate reconnaissance, phishing content, and vulnerability discovery. The defensive side benefits too, and it is genuinely unclear who is ahead—but the offensive gains are real and available to unsophisticated actors, which is the part that changes the threat landscape.

Autonomous weapons are deployed. Loitering munitions with automated target recognition have been used in several conflicts. The debate over whether to permit them has been overtaken by the fact of them.

What Is Credible but Not Yet Demonstrated

AI-assisted bioweapon design. Frontier labs now test their models specifically for uplift in this domain before release, which is itself informative: they consider the risk plausible enough to evaluate. Published evaluations have generally found limited uplift over what a determined actor could obtain from existing literature—but "limited" is not "none," and the trend runs the wrong way.

Loss of control of an autonomous system. No AI system has escaped meaningful human control. Systems have, however, pursued specified objectives in unintended ways often enough that the failure mode is well established at small scale.

Estimating the Unestimable

Expert surveys on catastrophic and existential risk produce wide dispersion. Toby Ord's estimate of roughly one in six for existential catastrophe this century is among the most cited and among the higher ones; surveys of AI researchers commonly produce median estimates of a few percent for AI-caused catastrophe, with enormous variance and a substantial minority at both extremes.¹

These numbers deserve a specific kind of skepticism. They are not frequencies derived from data—there is no reference class of civilizations to sample. They are structured intuitions, and they vary by an order of magnitude between serious people who have thought carefully. Treating any of them as a measurement is a mistake; treating the whole distribution as meaningless is a larger one. What the spread actually establishes is that informed people cannot rule the risk out, which is enough to justify effort without justifying panic.


Categories of Risk

The useful taxonomy distinguishes three failure modes, because they require different countermeasures.

Misuse is deliberate harm by a human actor: engineered pathogens, autonomous weapons aimed at civilians, mass disinformation, infrastructure attack. Perpetrators range from states to individuals. The governing dynamic is democratization—capability that once required a national program becomes available to a small group, then to one person. Countermeasures target access and attribution.

Accidents are unintended harm from systems working as built but not as intended: a pathogen escaping containment, an automated system triggering cascading infrastructure failure, an optimizer pursuing a specified objective through an unanticipated route. The governing dynamic is complexity—systems whose interactions exceed anyone's ability to model them fail in ways testing does not surface. Countermeasures target containment, monitoring, and reversibility.

Structural risk is harm from technologies functioning exactly as designed, reshaping society badly: surveillance enabling durable authoritarianism, power concentrating past the point of correction, automated persuasion degrading collective decision-making. Nothing malfunctions. Countermeasures are institutional rather than technical, and this category is consistently the most neglected because it has no incident to point at.

The fourth category, existential and catastrophic risk, cuts across all three. What distinguishes it is not probability but irreversibility. Ordinary cost-benefit reasoning assumes that mistakes can be corrected and that the expected value of many bets is what matters. Neither assumption holds for outcomes that end the ability to make further bets, which is why "low probability, high impact" understates the case: the relevant feature is that there is no second attempt.


AI-Specific Risks

Near-Term

The realistic near-term AI risks are not superintelligence. They are cheaper, faster versions of things that already go wrong.

Disinformation becomes economically trivial to produce, which shifts the constraint from generation to distribution. Cyber offense benefits from automation at every stage of the kill chain. Surveillance becomes comprehensive not because cameras improve but because analysis stops requiring human analysts—the binding constraint on mass surveillance was never collection, it was that someone had to watch the footage. Recommendation systems optimizing for engagement continue to select for outrage, because outrage is engaging.

None of these is exotic. All are extrapolations of present harms with the friction removed, and collectively they are more likely to matter this decade than anything in the next section.

The Alignment Problem

Alignment is the problem of getting a system to pursue what its principals actually want, including in situations nobody specified in advance.

It is hard for reasons that are not about AI at all. Humans do not fully know what they want. Stated preferences differ from revealed ones. Values conflict internally and across people. Any specification precise enough to optimize against will have edge cases where optimizing it produces something nobody intended—a problem familiar from every incentive scheme ever designed, and one that gets worse as the optimizer gets more capable.

Progress is real but partial. Reinforcement learning from human feedback made models substantially more usable and is now standard.⁸ Constitutional AI, which trains models against explicit written principles using AI-generated critique, reduces dependence on human labeling and makes the target values legible rather than implicit.⁹ Interpretability research has begun to identify meaningful internal structure in these systems rather than treating them as black boxes.

What none of these provides is a guarantee. They are techniques that measurably improve behavior on the distribution they were trained against. Whether they hold for systems substantially more capable than their overseers is unknown, and it is unknown in a way that cannot be resolved by testing on current systems.

Power Concentration

The most likely serious AI harm may involve no misbehavior at all.

Frontier capability requires capital, compute, and talent concentrated in a small number of organizations, mostly in two countries. Whoever controls those systems gains asymmetric advantage in intelligence analysis, persuasion, cyber operations, and economic production. This is a structural risk that arrives whether or not alignment is solved—an aligned system faithfully serving a narrow set of principals is precisely the concern.

Historically, illegitimate power has been constrained by the fact that it requires many cooperating people, and people can refuse. Automation weakens that constraint. A surveillance apparatus that once needed tens of thousands of informants needs considerably fewer.


Biological Risks

Biology is where the misuse case is most concrete, for a reason specific to the domain: biological agents self-replicate. A cyberattack requires continuous effort to propagate. A pathogen does not.

The barriers have fallen in every dimension. Gene editing that required a specialized laboratory is now accessible enough that basic CRISPR kits sell for around $150 and undergraduate courses perform gene editing routinely.³ DNA synthesis can be ordered commercially. The literature describing how pathogens achieve transmissibility and virulence is largely public, and the arguments for keeping it public are serious ones about scientific progress and defensive research.

The dual-use dilemma has no clean resolution. The 2011–2012 controversy over H5N1 transmissibility studies is the formative case: research demonstrating how avian influenza could become transmissible between mammals was defensively valuable and simultaneously a blueprint.² The resulting US funding moratorium was later lifted, then re-tightened. Two decades of debate have not produced a durable answer, which should be taken as evidence that a clean one may not exist.

Screening is the most effective available intervention and it is voluntary. The International Gene Synthesis Consortium coordinates screening of synthesis orders against dangerous sequences, and its members do this conscientiously.⁴ The gaps are structural: not all providers participate, benchtop synthesizers move capability outside the ordering system entirely, and screening against known sequences does not catch novel designs. This is the highest-leverage chokepoint in biosecurity, and it currently rests on an industry agreement rather than a legal requirement.

The Biological Weapons Convention has no verification mechanism. The 1972 treaty prohibits development, production, and stockpiling of biological weapons and has 183 parties.⁷ It contains no inspection regime, no enforcement provision, and no organization to administer it—verification protocols were negotiated and then abandoned. It is a norm rather than a control, and norms constrain states more than they constrain anyone else.

The relevant AI interaction is not that a model will design a pathogen unprompted. It is that AI compresses the expertise gap. The knowledge required to do this has always existed; what protected the world was that assembling it required years of specialized training. Tools that substitute for that training reduce the number of people capable of attempting it from a few thousand to a considerably larger number.


Risk Reduction

What Works

Chokepoints. The most effective interventions target physical bottlenecks rather than information. Synthesis screening works because DNA has to be manufactured somewhere. Compute governance is tractable because advanced chips are made in very few places. Information cannot be recalled; supply chains can be monitored.

Pre-deployment evaluation. Frontier labs now test models for dangerous capabilities before release, and the UK AI Safety Institute—established in 2023 as the first government body of its kind—conducts independent evaluation, with US and other national institutes following.⁶ This is the single most concrete governance development of the period. Its limitation is that evaluation catches capabilities that evaluators know to test for.

Defensive asymmetry where it exists. Some domains favor defense. Pathogen surveillance and rapid vaccine platforms benefit from the same biological AI that creates the threat, and mRNA platforms compress response time from years to months. Where defense scales faster than offense, accelerating defense is more effective than restricting offense.

What Is Weaker Than It Looks

Voluntary commitments. Labs have signed a series of safety pledges; the Asilomar AI Principles (2017) established the template.¹⁰ These have shaped norms and are not worthless. They are also unenforceable, unaudited, and abandoned under competitive pressure precisely when they would matter most.

Institutional capacity is thinner than the field's prominence suggests. The Center for AI Safety, MIRI, and the Alignment Research Center do serious work with small headcounts. The Future of Humanity Institute, which originated much of the framework this chapter uses, closed in 2024.⁵ The total number of people working full-time on technical alignment remains small relative to the number working on capability.

International coordination has no precedent that fits. Nuclear non-proliferation is the usual analogy and it is a poor one: fissile material is detectable, hard to produce, and useless for anything else. Model weights are copyable files, the inputs are dual-use, and the primary applications are civilian.


Second-Order Impacts

Security concerns reshape openness. Research norms built on free publication are being renegotiated under dual-use pressure. Each restriction is individually defensible and collectively they slow the defensive research that depends on the same openness.

Concentration is justified by safety. The argument that frontier development should occur only within a few well-resourced, safety-conscious organizations is genuinely reasonable and conveniently aligned with those organizations' commercial interests. Both things are true simultaneously, which makes the argument hard to evaluate and easy to abuse.

Attribution failure destabilizes deterrence. Deterrence requires knowing who attacked. AI-enabled operations are harder to attribute, and a world in which attacks cannot be traced is one where retaliation is either impossible or misdirected.

Preparedness decays predictably. Institutional attention to pandemic risk peaked in 2021 and has substantially receded, following the historical pattern exactly. The window for building defenses closes not when the risk passes but when attention does.


The Path Forward

Near-Term Likely (2026–2032)

Pre-deployment evaluation becomes standard practice and, in some jurisdictions, law. Compute thresholds serve as the regulatory trigger despite being a crude proxy.

Synthesis screening moves from voluntary to mandatory in at least one major jurisdiction—the most tractable and highest-value biosecurity action available, and one that could be taken immediately.

Incidents occur at a scale that demonstrates the risk without validating the worst forecasts: significant AI-enabled fraud, a serious infrastructure intrusion, an election meaningfully disrupted by synthetic media. Each produces a burst of regulatory attention.

No catastrophe. This is the modal outcome and it carries its own hazard, since a decade without disaster is readily interpreted as evidence that the concern was overblown.

Plausible (2032–2040)

A major incident—a costly AI system failure, a biosecurity breach, or a cyberattack with physical consequences—produces the first serious international governance response, as post-incident regulation always outpaces anticipatory regulation.

Alignment research matures into something closer to engineering, with measurable properties rather than empirical tuning.

Absolute risk rises even as management improves, because capability grows faster than governance.

Wild Trajectory (2040+)

The optimistic branch: alignment is substantially solved, governance holds, and the benefits described in this book are realized without catastrophe. This is a real possibility and deserves stating, since risk chapters tend to omit it.

The pessimistic branch: an engineered pandemic, a loss-of-control event, or an AI-enabled conflict causes harm at civilizational scale.

The most probable branch is neither: continued muddling, with close calls that are recognized only afterward, gradual institutional improvement, and permanent low-grade risk as the ordinary condition of a technologically capable civilization.


What Can Actually Be Done

If you build these systems: treat safety as a design constraint rather than a review stage. Red-team your own work with the assumption that it will be misused. Publish safety findings even when they are commercially inconvenient. And retain the judgment that some projects should not be built—a judgment that is worth very little unless someone occasionally exercises it.

If you influence policy: fund the unglamorous chokepoints. Mandatory synthesis screening, pathogen surveillance, and evaluation capacity in government are cheap, tractable, and effective. Build technical competence inside institutions rather than outsourcing assessment to the entities being assessed. Prepare response capability before it is needed, because it cannot be assembled during an emergency.

If you are neither: the honest answer is that individual action matters less here than in most chapters of this book, and pretending otherwise is condescending. What does matter is political attention. These risks are managed through institutions, institutions respond to sustained pressure, and sustained pressure is the one input ordinary citizens actually supply.


Conclusion

This book has been optimistic, and the optimism is warranted. The technologies described here can cure disease, decarbonize energy, educate people who have no access to teachers, and expand what humans are able to do. Those are not speculative benefits; several are already arriving.

But the same properties that make these technologies powerful make them dangerous, and not as a separate application—as the same capability, asked a different question. That is the structural fact this chapter exists to establish.

Three things follow from it.

The first is that the near-term risks are mundane rather than cinematic. Fraud, intrusion, surveillance, and manipulation will do more damage this decade than anything involving superintelligence, and they are being under-resourced precisely because they are boring.

The second is that the tractable interventions are known and unfunded. Mandatory synthesis screening, government evaluation capacity, and pathogen surveillance are not conceptually hard. They are cheap relative to the harms they address, and they are not being done at anything like the necessary scale.

The third is that irreversibility, not probability, is what should govern caution. Most technological mistakes get corrected. The subset that cannot be corrected deserves a different standard of evidence before proceeding—not because catastrophe is likely, but because the argument that recovers from an ordinary mistake does not apply.

The bet being made is that these systems can be built faster than they can be understood, and that understanding will catch up in time. That bet may well pay. It is being made by a generation that will not be present for most of the consequences, on behalf of one that has no say in it, and it is worth being honest that this is the arrangement.


Endnotes — Chapter 60

  1. Existential risk estimates vary enormously. Toby Ord's The Precipice (2020) estimates roughly a one-in-six chance of existential catastrophe this century. Surveys of AI researchers typically produce median estimates of a few percent for AI-caused catastrophe, with very wide dispersion. These are structured intuitions, not measurements.
  2. H5N1 transmissibility studies (2011–2012) demonstrated how avian influenza could become mammal-transmissible, sparking a formative debate over publication of gain-of-function research and leading to a US funding moratorium that was later lifted and subsequently re-tightened.
  3. CRISPR accessibility: basic gene editing kits are available for around $150, and undergraduate laboratory courses routinely perform gene editing—a capability that required specialized facilities a decade earlier.
  4. DNA synthesis screening: the International Gene Synthesis Consortium coordinates voluntary screening of orders against dangerous sequences. Coverage gaps include non-member providers, benchtop synthesizers, and novel sequences with no known signature.
  5. AI safety organizations: the Center for AI Safety (CAIS), the Machine Intelligence Research Institute (MIRI), and the Alignment Research Center (ARC) remain active. The Future of Humanity Institute (FHI), which originated much of the existential-risk framework, closed in 2024.
  6. The UK AI Safety Institute, established in 2023, was the first government body to conduct independent pre-deployment evaluation of frontier models; US and other national institutes followed.
  7. The Biological Weapons Convention (1972) prohibits development, production, and stockpiling of biological weapons and has 183 parties. It contains no verification mechanism, no inspection regime, and no implementing organization.
  8. RLHF (Reinforcement Learning from Human Feedback) is the primary method for aligning language model behavior with human preferences; it is standard practice and has known limitations, particularly regarding behavior outside the training distribution.
  9. Constitutional AI, developed at Anthropic, trains models to evaluate their own outputs against explicit written principles using AI-generated feedback, reducing reliance on human labeling and making target values legible.
  10. The Asilomar AI Principles (2017), signed by a large number of AI researchers, established the template for voluntary safety commitments in the field—influential in shaping norms, unenforceable by construction.