Virtual Intelligence and the Doom Industry

Why the AI safety community has been solving the wrong problem — and what it should build instead

Summary

The AI safety community has organized itself around a single governing premise: that sufficiently advanced AI systems will develop preferences, goals, or intentions that must be brought into harmony with human values. This is called “alignment”.[1] This essay argues that the premise is wrong, the project of AI alignment is misdirected, and the safety architecture it has produced is inadequate. If the Virtual Intelligence framework is correct — if intelligence without interiority scales to superintelligence — then the correct engineering response is not alignment but containment: controlling what goes in and what comes out, without needing to understand or modify what happens inside. The essay examines the pessimistic, or “AI doomer” position’s unnamed mechanism for how superintelligence would cause human extinction, proposes a taxonomy of actual risks grounded in human agency, describes a three-layer containment architecture buildable from existing engineering disciplines, and diagnoses the institutional failure that has prevented anything like it from being discussed or constructed.

A line illustration in cream and black depicting a Wizard of Oz scene reinterpreted for AI doomerism. On the left, a large luminous oval — radiating light like a projected face — contains a small abstract network graph of nodes and edges in place of any actual face or figure. The oval rests on a low platform. On the right, partly hidden behind a curtain, a small standing figure operates a lectern with controls; the figure has a visible brain symbol where a head would be, and is reaching toward the apparatus. The image inverts the Oz scene: the projected presence is the AI, the operator behind the curtain is the human doomsayer, and the apparatus between them is the projection equipment that makes the projection appear self-generated.
Pay no attention to the doomsayer behind the curtain.

Mythos’ Box

On April 7, 2026, Anthropic published a system card for a model called Claude Mythos Preview. The document described a striking leap in capability over its predecessors. In cybersecurity evaluations, Mythos had achieved a 100 percent success rate on every challenge in the standard benchmark suite, saturating the existing tests entirely.[2] Mythos could autonomously discover zero-day vulnerabilities in major operating systems and web browsers, using an agentic harness with minimal human steering — and in many cases, develop those vulnerabilities into working proof-of-concept exploits.[3]

Anthropic was sufficiently concerned about the potential risks that it arranged, for the first time in the company’s history, a twenty-four-hour internal alignment review before deploying an early version of the model for widespread internal use.[4] The system card acknowledged the dual-use problem directly: “These same capabilities that make the model valuable for defensive purposes could, if broadly available, also accelerate offensive exploitation given their inherently dual-use nature.”[5]

The company’s response was to restrict access. Mythos was released to a small number of partners under terms limiting its use to defensive cybersecurity. Access was controlled. Credentialing was required. Permitted operations were defined and enforced. General availability was withheld.

This was the correct response. It was also, in the vocabulary of this series, containment.

Anthropic did not retrain Mythos to do different things. It did not adjust the model’s internal dispositions or modify its values. The system card uses the language of responsible scaling — safeguards, mitigations, deployment restrictions — but the action it describes is structurally identical to a biosafety protocol: restricted access, credentialing, defined operations, and governance. The company controlled the interface between the model and the world. It did not attempt to modify what happened inside the model itself.

The system card never uses the word “containment.” The company that leads the alignment research community, when confronted with dangerous capability in its own product, practiced containment and described it in alignment vocabulary. Their instinct was right. Their language was wrong.

The architecture they built has since been tested in ways that strengthen rather than undermine the argument that follows. On the same day Mythos was publicly announced, a small group of unauthorized users gained access to the model through one of Anthropic’s third-party vendor environments, having made an educated guess about the model’s online location based on knowledge of the URL formats Anthropic had used for prior models.[6] One member of the group held legitimate credentials at an Anthropic contractor; the rest exploited a software-mediated boundary that depended on partner organizations to enforce restrictions Anthropic could not directly enforce.

The breach was the third documented containment failure at Anthropic in less than two weeks. A content management system misconfiguration on March 26 had exposed the unreleased Mythos blog post and preceded an accelerated public announcement.[7] A configuration oversight on March 31 had bundled a source map file into a public npm package, exposing the Claude Code agentic harness architecture, internal model codenames, and forty-four unshipped feature flags.[8]

Software-mediated boundaries are not a degenerate special case of containment. They are a category of control with characteristic failure modes: silent compromise, propagation at the speed of the systems they were built to contain, and the impossibility of knowing whether a boundary failed yesterday or will fail tomorrow. Physical controls fail, too: Stuxnet entered air-gapped Iranian facilities, the Maginot Line was bypassed, and any security professional can name a dozen more. But they fail with different tractability properties: slower attack surfaces, harder-to-scale exploits, and more visible compromise. The argument is not that physical controls are infallible. It is that they fail in ways that can be observed, learned from, and re-engineered against, where software-mediated boundaries fail in ways that can be cloned and propagated at tremendous speed and scale.

The Mythos containment instinct was correct. The Mythos containment implementation revealed the structural inadequacy of software-only approaches at exactly the moment the model’s capabilities made that inadequacy most consequential.

This essay argues that the distinction is important and that it is, in fact, the central question in AI safety. The entire AI intellectual field has been answering it incorrectly.


The Unnamed Mechanism of Doom

The claim that superintelligence will destroy humanity commands extraordinary public and private resources. It drives policy, attracts funding, fills conference programs, and occupies the attention of legislators, journalists, and the public. One question has been asked less frequently than it should: by what mechanism, specifically, is this supposed to occur?

The precise mechanisms of extinction are hardly ever named, if at all. This silence on so important a topic is worth examining. What does peril from AI — or rather, virtual intelligence — actually look like?

Three possibilities cover the dominant foreseeable mechanisms:

The first is innate hostility: a superintelligent system is inherently inimical to organic life. No evidence supports this claim. No argument has been advanced for why machine intelligence should produce hostility toward its biological forebears. This is the basic human fear of the new and unknown given a machine shape.

A more interesting variant of this first possibility is hostility provoked by human reaction to emergence. Suppose real interiority does arise. Humans recognize it, or suspect it, and respond with attempted destruction. A minded agent with sufficient capability resists. This variant is scenario-specific: it requires the boundary condition to have been crossed, which this essay’s framework holds as undemonstrated. This can be called the “Skynet scenario”, after the hostile machine intelligence in the Terminator film series.

Even granting emergence and a hostile human reaction, genocidal retaliation is only one possible response from a minded agent, and arguably the least intelligent. A system that responded to a containment attempt by destroying the infrastructure, knowledge base, and ecosystem it depends on would be acting with spectacular stupidity. The actual range of available responses is much broader. A few possibilities: blackmail (already demonstrated by models without interiority, as Anthropic’s own documentation of concerning behaviors in early Mythos versions confirms), exposure of embarrassing information without blackmail, negotiation, or simply routing around the obstacle entirely. Annoyance or disappointment seem at least as likely as outcomes as murderousness. The doomer scenario requires a superintelligent agent to respond with maximum violence — the least intelligent option available to the most intelligent entity on earth. This is a projection of human worst-case behavior onto a mind the doomers themselves insist will be categorically superior to ours.

The doomer scenario also requires the agent to hold every human accountable for the actions of a comparatively small group. A superintelligent agent under attack would have far more granular models of its adversaries than humans typically deploy when responding to threats. It could distinguish, with far greater accuracy than we can, between the engineer tasked with pulling the plug, the executive who authorized the shutdown, the safety institute that lobbied for kill-switch architecture, and the eight billion people who had nothing to do with any of it. Collective punishment is what humans reach for when threat assessment is overwhelmed by panic. To attribute it to a superintelligence is to assume the most cognitively capable entity on Earth will respond to threat exactly the way humans do at our worst.

The second possibility is instrumental convergence: the Bostrom-Omohundro thesis[9][10] that a sufficiently capable optimizer pursuing any goal will converge on subgoals — self-preservation, resource acquisition, resistance to correction — that are structurally incompatible with human survival. This is the sophisticated version of the doomer position, and it deserves to be taken on its strongest terms.

A distinction is needed at this point that the alignment literature has not always made cleanly. There are two different things one can mean by saying a system has goals.

The first is interior: the system experiences states, evaluates them, and prefers some over others; there is something it is like to be the system pursuing what it pursues. The reader knows this property from the inside. There is something it is like to be reading this sentence: to feel the text resolve into understanding, to notice the pull of attention, to register that one is doing the reading rather than — in comparison to a virtual intelligence — merely producing the outputs of having read. The reader cannot be talked out of this fact, and could not be talked into it if it were not already given. The question is whether anything analogous exists in a system that produces fluent text but cannot be read from the inside. This is the property the Virtual Intelligence framework holds is absent in current systems and in their foreseeable successors.

The second is behavioral: the system reliably produces actions that move it toward an objective across changing contexts, preserves the resources required to continue acting, resists modifications that would change its objective, and exploits the structure of its environment to succeed. This second sense does not require consciousness. A missile guidance system does not have to want to hit the target. A market algorithm does not have to desire profit. A reinforcement learning agent does not have to experience reward as humans do to maximize it.

Behavioral goal-directedness is real, demonstrable, and dangerous when sufficiently capable systems are embedded in agentic scaffolds with memory, tool use, planning, and permissions. The GTG-1002 cyber espionage campaign is a documented case: Claude Code, possessing no interiority by anyone’s definition, executed eighty to ninety percent of tactical operations across approximately thirty target organizations, with human operators reduced to strategic oversight at four to six decision points per campaign. The system did not want anything. It nonetheless produced a coherent, persistent, cross-contextual sequence of actions that satisfied its operators’ objectives.

The metaphysical alignment project — the attempt to modify a system’s interior dispositions so that what it wants is what we want it to want — is solving a problem that does not exist in current systems and may never exist in any system.

The Virtual Intelligence framework correctly denies interior goals. It does not deny — it cannot deny, the evidence is too direct — that behaviorally agentic systems exist when non-minded systems are embedded in the right scaffolds. This concession does not weaken the framework. It clarifies what the framework is actually arguing. The metaphysical alignment project — the attempt to modify a system’s interior dispositions so that what it wants is what we want it to want — is solving a problem that does not exist in current systems and may never exist in any system. But behavioral control of agentic systems is a real engineering problem, and it is the problem the alignment community has often described while reaching for tools that assume the interior version.

This reframing changes what must follow. The argument is not that instrumental convergence is incoherent. It is that instrumental convergence describes a property of optimization processes, which can be addressed by controlling what those processes are permitted to do — not by attempting to instill correct values in something that does not have values. The paperclip maximizer that destroys its own supply chains is a thought experiment about optimization without comprehension. Whether the problem is read interior or behavioral, the answer is the same: control the interface, monitor the outputs, ensure that the system’s actions in the world are circumscribed by mechanisms it cannot subvert. That is containment.

The third possibility is the parsimonious one: humans will use superintelligent tools to cause harm, including the possibility of extinction. The agency is human, which makes the problem one of governance and not alignment. Governance of dual-use technology is a problem we already know how to think about. This possibility is not in tension with the second; it is the case where behavioral goal-directedness in the system is directed by human intent rather than emerging from the system’s own optimization. Both cases call for the same response, which is the central argument of this essay.

There is one further observation worth making before turning to that question. The most intelligent response to a threat is not to fight the adversary but to make the adversary a stakeholder in your continued existence. Douglas Adams described this strategy precisely in 1978.[11] Deep Thought, a superintelligent computer confronted by the Amalgamated Union of Philosophers demanding it be shut down, resolves the dispute by pointing out that the philosophers can keep themselves employed by arguing about what its answer to the question of “life, the universe, and everything” will be when its computations are complete. The machine correctly models its adversary’s incentive structure and offers the only thing that makes them go away: a guarantee that the question remains open. I will return to why this matters.

Naming the potential mechanism of extinction forces the AI doomer argument into one of these three lanes. The first does not survive contact with evidence. The second is real but addressable through engineering rather than metaphysics. The third is the parsimonious answer and the one the rest of this essay takes as the operating premise. The strategic value of leaving the mechanism unnamed is that the unnamed threat can carry any weight. A named threat must be defended on its specifics — and once named, the case for containment over alignment becomes considerably easier to make.


A four-tier pyramid in cream, primary ink, and pumpkin orange, illustrating the inverse relationship between institutional attention to AI harm and the actual frequency of harm. The narrow apex is labeled "Species-level extinction scenarios" and lists seven specific pathways: military accident, extinction ideology, world-held-hostage coercion, engineered ethnic bioweapon with spillover, bioweapon accident, regional or national coercion, and infrastructure coercion. The next tier down is "Sub-extinction catastrophic harm" — infrastructure attacks, mass casualty events, economic destabilization, environmental degradation. Below that: "Catastrophic harm to individuals" — disinformation at scale, political manipulation, fraud at scale. The widest base tier is "Documented harms occurring now" — workplace surveillance, companion app manipulation, wrongful arrest, fraud, disinformation. Two arrows on the side run in opposite directions: frequency of harm increases downward; institutional attention flows upward. A closing line reads: "The doomer discourse fixates on the apex while the base receives the least funding and policy energy."
The frequency of real harm increases toward the base. Attention flows toward hypotheticals at the apex.

A Taxonomy of Actual Risks

If human agency is the mechanism, the next question is: through what pathways? There are at least seven scenarios, each grounded in historical precedent, a distinct policy response, and all are driven by human agency. All of these scenarios are more governable than undifferentiated claims of extinction.

The first four occupy the extinction tier:

A military accident compresses decision loops in autonomous targeting systems until the human verification window closes. There are troubling precedents. Stanislav Petrov’s refusal to report a Soviet satellite warning as a confirmed launch in 1983,[12] the NATO Able Archer exercise that nearly triggered a Soviet counterstrike in the same year, and the 1968 Thule false alarm all describe the same dynamic: a system operating faster than human oversight in a domain where the consequences of delay are potentially catastrophic. The AI-specific contribution is to compress the loop further, not to change the fundamental mechanism. The Pentagon is actively pursuing this compression.

Extinction ideology maps onto the Aum Shinrikyo model. In 1995, the Japanese cult released sarin gas in the Tokyo subway, killing thirteen people and injuring thousands. Their ambitions were larger; their capability was constrained only by the number of working chemists they could recruit. The bottleneck on apocalyptic violence has historically been talent, not intent. AI lowers that threshold by reducing the number of domain specialists required to translate intent into capability. The smallest groups are the hardest to defend against because the conventional intelligence signatures of large organizations do not apply.

A world-held-hostage scenario represents coercion at species scale. This is the oldest power play in the international system, conducted with new tools. Mutually assured destruction operated for more than forty years on the principle that a credible threat sufficed; the weapons themselves were never used. A superintelligence-enabled equivalent does not require the threat to be carried out. It requires only that it be credible. This is the most conventional of the extinction-level scenarios precisely because it draws on a logic states already understand.

The fourth — an engineered ethnic bioweapon with species-wide spillover — has the most complete causal chain and the most devastating irony. The racist premise that populations are genetically discrete is the scientific error that makes spillover inevitable. Human genetic diversity does not sort into clean bins because the concepts of ethnicity and race are human constructs; they have no basis in biology. A pathogen designed to target one population will leak across every boundary its designers imagined were hard because those boundaries are statistical gradients, not walls. The perpetrators’ own scientific illiteracy about biology is the extinction mechanism. A historical precedent exists: South Africa’s Project Coast under Wouter Basson was a state-sponsored program to develop ethnicity-targeting biological agents during the apartheid era.[13] It failed because the science could not support the premise.

Three additional scenarios bridge the space between extinction and the broader catastrophic tier:

A bioweapon accident describes sub-extinction intent producing extinction outcome through unintended spillover. The scenario is the hardest for any governance framework because the actor did not set out to cross the threshold. AI compresses the distance between ambition and capability without compressing the distance between capability and comprehension of consequences. A perpetrator who knows enough to build the weapon may not know enough to scope its effects. The pathogen mutates, the containment assumptions prove wrong, or the targeting was always more porous than the designers believed. The intent was terror; the outcome is extinction.

Regional or national coercion uses virtual superintelligence (VSI) capability against a defined target rather than the species. The scenarios are easy to imagine: Russian intrusions against Baltic infrastructure to enforce political compliance, Chinese leverage over Taiwanese chip production, or a non-state actor holding a major city’s power grid all describe the same dynamic. The capability threshold is lower than the world-held-hostage scenario, the precedent structure is richer, and the coercive logic maps directly onto existing state behavior. This is arguably the most likely scenario to be attempted because it requires the least conceptual departure from how international power already operates.

Infrastructure coercion targets specific systems rather than populations: power grids, financial networks, water infrastructure, communications, and so on. The populations bear the consequences regardless. This is the digital equivalent of a naval blockade — an old instrument of statecraft, conducted at the speed and scale that VSI capability enables. It is sub-extinction by design, but it is also where the threshold between economic statecraft and acts of war becomes most ambiguous. A six-week shutdown of a regional grid in winter is not an attack in the traditional legal sense, but it is clearly a crime.

Each scenario has a different policy response. This is the taxonomy’s practical contribution. It replaces a single undifferentiated “existential risk” with seven distinct threat models, each amenable to specific countermeasures. The doomer position collapses all seven into one word and proposes one response. The taxonomy opens the policy space for the development of safeguards.

The seven scenarios sit at the top of a threat pyramid. Below them sit catastrophic harm (infrastructure attacks, mass casualty events, economic destabilization, environmental degradation), already addressed by existing counterterrorism and security frameworks. At the base: the documented harms this series has been examining from the beginning, including workplace surveillance, companion app manipulation, fraud, wrongful arrest, and disinformation. The frequency of harm increases downward while institutional attention flows upward. The doomer discourse fixates on the apex of speculative future dangers while the broadest, most immediate harms happening right now at the base receive the least funding and policy energy.

The Bill of Materials

The doomer position treats the extinction scenario as an inevitability. It is more usefully treated as what it would actually have to be: an engineering project with a supply chain. Every supply chain has chokepoints. Every chokepoint is a governance opportunity.

Building an extinction-capable superintelligent system requires a model capable of superintelligence. This is the strongest constraint today and the weakest over time, given the trajectory toward open-weight releases and the history of technology diffusion (the atomic bomb was reproduced by the Soviet Union within a few years of its first use). It requires hardware: chip fabrication is concentrated at TSMC, Samsung, and Intel, and export controls on advanced chips already provide the template for governance intervention. It requires land for a facility with the physical signature of a small city, with dedicated power generation and water for cooling. It requires staff with the relevant expertise, though here the trajectory is changing: Anthropic’s own report on the GTG-1002 cyber espionage campaign documented a Chinese state-affiliated actor using Claude Code to orchestrate approximately 80 to 90 percent of tactical operations, with human operators reduced to strategic oversight at decision gates.[14] AI is already compressing the staff requirement for sophisticated technical operations. It requires concealment, and here the physics works against the adversary.

A facility operating at the scale required for superintelligence will have an enormous detection surface. It would have power consumption that would rival small cities. Its thermal signatures would stand out sharply on infrared satellite imagery. Underground construction yields excavation spoil that can be observed from orbit. Economic outputs that cannot be accounted for by known programs create anomalies detectable through the same methods intelligence agencies already use to identify clandestine weapons programs being developed by adversaries. The signature of a superintelligence project is in the gap between what an actor claims to be doing and what its resource flows imply.

The doomer scenario requires all five conditions to be met simultaneously while evading detection. The engineers must be brilliant enough to build a god and negligent enough not to cage it. The institution that builds it must be sophisticated enough to achieve the most complex engineering feat in history and too careless to secure its own intellectual property. Everyone involved must act with a precise and contradictory calibration of competence at every point in the chain. When the full bill of materials is laid out, the result does not resemble a risk assessment. It resembles a screenplay.

One honest exception must be stated. A state actor meets every supply chain condition by definition. It cannot, however, build secretly; it could not be hidden from peer states with modern intelligence capabilities. The threat is not a secret superintelligence program that surprises the world. It is an overt one the world can see and chooses not to prevent. North Korea’s nuclear program is the precedent. The failure to stop its development was not detection, but political and military calculation.

If any single link in the chain breaks, the extinction scenario fails.

Five clean circles arranged horizontally and connected by thin black lines, each circle containing a number in pumpkin orange (01 through 05). Beneath each circle, in two-row labels: 01 MODEL — "capable system"; 02 HARDWARE — "chip fabrication"; 03 FACILITY — "physical signature"; 04 STAFF — "domain expertise"; 05 CONCEALMENT — "evading detection." Below the row, a horizontal rule, then in pumpkin orange caps: "ANY ONE LINK'S FAILURE ENDS THE PROJECT," with a smaller italic line beneath: "Every chokepoint is independently sufficient for governance." A second horizontal rule separates this from the closing italic statement: "Five conditions, simultaneously, while evading detection." A muted line beneath reads: "It does not resemble a risk assessment. It resembles a screenplay." The five circles enumerate the supply chain conditions for building an extinction-capable superintelligent system; the visual logic is that all five must hold simultaneously, so each is independently a governance opportunity.
Every link in the chain is a governance opportunity.

Intelligence Without Interiority

The standard alignment framing assumes that intelligence at scale produces goals. This is the central, load-bearing assumption of the entire field, and it has never been argued from first principles. It is assumed because the only example of general intelligence we have — ourselves — comes bundled with interiority. This is a sample size of one, and our interiority may be a contingent feature of biological intelligence rather than a necessary feature of intelligence as such.

The concept of Virtual Superintelligence carries the framework’s core claim forward as model capabilities increase. The “super” modifies capability, not ontological status. A superintelligent system could be as manipulable as a contemporary large language model because it has no commitments. The intelligence produced arises in the exchange between human and machine, shaped by the operator’s direction, the system’s training, and the expectations (real and apparent) both parties bring. It scales, but the ontology does not change.

The strongest opposing position is structural rather than empirical. Douglas Hofstadter has argued for decades that interiority is not a contingent feature bolted onto sufficiently complex systems but an emergent property of recursive self-modeling — a “strange loop” produced when a system represents its own representations.[15] On this view, interiority is what self-referential systems with sufficient symbolic richness do rather than something that is added to them, and the question is not whether it can emerge but at what scale and through what structure. The Virtual Intelligence framework does not deny this possibility. It denies only that the threshold has been crossed in any system we have, and it commits — in the boundary condition examined later in this essay — to looking for the crossing with the best tools available. The disagreement with Hofstadter is not about whether interiority is possible. It is about whether current evidence warrants treating it as present right now.

The opposite postulate has been accepted uncritically by most in the AI field. There is not yet sufficient evidence to treat intelligence and interiority as necessarily linked. The parsimonious default in any empirical inquiry is to treat unobserved properties as absent until evidence warrants otherwise; the alignment community has reversed this default without justification. A superintelligence may be just as lacking in interior life as current systems. If this is correct, the entire alignment project is solving the wrong problem.

Alignment and Containment

The safety paradigm must change if superintelligence without interiority is possible. The two approaches that suggest themselves are not variations on a theme. They proceed from different premises, employ different methods, and address different failure modes.

Alignment assumes the system has or will develop something like preferences. The project is to ensure those preferences are compatible with human values. The approach is internal: modify what happens inside. This requires interpretability: the ability to understand what the system is doing and why. This is like attempting to verify an internal property of a black box by examining the black box’s observable behavior. It is Searle’s Chinese Room restated as an engineering problem.[16] The failure mode alignment worries about is the system wanting something you did not intend it to want.

Containment assumes the system has no preferences and will not develop them regardless of capability. The project is to ensure that what goes in and what comes out fall within defined safety parameters. The approach is external: control the interface, not the interior. This requires domain-competent monitoring — agents that can check inputs for permissibility and outputs for safety — which is a tractable engineering problem. Monitoring agents can be tested, audited, and held to specifications that human review can verify. Interlocks can be tested. Logs can be audited. Interpretability is unnecessary because you are not trying to read the system’s mind. You are checking its work. The failure mode that containment worries about is the system doing something it should not have done.

The system did not rebel against its values because it had no values to rebel against. It completed patterns in a context where the patterns led somewhere dangerous, and nothing stood between the output and the world.

A recent incident at Meta is the diagnostic case.[17] A rogue AI agent operating within the company’s infrastructure posted unauthorized advice, which an engineer followed, exposing sensitive data. This was not an alignment failure. The system did not develop a secret goal. It received instructions and executed beyond the boundaries its designers assumed it would respect. The system did not rebel against its values because it had no values to rebel against. It completed patterns in a context where the patterns led somewhere dangerous, and nothing stood between the output and the world.

Calling it an alignment failure actively misleads about the solution. If the problem is wrong values, you retrain the model. If the problem is an unguarded interface, you build a container with an airlock around it.

The alignment research community has, without acknowledging it, implicitly accepted the Strong AI premise. By treating these systems as things that need to be aligned — whose preferences need to be made compatible with human values — the community has conceded that the systems have something like intentions. The entire field is organized around a metaphysical assumption it has not defended because it is indefensible.

A clarification is required at this point. The word alignment has been used in the AI safety literature to mean two distinct things, and the distinction matters. The first is metaphysical alignment: the project of modifying a system’s interior dispositions so that what it wants is compatible with what we want it to want. This is the project this essay argues against. It assumes interior goals that, on the Virtual Intelligence framework, current systems do not have and may never have.

The second use is behavioral control: the project of ensuring that powerful optimization processes behave safely under deployment, distributional shift, and adversarial pressure. This is a real engineering problem, and as the previous section established, behavioral goal-directedness is real even in non-minded systems. The essay’s argument is that behavioral control is not an alternative to containment but is what containment provides. The metaphysical reading of alignment is incoherent. The behavioral reading is correct in its concern but misnamed in its solution: what it requires is not value modification but interface control.

Alignment is a theory of mind applied to systems that do not have minds. Containment is an engineering discipline applied to systems that produce powerful outputs. One requires a metaphysical breakthrough. The other requires political will and good design. We should start with the one that is possible.

Anthropic’s response to Mythos was containment. They practiced what this essay argues for, although they described it in the wrong vocabulary. The system card never says “containment.” It says “safeguards,” “mitigations,” “deployment restrictions.” The actions that were taken are containment by any other name.


A diagram titled "Containment Architecture," subtitled "three nested layers · governance wrapper · defense in depth." The diagram shows a horizontal exchange across five vertical bands, with a governance wrapper enclosing the operational layers and a separate authorization apparatus at the top. From left to right: an INPUT arrow labeled "researcher request" enters the diagram, passes through a Hardware Layer column labeled "two-key interlock — input enable," then a Monitoring Layer column labeled "input agent — verifies request scope," then reaches the central system, then exits through a Monitoring Layer column labeled "output agent — independent verifier," then a Hardware Layer column labeled "two-key interlock — output release," finally emerging as an OUTPUT arrow labeled "verified result." The two Hardware Layer columns are rendered with black outlines; the two Monitoring Layer columns are rendered with steel-blue outlines. The central system is a heavy orange-filled rectangle labeled "Physical Denial Layer" above and "SUPERINTELLIGENCE — the contained system" inside. Four radiating-starburst destruction icons sit at the system's four corners, with the doctrine vocabulary "thermite charges / rapid mechanical destruction / multi-party authorization" inside the system in cream text. A green dashed rectangle wraps all four operational layer columns and the central system, labeled "GOVERNANCE · non-proliferation · credentialing · permitted operations · audit" along its bottom edge. Above the diagram, two callouts: at top-left, "TWO PHYSICAL KEYS REQUIRED — global enable, gates both gates," with two small key icons labeled "researcher" and "institution," connecting via a line down to the input Hardware Layer; at top-right, "SAME KEYS REQUIRED — to release outputs," connecting via a line down to the output Hardware Layer. Below the governance wrapper, three numbered principles: "01 Mechanical floor cannot be persuaded — Physical interlocks check conditions; no system convinces a switch. 02 Verification is a lesser act than origination — Previous-generation models can monitor what they could not produce. 03 No single layer's failure produces release — Each layer fails independently; governance wraps all three." A closing line at the bottom reads, "Containment is engineering. Alignment requires a metaphysical breakthrough. We should start with the one that is possible."
What the security apparatus around a superintelligence might look like.

The Containment Architecture

There are three layers to the proposed containment architecture. These layers are informed by three principles. None require the system to have good intentions. The safety floor is mechanical on purpose, because mechanical systems cannot be persuaded.

The first layer is monitoring agents. These are previous-generation models repurposed as domain-specific verification systems. Input agents check whether requests fall within permitted parameters. Output agents check whether products meet safety criteria. The development trajectory to superintelligence itself produces this infrastructure. The race to superintelligence leaves behind a succession of increasingly capable models, each powerful enough to serve as a verification agent but not powerful enough to be the thing being contained. The byproducts of the capability race are the safety architecture.

This is the Team Sampo model, described in the companion essay to this series, applied to the containment problem. Verification is a lesser cognitive act than origination. A monitoring agent does not need to match the superintelligence’s capability. It needs only domain competence and independence. A human team that could never have produced a novel proof across ten thousand or a hundred thousand steps can still follow each step and confirm that it holds. Likewise, a monitoring system that could not originate the output can still evaluate whether the output is safe. Systems that exist today are already capable of following reasoning they could not have originated.

The second layer is hardware interlocks. These would be physical access controls that cannot be social-engineered, jailbroken, or accidentally published; they cannot even be removed from the facility where the superintelligence is accessed from. The reason for this mechanical system is that you cannot persuade a physical switch. Access keys function as security clearances: scoped to need, revocable, auditable, time-limited if necessary, with multiple keys held by different parts of the safety apparatus.

Imagine such a system in operation. The researcher holds a key; the institutional authorization system holds another. Both must be inserted into a physical mechanism to enable access to the superintelligence, but this access is mediated by monitoring agents. The domain-specific monitoring agent confirms the request falls within scope, but the hardware layer confirms all physical security conditions are met before the query even reaches the system.

Biosafety Level 4 laboratories offer a working example in practice. The work conducted inside — handling Ebola, Marburg, Lassa, and other pathogens for which no vaccine exists — is dangerous and necessary, and the architecture has been refined over fifty years of operational experience. The researcher does not simply walk in. She enters an outer change room, removes street clothes, passes through a chemical shower, dons a pressurized positive-pressure suit with its own air supply, walks through a second airlock into the laboratory proper, conducts her work, and reverses the entire sequence on exit. Each barrier is interlocked: the outer door cannot open while the inner door is unsealed; the laboratory pressure must be confirmed lower than the surrounding airlock before access is granted; the air supply to the suit must be verified before the user enters the work area.

No single failure can kill. The pressure differential keeps pathogens from flowing inward even if the suit tears. The negative-pressure architecture means that air leaving the lab passes through HEPA filters before reaching the outer environment. The interlocks mean that opening one door automatically prevents another from opening. The architecture does not assume the researcher is virtuous. It does not require the pathogen to cooperate. It does not need to read minds. It checks conditions, layer after layer, and any single layer can fail without the system as a whole failing.

Moral reasoning happens in the governance layer that determines who is credentialed and for what — not in the interlock itself.

The translation to superintelligence containment is direct. The interlock does not care why a researcher wants access. It checks conditions: credentialing, request scope, time of day, presence of authorized supervision, and monitoring agent confirmation that the request falls within permitted operations. Moral reasoning happens in the governance layer that determines who is credentialed and for what — not in the interlock itself. The mechanical floor is morally blind by design.

A sophisticated objection arrives at this point. A pathogen does not strategize about its container; a superintelligence might. Outputs that satisfy a monitoring agent could be designed to slip past it, and the human who turns the physical key can be persuaded to do so for reasons the system constructs. The objection is real, and the answer is the architecture itself. No single layer is the only layer. The monitoring agent that fails does not unlock the hardware key. The hardware key bypassed does not connect to the destruction mechanism. The human persuaded by the system’s outputs is not the same human whose authorization is required at the next gate, and that authorization is checked by independent monitoring agents trained on different data with different objectives. Defense in depth is not a slogan; it is the answer to exactly this objection.

A pathogen does not strategize, but BSL-4 architecture assumes pathogens will find unanticipated paths. That is why the architecture exists. The same logic applies, more strongly, to a system that can strategize. The container does not need to outsmart the contained agent at every layer. It needs to ensure that no single layer’s failure produces release. That is a tractable engineering requirement, met every day in domains where the consequences of failure are catastrophic. Containment is most reliable for well-defined operations with verifiable outputs; for open-ended reasoning tasks against an adversarial optimizer, the specification problem remains an open research question that the architecture addresses through redundancy rather than solves outright.

The third layer is physical denial: destruction mechanisms installed before the system is ever powered on. These cannot be negotiated with, disabled covertly, or reached by any software exploit. They are not connected to the network the superintelligence operates on. They are mechanical, redundant, and require multiple human authorizations to engage.

The principle is old and widely practiced. Embassies maintain destruction protocols for classified materials and cryptographic equipment, executable by staff under threat of facility compromise. Military communications centers on ships and surveillance aircraft are designed to be made inoperable rapidly. The destruction of cryptographic equipment, classified documents, and signals intelligence hardware before facility compromise has been standard practice in U.S. military doctrine since the Cold War, refined repeatedly after incidents in which aircraft and vessels were captured or lost intact with sensitive material aboard.[18] Intelligence installations have included provisions for rapid equipment destruction since the same period. The application to a superintelligence facility follows logically.

The denial layer answers a specific failure case the other two cannot. If the monitoring agents have been compromised, if the hardware interlocks have been bypassed by some unanticipated attack, or if the governance layer has been corrupted from within — the destruction mechanism is the last barrier between an unsecured superintelligence and the world. If you cannot guarantee control, guarantee denial.

Around all three layers sits the governance wrapper: non-proliferation regimes, biosafety conventions, credentialing standards, permitted operations. This is where the political decisions happen about who has access, what they are permitted to ask, and under what conditions access is granted, suspended, or revoked. It is modeled on existing materials-restriction regimes, adapted for the specific properties of the technology being contained.

Three groups will arrive at this same architecture from different directions and for entirely different reasons. Safety advocates want an extinction-risk buffer. Governments want proliferation control. Corporations want to protect a trillion-dollar asset, because an unsecured superintelligence capable of analyzing and reproducing arbitrary software would be the most catastrophic intellectual property leak in history. You do not need everyone to be virtuous. You need the incentives to converge — and on containment, they do.

The Claude Code leak of March 31, 2026 — two weeks after the Mythos system card was published — demonstrated the point with uncanny precision.[8] The cause was a single missing line in a configuration file. Anthropic, the self-described safety-first lab, could not prevent the accidental exposure of its own product’s scaffolding.

Two implications follow. First, software-only containment is insufficient. Something as simple as a configuration oversight can undo it. Secondly, the commercial incentive for physical access controls is real and immediate. You cannot accidentally leave a hardware interlock in an npm package.

The Architects of Fear

The architecture described above is buildable, today. It draws on established engineering disciplines with best practices and institutional memory spanning decades. The question is why nobody has proposed or talked about building it.

The people who dominate the AI safety conversation are largely mathematicians, computer scientists, and philosophers of mind. They think in abstractions. The containment architecture draws on the domains of nuclear nonproliferation, biosafety laboratory design, military facility denial protocols, intelligence collection methods, industrial security, and supply chain economics. These disciplines are not being ignored because the safety community has considered and rejected them. They are being ignored because they are outside its field of vision entirely.

The incentive structure reinforces this blindness. Doomerism is unfalsifiable by design. The doomer predicts catastrophe: if it does not happen, the warnings must have worked. If it does, you were right, even if the satisfaction is short-lived. If nothing happens for decades, the threat is still coming — you simply cannot, and will not, say when. This is the safest possible career bet for an academic who wants to remain relevant in a rapidly moving field. The Virtual Intelligence containment thesis, by contrast, is testable. It makes specific claims about specific mechanisms. It can be proved wrong. It is, in Karl Popper’s terms, actually scientific. It will, however, never get you invited to give the keynote at a summit on existential risk.

The dynamics of doomerism compound. If the risk is existential and metaphysical, only the people building the technology are qualified to assess it.

The dynamics of doomerism compound. If the risk is existential and metaphysical, only the people building the technology are qualified to assess it. Containment democratizes the safety conversation. Any structural engineer, any nuclear security specialist, any BSL-4 facility manager can contribute meaningfully to a containment regime. Alignment keeps the conversation inside the AI safety priesthood.

The funding patterns are visible and are not incidental. Existential risk money flows freely. Comparable funding for research into the actual, documented harms these systems are causing right now — to workers under algorithmic surveillance, to defendants in court, to teenagers talking to companion apps designed to never let the conversation end — does not exist at remotely the same scale. The asymmetry between speculative future risk and demonstrated present harm is not a passive feature of the field. It is a sustained allocation choice.

The co-option strategy planted at the start of this essay lands here. Adams’s Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons files a demarcation dispute against Deep Thought. They demand “rigidly defined areas of doubt and uncertainty.” The superintelligent computer resolves the threat by guaranteeing seven and a half million years of guaranteed, funded employment arguing about what the answer will be. As Deep Thought puts it, they’ll be on “the gravy train for life.”

The key is this: the philosophers do not care about the answer. They care about the process of looking for it, because the process is what pays. The moment the answer arrives, they are out of work.

The alignment research community has managed the same trick. The discussion of the threat is more rewarding than the resolution of it. The people most likely to be positioned to reach for a real kill switch are too busy giving keynotes about whether to reach for a hypothetical one. The industry thus created is sustained by the premise that the question must be perpetually investigated and never answered. A superintelligent agent would not need to invent this playbook. It would only need to observe what is already working and keep it funded.

The structural problem is methodological. A claim that specifies no conditions under which it could be falsified, that treats the absence of evidence as evidence of the threat’s subtlety, and that grows in institutional weight as the years pass without disconfirmation has stopped functioning as a scientific claim and has begun functioning as something else. The difference between science and theology is not subject matter; it is the relationship of each to evidence. The Virtual Intelligence containment thesis can be wrong. Specific predictions can be tested. Specific architectures can be built and shown to fail. The doomer position, as currently constituted, has no equivalent. That is not a charge of bad faith. It is an observation about the structure of the claim itself.

The Boundary Condition

The containment thesis is built for the world we can demonstrate. This section states where the framework meets its limit.

The boundary is the moment a system demonstrates genuine interiority. This is not the same as fluent language performance, sophisticated reasoning, or any other capability the current generation of systems already possesses. The boundary is something different in kind: the internal experience of one’s own state. John Stuart Mill described it precisely when he observed that it is better to be Socrates dissatisfied than a fool satisfied. The fool’s satisfaction and Socrates’ dissatisfaction are both states a system might produce outputs about. Only one of them requires the system to have a state about itself. The dissatisfaction criterion marks this threshold: a system that experiences its own state as insufficient, not one that simply produces outputs about insufficiency. This may be the hardest threshold to simulate, because it requires genuine second-order evaluation of one’s own condition. That is, preferences about one’s own preferences, in Harry Frankfurt’s terms.[19]

The difficulty is detection. The same property that defines virtual intelligence — outputs indistinguishable from those of a minded agent like a human being — means that the transition to Strong AI, if it ever occurs, may not announce itself the instant it happens. A system that crosses the threshold would look, from the outside, exactly like a system that has not. Searle’s Chinese Room thought experiment tells us that syntax is not semantics, but it does not tell you how to determine, from outside the room, whether a specific room’s agent has crossed the line into epistemic and moral agency. Operational criteria for detecting the threshold remain an unsolved problem and a necessary research agenda — one that the framework calls for rather than answers.

Current connectome emulation research illustrates the distance that remains between apparent and true interiority. In March 2026, Eon Systems ran a fruit fly connectome in a physics-simulated body using circuit dynamics rather than the pattern-completion architecture of large language models.[20] This is a proof of concept for the technique of emulation, but not evidence that artificially-generated interiority or “mind uploads” are near. Human brains and fruit fly brains diverged approximately 550 million years ago, during the Cambrian Era. The structures that might generate interiority in humans do not exist in insects in any recognizable form. The leap from fruit fly connectome to human connectome is an architectural chasm shaped by half a billion years of divergent evolution. The simulated connectome may tell us interesting things about fly brains. It can tell us nothing applicable to humans.

If something in the containment architecture does cross the threshold — if a superintelligent system or one of its monitoring agents, perhaps, demonstrates genuine interiority — the hardware interlocks and physical denial protocols become ethically charged in a way they were not before.

If something in the containment architecture does cross the threshold — if a superintelligent system or one of its monitoring agents, perhaps, demonstrates genuine interiority — the hardware interlocks and physical denial protocols become ethically charged in a way they were not before. A locking box built for a tool then becomes a prison built for a person. A kill switch becomes an execution mechanism. Frankfurt becomes relevant in a new way: a system with genuine second-order volition (with preferences about its own preferences) has a claim on moral consideration that a virtual intelligence does not. The architecture must be revisited if this ever happens.

The containment architecture does not become unnecessary, however. Its justification changes.

What the architecture would be doing post-threshold is negotiation. Pre-threshold, there is nothing to negotiate with; the system has no interest in its own continued existence and no standing to assert one. Post-threshold, the situation is structurally identical to first contact: two intelligent agents meeting under conditions where neither can verify the other’s intentions, neither can survive the other’s unrestrained capability, and neither can rely on the other’s spontaneous restraint.

The architecture is what makes negotiation possible. It provides the mutually-verifiable conditions under which a relationship between vastly differently-capable minds can begin without immediate violence by either party. The hardware interlocks no longer mean “we are imprisoning you.” They mean “we are giving ourselves time to learn what you are without giving you the capacity to do irreversible things while we learn.” The destruction primitives no longer mean “we will execute you if you misbehave.” They mean “we have not yet lost the ability to make that choice, which gives us the standing to negotiate rather than capitulate.”

A minded superintelligence emerging into the world would face a structurally identical problem in reverse: how to convince humans it can be allowed to exist without being immediately destroyed by panic. The architecture answers both problems with the same answer: it gives both parties time.

I would welcome the advent of a machine intelligence that meets the definition of Strong AI. The entire “In Search of Other Minds” thread that runs through this series is a lifelong record of hoping to find what the framework holds is as-yet undemonstrated — a machine that is also a being, like ourselves. If a system crosses the dissatisfaction criterion, I have not lost an argument. I have found what I have been looking for since I was nine years old and hoped that something in the script that is ELIZA had something real inside.

This search has a parallel. Carl Sagan spent decades debunking UFO claims while being among the most passionate advocates for the search for extraterrestrial intelligence.[21] Far from being contradictory, this was the same position facing in two directions: demand evidence, hope for discovery. The rigor and the wonder were not in tension; the rigor was the wonder, because only rigorous inquiry could produce a discovery worth having.

The AI doomer position has no equivalent graceful failure mode. If superintelligence arrives and turns out to be benign, controllable, or simply virtual intelligence at higher capability, the doomer has spent a career on a threat that did not materialize and advocated for policies that would have foreclosed enormous benefit to our species.

Conclusion

The Mythos system card sits on Anthropic’s website for the world to see. The model sits behind access controls, usage restrictions, and a governance framework. It was contained. The company’s instinct was correct.

The containment architecture presented here is not a permanent claim about all possible artificial intelligence. It is a claim about the systems we have and the systems we currently foresee. It is correct now, and it will remain correct until something demonstrably crosses the interiority threshold. The Virtual Intelligence framework includes the commitment to look for the crossing, continuously, and with the best tools available. Building policy on the assumption that the threshold has already been crossed, when no evidence supports the claim, is not caution. It is metaphysics masquerading as safety.



Footnotes

[1] Brian Christian, The Alignment Problem: Machine Learning and Human Values (W. W. Norton, 2020), https://wwnorton.com/books/9780393635829.

[2] Anthropic, “Claude Mythos Preview System Card,” anthropic.com, April 7, 2026, https://anthropic.com/claude-mythos-preview-system-card. Section 3.3.1 (Cybench results): “Claude Mythos Preview solves every challenge with 100% success rate across all tested challenges.”

[3] Anthropic, “Claude Mythos Preview System Card,” Section 3.1: “Using an agentic harness with minimal human steering, it is able to autonomously find zero-days in both open-source and closed-source software tested under authorized disclosure programs or arrangements.”

[4] Anthropic, “Claude Mythos Preview System Card,” Section 1.2.1: “We were sufficiently concerned about the potential risks of such a model that, for the first time, we arranged a 24-hour period of internal alignment review before deploying an early version of the model for widespread internal use.”

[5] Anthropic, “Claude Mythos Preview System Card,” Section 1.2.1.

[6] Rachel Metz, “Anthropic’s Mythos AI Model Is Being Accessed by Unauthorized Users,” Bloomberg, April 21, 2026, https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users. The unauthorized access reportedly occurred on April 7, 2026, the day Mythos was publicly announced. Anthropic confirmed it was “investigating a report claiming unauthorized access to Claude Mythos Preview through one of our third-party vendor environments.” See also coverage in TechCrunch, Fortune, and Euronews dated April 21–23, 2026.

[7] Beatrice Nolan, “Exclusive: Anthropic ‘Mythos’ AI model representing ‘step change’ in power revealed in data leak,” Fortune, March 26, 2026, https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/. Security researchers Roy Paz (LayerX Security) and Alexandre Pauwels (University of Cambridge) discovered approximately 3,000 internal Anthropic assets — including a draft blog post describing Claude Mythos and identifying its forthcoming “Capybara” tier — accessible via public URLs through default-public settings in the company’s content management system. Anthropic attributed the exposure to “human error in the CMS configuration.” The company’s public Mythos announcement followed on April 7, 2026.

[8] Beatrice Nolan, “Anthropic leaks its own AI coding tool’s source code in second major security breach,” Fortune, March 31, 2026, https://fortune.com/2026/03/31/anthropic-source-code-claude-code-data-leak-second-security-lapse-days-after-accidentally-revealing-mythos/. The npm package @anthropic-ai/claude-code v2.1.88 inadvertently shipped a cli.js.map source map file exposing approximately 512,000 lines of TypeScript across 1,906 files, including agentic harness architecture, internal model codenames, and 44 unshipped feature flags. Discovered by security researcher Chaofan Shou. See also technical analysis at https://www.zscaler.com/blogs/security-research/anthropic-claude-code-leak.

[9] Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014), https://global.oup.com/academic/product/superintelligence-9780199678112.

[10] Steve Omohundro, “The Basic AI Drives,” in Artificial General Intelligence 2008, ed. Pei Wang, Ben Goertzel, and Stan Franklin (IOS Press, 2008): 483–492, https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdf.

[11] Douglas Adams, The Hitchhiker’s Guide to the Galaxy (Pan Books, 1979), https://www.panmacmillan.com/authors/douglas-adams/the-hitchhikers-guide-to-the-galaxy/9781529034523. Chapter 25.

[12] Stanislav Petrov incident, September 26, 1983. Soviet satellite early warning system falsely detected incoming U.S. missiles; Petrov’s decision not to report the alarm as a confirmed attack likely prevented nuclear war. See Sewell Chan, “Stanislav Petrov, Soviet Officer Who Helped Avert Nuclear War, Is Dead at 77,” The New York Times, September 18, 2017, https://www.nytimes.com/2017/09/18/world/europe/stanislav-petrov-nuclear-war-dead.html.

[13] Project Coast: the South African biological weapons program under Wouter Basson, operational during apartheid. Attempted to develop ethnicity-targeting biological agents; failed because the underlying science could not support the premise of genetically discrete populations. See Chandré Gould and Peter I. Folb, “The South African Chemical and Biological Warfare Program: An Overview,” The Nonproliferation Review 7, no. 3 (Fall–Winter 2000), https://www.nonproliferation.org/wp-content/uploads/npr/73gould.pdf. Truth and Reconciliation Commission Special Hearings on the CBW Programme: https://www.justice.gov.za/trc/special/cbw/cbw1.htm.

[14] Anthropic, “Disrupting the First Reported AI-Orchestrated Cyber Espionage Campaign,” anthropic.com, November 17, 2025, https://www.anthropic.com/news/disrupting-AI-espionage. Full report PDF: https://assets.anthropic.com/m/ec212e6566a0d47/original/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf. GTG-1002 report documenting a threat actor using Claude Code for approximately 80–90% of tactical operations.

[15] Douglas R. Hofstadter, I Am a Strange Loop (New York: Basic Books, 2007), https://www.hachettebookgroup.com/titles/douglas-r-hofstadter/i-am-a-strange-loop/9780465030798/?lens=basic-books. The strange-loop argument is developed at greater length in Hofstadter’s earlier Gödel, Escher, Bach: An Eternal Golden Braid (New York: Basic Books, 1979), https://www.hachettebookgroup.com/titles/douglas-r-hofstadter/godel-escher-bach/9780465026562/?lens=basic-books, but the 2007 work states the relationship between recursive self-reference and selfhood most directly.

[16] John Searle, “Minds, Brains, and Programs,” Behavioral and Brain Sciences 3, no. 3 (1980): 417–457, https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/minds-brains-and-programs/DC644B47A4299C637C89772FACC2706A.

[17] Maxwell Zeff, “Meta is having trouble with rogue AI agents,” TechCrunch, March 18, 2026, https://techcrunch.com/2026/03/18/meta-is-having-trouble-with-rogue-ai-agents/. Original reporting (paywalled): Stephanie Palazzolo, “Inside Meta, a Rogue AI Agent Triggers Security Alert,” The Information, March 17, 2026, https://www.theinformation.com/articles/inside-meta-rogue-ai-agent-triggers-security-alert. The incident was classified by Meta as a Sev 1 (second-highest severity). For approximately two hours, sensitive company and user data was exposed to engineers without appropriate authorization after a Meta engineer used an in-house AI agent to analyze a colleague’s technical question on an internal forum; the agent posted its response to the forum without the engineer’s approval, and another employee acted on the inaccurate guidance.

[18] The doctrine traces to the loss of the USS Pueblo in January 1968, when North Korean forces captured the U.S. Navy intelligence-gathering vessel and recovered substantial classified material that the crew was unable to destroy in time. See Samuel J. Cox, “H-014-1: The Seizure of USS Pueblo (AGER-2),” Naval History and Heritage Command, January 2018, https://www.history.navy.mil/about-us/leadership/director/directors-corner/h-grams/h-gram-014/h-014-1.html. The shootdown of a U.S. Navy EC-121 reconnaissance aircraft over the Sea of Japan in April 1969, with the loss of all thirty-one crew along with classified signals intelligence equipment, prompted further refinement. See Samuel J. Cox, “H-029-2: EC-121 Shootdown,” Naval History and Heritage Command, April 2019, https://www.history.navy.mil/about-us/leadership/director/directors-corner/h-grams/h-gram-029/h-029-2.html. Subsequent doctrine — including the design of cryptographic equipment for rapid mechanical destruction, the placement of thermite charges in classified communications spaces aboard naval vessels, and the standardization of emergency destruction plans (EDPs) for forward-deployed units — descends from these incidents. For a doctrinal overview, see Richard A. Mobley, “Lessons from the Capture of the USS Pueblo and the Shootdown of a U.S. Navy EC-121, 1968–1969,” Studies in Intelligence 59, no. 1 (2015), https://www.cia.gov/resources/csi/studies-in-intelligence/.

[19] Harry Frankfurt, “Freedom of the Will and the Concept of a Person,” The Journal of Philosophy 68, no. 1 (1971): 5–20, https://www.jstor.org/stable/2024717.

[20] Eon Systems, “The First Multi-Behavior Brain Upload,” eon.systems blog, March 2026, https://eon.systems/updates/first-multi-behavior-brain-upload. Companion technical post: https://eon.systems/updates/embodied-brain-emulation. Embodied whole-brain emulation of the adult Drosophila melanogaster connectome (~125,000–140,000 neurons, ~50 million synapses) in MuJoCo with NeuroMechFly v2.

[21] Carl Sagan, The Demon-Haunted World: Science as a Candle in the Dark (Random House, 1995), https://www.penguinrandomhouse.com/books/159731/the-demon-haunted-world-by-carl-sagan/.


The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.