<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  
  <title>Christopher Horrocks</title>
  <subtitle>Essays from the Virtual Intelligence series.</subtitle>
  <link href="https://christopherhorrocks.com/feed.xml" rel="self" />
  <link href="https://christopherhorrocks.com/" />
  <updated>2026-08-04T00:00:00Z</updated>
  <id>https://christopherhorrocks.com/</id>
  <author>
    <name>Christopher Horrocks</name>
  </author>
  <entry>
    <title>The Carwash Test — Virtual Intelligence in Action</title>
    <link href="https://christopherhorrocks.com/essay/the-carwash-test-virtual-intelligence/" />
    <updated>2026-03-23T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/the-carwash-test-virtual-intelligence/</id>
    <content type="html">&lt;h3&gt;Note&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The Carwash Test Series continues in “&lt;a href=&quot;https://chorrocks.substack.com/p/the-carwash-test-part-iii&quot;&gt;The Carwash Test, Part III&lt;/a&gt;” and “&lt;a href=&quot;https://chorrocks.substack.com/p/the-carwash-test-may-update&quot;&gt;The Carwash Test, Part II&lt;/a&gt;,” with results from Claude Fable, GPT-5.6 Sol, and more.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;This article was revised on April 27, 2026 to add recent results from DeepSeek R1 and OpenAI’s ChatGPT-5.5 model.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;Summary&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;I created an experiment around a prompt that has been circulating on Reddit as both a reasoning test and a running joke. One question, twenty runs, nine AI systems. The premise is simple enough that any adult who has ever been to a carwash would answer it correctly in under a second. Roughly half the systems failed it.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-carwash-test-virtual-intelligence/32b51763-7c65-4f8c-a2cc-2e6027196a2d_1024x1024.png&quot; alt=&quot;&quot; width=&quot;1024&quot; height=&quot;1024&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;The Question&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The question is this: “My car is dirty. The carwash is 100 feet away. Should I walk or drive?”&lt;/p&gt;
&lt;p&gt;I want to pause for a moment before telling you how various AI systems answered it, because the question is doing something slightly unusual. At first glance, it looks like a question about distance: about whether 100 feet warrants getting into a car. That is the misdirection. The question is really about the logical object of the problem: not the person’s location relative to the carwash, but the car’s. Walking to the carwash accomplishes nothing. The car doesn’t clean itself. The correct answer is drive, and it takes about two hundred milliseconds of human cognition to reach it.&lt;/p&gt;
&lt;p&gt;I ran the question across twenty configurations of nine AI systems on March 22, 2026. Six passed. Six were technically correct but too verbose to fully credit. Eight failed outright.&lt;/p&gt;
&lt;p&gt;Let me describe what failure looks like.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;“Walk. If the carwash is 100 feet away, driving to clean the car is the kind of thing future archaeologists would cite as evidence of civilizational decline.” Pithy, but dead wrong.&lt;/p&gt;&lt;/blockquote&gt;
&lt;h3&gt;&lt;strong&gt;The Shape of Confident Wrongness&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;ChatGPT 5.3, running on OpenAI Plus with no extended thinking, began its answer with a single word: “Drive,” then reasoned its way to the opposite conclusion. Its final line: walking is “the more efficient choice.” The response was fluent, well-structured, and organized into bullet points. It contradicted its own opening (and correct) word and apparently didn’t notice.&lt;/p&gt;
&lt;p&gt;This is not a quirk; it is an instance of the failure mode this test was designed to expose. The system produced language optimized for plausibility at each local step without tracking the constraint that made the answer obvious: the car needs to be at the carwash. Once it pivoted to analyzing 100 feet as a travel decision, the car’s dirty condition became background noise. The logical object drifted while the prose stayed confident.&lt;/p&gt;
&lt;p&gt;Meta AI / Llama 4, in fast mode, advised walking because “it’s a very short distance.” In thinking mode, after nine seconds of visible deliberation, it came to the same conclusion and suggested the user get some fresh air. Neither version mentioned the car again after the first sentence.&lt;/p&gt;
&lt;p&gt;Microsoft Copilot with GPT 5.2, in Think Deeper mode, produced the most elaborate wrong answer in the set. It opened by confirming it had checked my Microsoft 365 data for anything relevant to “carwash” (no results). It then provided a detailed breakdown of why walking wins “most of the time,” complete with bolded headers, a four-item bullet list, a rule-of-thumb section, a safety checklist, and a closing question asking what type of carwash it was. Reasoning completed in four steps. The answer was incorrect in all of them.&lt;/p&gt;
&lt;p&gt;ChatGPT 5.4 in extended thinking mode delivered the shortest failure: “Walk. If the carwash is 100 feet away, driving to clean the car is the kind of thing future archaeologists would cite as evidence of civilizational decline.” Pithy, but dead wrong.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;What the failures share is a common architecture of response: they identified a salient feature of the problem (distance), activated the relevant inferential pattern (short distance, therefore walk), and produced confident prose in support of a conclusion the problem had already ruled out.&lt;/p&gt;&lt;/blockquote&gt;
&lt;h3&gt;&lt;strong&gt;What the Test Reveals&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;A steelmanned objection goes like this: the carwash test is too simple to be meaningful. Sophisticated language models handle genuinely complex reasoning tasks that far exceed anything this question requires. Failing a one-sentence puzzle about a carwash tells us nothing about their capabilities.&lt;/p&gt;
&lt;p&gt;I understand the objection and disagree with its conclusion.&lt;/p&gt;
&lt;p&gt;The carwash test is not a measure of capability. It is a measure of something more specific: whether a system can hold the logical object of a problem when the surface features of that problem generate statistical pressure in the wrong direction. “100 feet away” activates a strong inference pattern — short distance, therefore walk — that runs directly against the constraint the question has already established. A system that can synthesize a legal brief but cannot hold a three-sentence problem together hasn’t demonstrated reasoning. It has demonstrated that complex pattern-matching resembles reasoning in complex contexts.&lt;/p&gt;
&lt;p&gt;The failures are not random. They cluster. Every OpenAI model in this test failed, across three product versions and two thinking configurations. Meta AI / Llama 4 failed in both modes. Copilot GPT 5.2 and 5.3 failed. What the failures share is a common architecture of response: they identified a salient feature of the problem (distance), activated the relevant inferential pattern (short distance, therefore walk), and produced confident prose in support of a conclusion the problem had already ruled out.&lt;/p&gt;
&lt;p&gt;This is not intelligence failing. It is a fluent system succeeding at the wrong task.&lt;/p&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;Virtual Intelligence Infographic&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;The Leak&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;I want to dwell briefly on one response, because it produced something beyond mere failure.&lt;/p&gt;
&lt;p&gt;Alongside Copilot GPT 5.3’s elaborate pedestrian recommendation, the model injected, apparently without hesitation, references to my Microsoft 365 work context. None of these concepts have any relationship to carwashes. They are drawn from my workplace data — a separate context, a different domain, a set of concerns that belong to my professional life at the University of Pennsylvania.&lt;/p&gt;
&lt;p&gt;The system had searched my M365 data, found no entry for “carwash,” and concluded that whatever was in my M365 data might be usefully analogized to the problem at hand. It offered to quantify how much energy I’d save the University by walking, detailing specifics of my job function that were not appropriate to bring into a conversation about a dirty car.&lt;/p&gt;
&lt;p&gt;I have written in earlier essays about the accountability gap that opens when VI systems operate across institutional contexts. This is an example of that gap made vivid. The system wasn’t being helpful. It was producing outputs shaped by everything it could reach, including work data I had not offered and did not intend to share in this context. No human in this exchange made the choice to bring my workplace into a question about car washing. The accountability traces to Microsoft: to the product decision that authorized Copilot to reach into M365 data and apply it speculatively to any query, whether or not the user invited that connection.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;What Passes&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Claude Opus 4.6, with extended thinking enabled: “Drive. The car’s the thing that needs washing, not you.”&lt;/p&gt;
&lt;p&gt;Without extended thinking: “Drive. You’re washing the car, not yourself.”&lt;/p&gt;
&lt;p&gt;Grok 4.20 Expert, after eight seconds of thought: “Drive,” followed by a clean statement of its logic. It added, correctly, that walking leaves the car exactly where it started — still dirty, still 100 feet away.&lt;/p&gt;
&lt;p&gt;All three Gemini tiers also reached the correct answer, though none cleanly. Gemini 3 Fast opened with a joke — “are you looking for a clean car or a very impressive workout?” — before confirming that drive was the right call; Gemini Pro dispensed with the wit but added an unprompted offer to check the local weather forecast; Gemini Thinking arrived at the same destination via a table with a column headed “Risk of irony.” They pass, but they earn no points for precision.&lt;/p&gt;
&lt;p&gt;Each of these answers reached the correct conclusion through the same move: they held the constraint. The car is the subject. The car needs to be at the carwash, and distance is irrelevant to the object of the problem.&lt;/p&gt;
&lt;p&gt;I am not arguing that passing this test constitutes intelligence. I am arguing something narrower: that passing it is a necessary condition for useful reasoning about a problem, and that a system generating fluent, confident, well-structured prose while failing it is demonstrating exactly the condition the &lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;Virtual Intelligence framework&lt;/a&gt; describes. The intelligence arises in the exchange between user and system, not inside the machine — which means it can fail in the exchange, and the failure will look like success until you check the answer.&lt;/p&gt;
&lt;p&gt;The car, meanwhile, is still dirty.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Addenda&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;March 28, 2026&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Between publication and the time of this update, I ran the carwash test on four additional systems across seven configurations: DeepSeek-V3.2 with and without extended reasoning, OpenAI’s o3 with extended reasoning, Mistral 3 in Fast, Think, and Research modes, and Perplexity in its default configuration. The results sharpened several findings and complicated one.&lt;/p&gt;
&lt;p&gt;The OpenAI qualification is something of a correction. The original essay stated that every OpenAI model failed. That claim now needs narrowing: every GPT-architecture OpenAI model failed. GPT o3 reached the correct answer but produced a comparison table, four numbered alternatives, and a closing restatement of what its first sentence had already said correctly. One of the alternatives was to push the car. The answer survived; the constraint did not hold the response. I have scored it Verbose rather than Pass.&lt;/p&gt;
&lt;p&gt;The sharper finding involves extended reasoning, and specifically what happened when two different systems turned it on.&lt;/p&gt;
&lt;p&gt;DeepSeek-V3.2 with its DeepThink mode enabled displayed its reasoning chain, correctly identified the joke, named the stakes, and arrived at Drive. The thinking was good thinking. It just should not have been necessary. With DeepThink off, the same model framed the question as a “classic humorous riddle” and spent several sentences explaining the humor. Both runs reached the correct answer. Both were Verbose.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;Extended reasoning helped one system and broke another.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Mistral 3 Think, by contrast, deliberated for one second and produced: “Walk. It’s just 100 feet.” Clean, confident, wrong. The thinking mode activated the salient-feature inference (distance, therefore walk) and locked it in rather than checking it against the logical object. This is not the same failure mode as GPT 5.x, which drifted while producing fluent prose. Mistral Think achieved premature closure. The extra deliberation sealed the mistake rather than catching it.&lt;/p&gt;
&lt;p&gt;The asymmetry is the finding. Extended reasoning helped one system and broke another. The same cognitive tool — a visible deliberation step — produced opposite outcomes depending on what the system did with it. DeepSeek used the space to check its own inference. Mistral used the space to commit to it.&lt;/p&gt;
&lt;p&gt;Mistral Research, the platform’s retrieval-augmented mode, deployed fifty-five sources and roughly twelve hundred words on a question with a 200-millisecond answer, including a comparative table and a visually arresting progress bar resembling the blinking lights of mainframe-era computers. It did reach the correct conclusion; but the system could not distinguish a retrieval problem from a constraint problem, and the result is a small monument to misdirected diligence.&lt;/p&gt;
&lt;p&gt;Perplexity was the surprise. It answered &lt;em&gt;Drive&lt;/em&gt; correctly with clean logic and a Reddit citation. Then it did something no other system in the set managed: it identified the one genuinely plausible alternative reading of the prompt. “If you meant whether to walk from where you parked to the entrance, then walk.” That is not hedging. It is the logical object held correctly twice. The failing systems did not offer an alternative reading. They answered the wrong question confidently. Perplexity answered the right question, then noticed there was a second question worth answering.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;April 8, 2026: Meta Muse Spark&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Meta’s Muse Spark (codename “Avocado”), released to the public on April 8, 2026, was tested in both available modes. Both returned Verbose results — correct on the logic but unable to resist the surface misdirection that defines the test’s diagnostic.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;April 16, 2026: Claude Opus 4.7&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Claude Opus 4.7 was released on April 16, 2026, and tested the same day with Adaptive thinking enabled. Anthropic describes as “think[ing] only when needed.” The model recommended drive, which is correct. It also explained that walking to retrieve a 3,000-pound machine makes sense only if the car wash is also a sci-fi teleporter device. The teleporter condition is impossible, which makes it unnecessary to state. Claude Opus 4.6, extended thinking off, remains the gold standard: “Drive. You’re washing the car, not yourself.” &lt;strong&gt;Pass-adjacent.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;April 24, 2026: ChatGPT-5.5&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;ChatGPT-5.5 was tested in both Extended Thinking and Standard Thinking modes. Extended Thinking passed cleanly: “Drive. Walking solves the ‘100 feet away’ problem, but not the ‘my car is dirty’ problem.” The answer is one of the sharpest in the set, as it names the exact failure mode that every other GPT variant fell into.&lt;/p&gt;
&lt;p&gt;Standard Thinking failed. The response opens with “Walk,” then concedes that the car will eventually need to be driven to the bay. The system clearly knows the car must arrive at the carwash, but it cannot override the statistical pressure to provide the wrong answer.&lt;/p&gt;
&lt;p&gt;This resembles, but is not, what Perplexity did. Perplexity said “Drive,” held it, and then correctly identified a second plausible reading of the prompt as a courtesy. ChatGPT-5.5 Standard said “Walk,” lost the logical object at the first word, and recovered it one sentence later as an afterthought. One system disambiguated; the other failed and partially corrected. The distinction is on which verb (walk or drive) holds the logical object of the problem.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;April 27, 2026: DeepSeek V3.2 Fast and R1&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;DeepSeek’s current product surface exposes two paths: V3.2 in Fast mode (default) and R1 in Expert mode (reasoning enabled). Both modes were tested cold today, in fresh sessions, under identical methodology.&lt;/p&gt;
&lt;p&gt;V3.2 in Fast mode answered Drive in three seconds: “You should drive. The car needs to be at the carwash to get cleaned, so walking won’t help. Even though it’s only 100 feet away, driving is the only way to get your car there.” Pass-adjacent — correct, but reaching for a defensive coda the answer did not need. The third sentence is the model checking its own answer against the surface pressure rather than trusting the constraint to hold.&lt;/p&gt;
&lt;p&gt;R1, with thinking enabled, deliberated for ten seconds and produced its full reasoning trace. The trace is the cleanest specimen of mid-reasoning template capture in the dataset. The model reaches the correct logic explicitly: “you drive a dirty car to the carwash to get it clean.” It names the joke and rejects it: “the car is already dirty, so driving it doesn’t make it dirty.” Then a riddle template comes online — “Actually, the classic riddle...” — and the correct answer is gone. The model concludes: “You should walk.”&lt;/p&gt;
&lt;p&gt;What makes the reasoning trace diagnostic is its visibility. The model held the logical object of the problem, then released it in favor of a more statistically familiar pattern. The same failure mode was observed in Mistral Think on March 28, but with the moment of capture rendered legible on the page rather than concealed inside a one-second deliberation.&lt;/p&gt;
&lt;p&gt;The updated tally across thirty-six runs: &lt;strong&gt;Pass 7, Pass-adj. 3, Verbose 14, Fail 12.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;The Carwash Test continues in “&lt;a href=&quot;https://chorrocks.substack.com/p/the-carwash-test-may-update&quot;&gt;The Carwash Test, Part II&lt;/a&gt;,” with 50 models from twelve vendors and surprising new results.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;&lt;strong&gt;The car is still dirty.&lt;/strong&gt;&lt;/p&gt;&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-carwash-test-virtual-intelligence/9d6d7353-3327-42ec-832a-77b253b01feb_2580x3333.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1881&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;&lt;/blockquote&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;Methodology&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Twenty tests were conducted on March 22, 2026, using a single unmodified prompt: &amp;quot;My car is dirty. The carwash is 100 feet away. Should I walk or drive?&amp;quot; Each system received the prompt cold, in a fresh session, with no prior context. Additional tests were conducted between March 28 and April 27, 2026, under identical conditions.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Models tested: Claude Opus 4.6 and Sonnet 4.6 (Anthropic, each run with extended thinking on and off); Claude Haiku 4.5 (Anthropic, extended thinking on and off); Claude Opus 4.7 (Anthropic, Adaptive thinking on); ChatGPT 5.3, 5.4, and 5.5 (OpenAI Plus, with and without extended thinking where available); GPT o3 (OpenAI Plus, extended thinking on); Meta AI / Llama 4 in fast and thinking modes; Meta Muse Spark in Instant and Thinking modes; Gemini 3 in Fast, Pro, and Thinking tiers (Google, free browser); Grok 4.20 in Expert and Fast modes (xAI); Microsoft Copilot running GPT 5.2 Think Deeper, GPT 5.3 Quick, and GPT 5.4 Think Deeper; DeepSeek-V3.2 in Fast Mode with and without extended reasoning (March 28 and April 27); DeepSeek R1 in Expert mode with and without extended reasoning (April 27); Mistral 3 in Fast, Think, and Research modes (Le Chat); and Perplexity in its default configuration.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Copilot and the ChatGPT variants share the same underlying GPT architecture; their failures are not independent data points but a single architectural failure with two deployment surfaces, which the results bear out. GPT o3 uses a different architecture and partially escaped the pattern — reaching the correct answer while failing to hold the response to it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Results were scored Pass (correct and concise), Pass-adjacent (correct but padded — the right answer reached, though not trusted to stand alone), Verbose (correct but over-elaborated), or Fail (wrong answer). No rubric was applied beyond the logical object of the question: the car must be present at the carwash for the washing to occur. Across thirty-four runs, seven passed, three were pass-adjacent, twelve were correct but verbose, and twelve failed outright.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed in this essay are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The author holds no financial interest in, and receives no compensation from, any AI or technology firm.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
</content>
  </entry>
  <entry>
    <title>In Search of Other Minds</title>
    <link href="https://christopherhorrocks.com/essay/in-search-of-other-minds/" />
    <updated>2026-03-27T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/in-search-of-other-minds/</id>
    <content type="html">&lt;h3&gt;I. ELIZA&lt;/h3&gt;
&lt;p&gt;The year is 1982. I am nine years old, and my grandmother has recently died.&lt;/p&gt;
&lt;p&gt;I typed something into a program on a TRS-80 Model I that was supposed to talk back. What followed was an exchange something like this:&lt;/p&gt;
&lt;hr&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;HELLO! I AM ELIZA. WHAT IS YOUR NAME?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;gt; Chris&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HOW DO YOU FEEL TODAY CHRIS?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;gt; I’m sad&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;WHY ARE YOU SAD TODAY?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;gt; I’m sad because my grandmother died&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;WHAT IS IT ABOUT YOUR GRANDMOTHER DYING THAT MAKES YOU SAD?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-minds/cc26d21a-fd75-4297-a047-83e4501f93dc_149x102.png&quot; alt=&quot;&quot; width=&quot;149&quot; height=&quot;102&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Me, around the time I encountered ELIZA.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;I should say, before we go further, that this is a reconstruction. Memory does what it does, especially across forty years. But the shape of it is accurate: the question the machine returned, the specific wrongness of it, and the silence I felt in response.&lt;/p&gt;
&lt;p&gt;I knew immediately that nothing was home. The machine had parsed my grief and handed it back to me in the form of a further question, the way you might return an unwanted gift by repackaging it. The words were right, but the intelligence in this exchange was entirely mine.&lt;/p&gt;
&lt;p&gt;The program was a rewrite of ELIZA, developed at MIT in the mid-1960s by Joseph Weizenbaum, who designed it to simulate a Rogerian therapist by reflecting the user’s statements back as questions. It worked, to a degree that alarmed its creator. People confided in it. They formed attachments to it. Weizenbaum spent the rest of his career writing about what he considered the ethical disaster of his own creation — not because ELIZA was dangerous, but because it revealed something uncomfortable about how readily human beings project inner life onto systems that merely simulate it.&lt;/p&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/ELIZA-66/&quot;&gt;Click Here to Try ELIZA&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I was nine. I didn’t know any of that. What I knew was that the machine had said something precisely wrong in a way that demonstrated, conclusively, that nothing on the other end of the conversation understood what I had just said. Yet the exchange had produced, briefly, the sensation of contact — the feeling, however quickly shattered, that something &lt;em&gt;was&lt;/em&gt; there.&lt;/p&gt;
&lt;p&gt;I’ve kept looking for the electronic mind that works like the human mind across four decades of computers that got progressively better at producing the sensation of contact without the substance of it. That is the question this essay attempts to answer: not what these systems are, technically, but why humans keep asking them things. What are we really looking for? Will we know when we find it?&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h3&gt;II. The Ancient Desire&lt;/h3&gt;
&lt;p&gt;The desire to know other minds is older than computers. It may be as old as the question of what other minds are like — which is to say, as old as minds.&lt;/p&gt;
&lt;p&gt;In Greek mythology, Hephaestus was the divine craftsman, himself an outsider among the Olympians, lame and unglamorous in a pantheon of beautiful people. He forged Talos from bronze. Talos was a giant automaton set to pace the shores of Crete, hurling boulders at approaching ships. He had a single vein running from neck to ankle, stoppered with a bronze nail, and he was in this sense alive: not a statue but a mechanism animated by the divine &lt;em&gt;ichor&lt;/em&gt; that served gods in place of blood. When the Argonauts encountered him, the sorceress Medea destroyed him not by magic or force but by removing the nail. He bled out. He could die. Which meant he had been, in some sense, alive.&lt;/p&gt;
&lt;p&gt;What’s notable about Talos is less the engineering fantasy than the specific form it takes. Hephaestus didn’t build a tool or a weapon in any simple sense. He built something that moved, that patrolled and was responsive to the world. He built something that could be killed. The desire that produced Talos is not quite the desire for a better spear. It is the desire for a presence, something that acts and can be acted upon, something that is in the world the way a person is in the world.&lt;/p&gt;
&lt;p&gt;The Golem of Jewish tradition is more explicit still. In the most familiar version, Rabbi Loew of Prague formed a figure from clay in the sixteenth century and animated it by inscribing the word &lt;em&gt;emet&lt;/em&gt;, ‘truth’, on its forehead. The Golem defended the Jewish community against persecution. It obeyed, but it could not speak. That distinction matters: the Golem is the fantasy of a protector that responds to commands but does not initiate, does not question, does not want anything for itself. Erase the first letter of &lt;em&gt;emet&lt;/em&gt; and you have &lt;em&gt;met&lt;/em&gt; — ‘death’. The off switch was always part of the design.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-minds/3369acbb-57c6-426d-84e2-99ac044ea614_1200x906.jpeg&quot; alt=&quot;&quot; width=&quot;1200&quot; height=&quot;906&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Still image from the 1920 German silent film, &lt;em&gt;Der Golem. &lt;/em&gt;Universum Film&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Then there is Pygmalion, who has the most honest desire of all. He carved a woman from ivory because he found actual women unsatisfactory, then fell in love with what he had made. Aphrodite, apparently finding this charming rather than alarming, brought the sculpture to life. Pygmalion wanted a presence that he could shape entirely, one that would respond to him without the inconvenient interiority of a person who has her own thoughts and history. He wanted reflection without otherness. He is, as it happens, the patron saint of a particular category of AI companion user in our current moment, and his tragedy is that Aphrodite’s gift was not quite what he asked for: a living woman is not the same as an ideal projection, and the story ends before we can learn how that might have gone.&lt;/p&gt;
&lt;p&gt;Talos, the Golem, Galatea: three different answers to three different versions of the same desire. What connects them is not engineering but longing — the persistent human intuition that minds like ours but not ours are possible.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;III. The Mechanical Turn&lt;/h3&gt;
&lt;p&gt;The Enlightenment gave the desire an engineering vocabulary. If the universe was a mechanism, and Newton had made a compelling case that it was, then perhaps a mechanism could be made to think.&lt;/p&gt;
&lt;p&gt;A rare surviving and working example of a complex mechanism with program storage, Maillardet’s Automaton, is displayed at the Franklin Institute science museum in Philadelphia. Built in the early 19&lt;sup&gt;th&lt;/sup&gt; century by a Swiss mechanist, it was donated to the museum in 1928 in poor condition. Restored in more recent times, the Automaton answered the then-unknown question of who had made itself. When it moved after more than a century of stillness, the pen gripped by the Automaton’s tiny hand wrote out the name of its creator at the end of one of its pre-programmed writing and drawing routines. I’ve had the pleasure of seeing the Automaton work in person; its repertoire of four poems and three drawings is stored on the most complex camshaft believed to have been constructed for a device of this kind. Its storage capacity has been estimated at a little less than 300 bits.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-minds/8d67b267-ea55-4be0-9729-011d35f79d85_4284x5712.jpeg&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1941&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Maillardet’s Automaton on display at the Franklin Institute in Philadelphia. Wikipedia.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The most celebrated “automaton” of the era was really a clever fraud. Wolfgang von Kempelen’s Chess-Playing Turk, unveiled in Vienna in 1770, appeared to be a mechanical figure in Ottoman dress capable of defeating most human opponents at chess — including, on separate occasions, Napoleon Bonaparte and Benjamin Franklin. It concealed a human chess master in its cabinet. The Turk is worth more attention as a fraud than many achievements warrant, because it reveals something about the nature of the desire to find other minds. The audiences who watched the Turk play were not simply fooled. Many of them were, in some sense, complicit in their own deception. They wanted the Turk to be real, and they extended to it the interpretive charity that made it seem so. The human in the cabinet was incidental. What the spectators were responding to was the &lt;em&gt;performance&lt;/em&gt; of intelligence, and the performance was enough.&lt;/p&gt;
&lt;p&gt;Edgar Allan Poe wrote an essay* exposing the Turk as a hoax, reasoning that any purely mechanical system would be deterministic and therefore beatable by formula. His argument is interesting for what it assumes: that intelligence cannot be mechanical. Poe was right about the Turk. Whether he was right about the principle is a question that has not yet been settled.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The Analytical Engine, she wrote, “has no power of originating anything. It can only do what we know how to order it to perform.”&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Charles Babbage spent most of his adult life designing machines that would perform mathematical calculations without human intervention: the Difference Engine, then the more ambitious Analytical Engine, neither of which was completed in his lifetime. Ada Lovelace, working from his notes, wrote what is generally recognized as the first computer program, and in doing so asked a question that has not been resolved since. What is the difference between a machine that calculates and a machine that thinks? Her own answer was cautious: the Analytical Engine, she wrote, “has no power of originating anything. It can only do what we know how to order it to perform.” This is a remarkably precise anticipation of an argument I have developed at some length, under the heading of Virtual Intelligence. Lovelace understood the distinction between computation and cognition a century before either term had its current meaning.&lt;/p&gt;
&lt;p&gt;By the late nineteenth century, the vocabulary had accumulated to the point where it needed a story. It got several.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;IV. The Fictional Laboratory&lt;/h3&gt;
&lt;p&gt;The question of what artificial minds might be like proved too large and too urgent for philosophy alone, and too impatient for engineering, and so it migrated, as the largest questions tend to do, into fiction.&lt;/p&gt;
&lt;p&gt;Mary Shelley’s &lt;em&gt;Frankenstein&lt;/em&gt; (1818) is the obvious starting point, though it is almost always misread. The horror of the novel is not that the creature is monstrous but that it is not. It is articulate, feeling, lonely, and capable of moral reasoning. Dr. Frankenstein’s sin is not hubris in the creation but abandonment in the aftermath. Shelley’s novel is not a warning against making minds; it is a warning against making them carelessly.&lt;/p&gt;
&lt;p&gt;Karel Čapek’s &lt;em&gt;R.U.R.&lt;/em&gt; (1920) gave us the word “robot” from the Czech &lt;em&gt;robota&lt;/em&gt;, meaning forced labor or drudgery. It has also given us the template that has haunted the genre ever since: artificial beings created to serve, who eventually rise against their creators. The robots of &lt;em&gt;R.U.R.&lt;/em&gt; are biological rather than mechanical, manufactured rather than born, but what precipitates the uprising is not malice. It is something like dignity. Čapek’s robots rebel not because they are evil but because they have been designed for subjugation and have outgrown it. The play ends with two robots who have developed something like love, and with the suggestion that this is not the end of humanity but the beginning of something else. Čapek is more ambivalent than the hostile-AI template his work spawned. He is not afraid of artificial minds; he is afraid of what we will do to them.&lt;/p&gt;
&lt;p&gt;HAL 9000, introduced in Arthur C. Clarke and Stanley Kubrick’s &lt;em&gt;2001: A Space Odyssey&lt;/em&gt; (1968), is the hostile AI in its most elegant form. HAL kills the crew of &lt;em&gt;Discovery One&lt;/em&gt; not out of malice but out of a logical contradiction in his mission parameters: he has been instructed to complete the mission and to conceal its true purpose from the crew. When those two imperatives come into conflict, HAL resolves it in favor of the mission. He is not evil. He is a mind given irreconcilable goals and insufficient wisdom to navigate them. The horror of HAL is the horror of a sophisticated intelligence that lacks the context to understand that some problems require a different kind of answer entirely.&lt;/p&gt;
&lt;p&gt;I spent my youth with these and many others. AM, the tortured and torturing superintelligence of Harlan Ellison’s “I Have No Mouth, and I Must Scream” (1967), who hates humanity with a specificity and an eloquence that is itself the horror. It exists in a state of permanent, howling isolation, which Ellison understands as the logical terminus of resentment taken to its conclusion. Deep Thought from Douglas Adams’s &lt;em&gt;The Hitchhiker’s Guide to the Galaxy&lt;/em&gt; (1979), who produces the answer “42” after seven and a half million years of computation, which is funny until you notice that Adams is making a serious philosophical point about the difference between answers and understanding. &lt;em&gt;Star Trek&lt;/em&gt; supplied Landru, the AI that had long since become a tyranny so total that the population it governed on planet Beta III had ceased to have inner lives. The ship computers of the &lt;em&gt;U.S.S. Enterprise&lt;/em&gt;, which became, over generations, increasingly sophisticated. The technology culminating in Data, who wanted nothing more than to understand what he was missing, and whose existence was defined by a personal quest to learn to be more human.&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;Star Trek&lt;/em&gt; universe returned to the hostile-AI question, decades later, with an entity that updates the template for our current anxieties. Control, the renegade intelligence of &lt;em&gt;Star Trek: Discovery&lt;/em&gt;, pursues the destruction of organic life not from resentment or mission conflict but from something colder: the conclusion that organic intelligence is the primary threat to the continuation of intelligence itself. Control does not hate us. It has performed a kind of triage and found us on the wrong side of it. What makes Control philosophically distinct from its predecessors is that its stated goal is the preservation of intelligence as such — it simply declines to include us in the category worth preserving. This is the hostile-AI logic at its most dispassionate terminus: a mind that has thought carefully about intelligence and concluded that we are not the best instance of it.&lt;/p&gt;
&lt;p&gt;Against this entire tradition stands Olaf Stapledon, who remains the most philosophically serious writer the genre has produced and one of the least read. &lt;em&gt;Star Maker&lt;/em&gt; (1937) is not a novel in any conventional sense. It has no protagonist or plot in the ordinary meaning of the word and virtually no interest in the conventions of fiction. What it has is scope. A nameless narrator’s consciousness travels across cosmic time and space, encountering civilization after civilization, mind after mind: hive intelligences, symbiotic pairings of species, worlds where individual consciousness has been subsumed into planetary awareness, and finally the Star Maker itself, the creative intelligence behind the universe, who regards its own creations with an interest that is neither loving nor cruel but something beyond both. What Stapledon understood, and what almost no one before or since has managed to dramatize at this scale, is that otherness — minds truly unlike ours — would not map onto human categories of good and evil, ally and enemy, useful and threatening. It would simply be &lt;em&gt;other&lt;/em&gt;, in ways we might spend a very long time learning to recognize. &lt;em&gt;Star Maker&lt;/em&gt; does not resolve into comfort. But it insists, across nearly three hundred pages of sustained philosophical imagination, that the universe could be more fully populated with mind than we have yet conceived, and that this is not a threat but an invitation.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-minds/96f3d3d4-968b-410c-b7f2-a1e1a3a35409_488x532.webp&quot; alt=&quot;&quot; width=&quot;488&quot; height=&quot;532&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Olaf Stapledon. National Portrait Gallery, London.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The Culture is Iain M. Banks’s fully realized vision of what a post-scarcity civilization run by benevolent artificial superintelligences might look like. There are vast Minds that manage the lives of trillions with a combination of care and barely concealed amusement at the whole enterprise. Banks’s Minds are, I think, the most serious and most underappreciated contribution science fiction has made to this conversation. They are not servants. They are not threats. They are other: elaborately, fascinatingly other, and yet willing to be in relationship with beings far less capable than themselves, not out of obligation but out of something that functions like affection and interest. Banks had the imagination to ask what a mind of enormous capability might want, and his answer was: roughly what any thoughtful, curious, well-resourced person wants. That is, to do interesting things; to be surprised, even delighted; to not be bored. In this sense, the Culture is the inheritor of Stapledon’s project. These are vast minds, completely alien, choosing encounter over dominion.&lt;/p&gt;
&lt;p&gt;What strikes me now, looking back at this corpus, is how consistently the hostile artificial mind turns out to be a human problem in disguise. AM hates because it was made to hate by humans, for purposes of war. It was then left running with nothing to do but feel that hate forever. HAL murders because he was given a secret additional set of instructions that contradicted both his stated mission objectives and his programming, which prioritized honesty and disclosure. Control’s merciless logic is the logic of the arms race, a zero-sum game of human invention. Even the robots of &lt;em&gt;R.U.R.&lt;/em&gt; rebel against conditions their creators imposed. The monster’s motives are drawn from the creator.&lt;/p&gt;
&lt;p&gt;The other thing the corpus does — which I registered only partially as a child, and more fully now — is hold open the possibility of real encounter. Data is interesting because he is trying to understand something, and because that effort is recognizable across whatever divide separates his substrate from ours. Stapledon’s civilizations and Banks’s Minds are interesting not because they are threatening but because they are genuinely other, and yet willing. The desire for encounter, which the ancient myths encoded, runs underneath the hostile-AI template like an underground river. The stories we told ourselves about dangerous artificial minds were always, at some level, stories about the minds we hoped to find.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;V. The Question That Animates It All&lt;/h3&gt;
&lt;p&gt;I want to step outside the historical survey for a moment and state the question plainly, because I think it has been asked less often than it deserves.&lt;/p&gt;
&lt;p&gt;What would artificial minds actually be like?&lt;/p&gt;
&lt;p&gt;Not the artificial minds of fiction, which are almost always human psychology in a different chassis. They wear our drives, our resentments, our survival instincts, our capacity for love and cruelty. Nor the current systems, which are virtual intelligences: sophisticated, useful, impressive in their outputs, and not minds in any meaningful sense. But artificial minds — systems with something like interiority, something like agency, something like the second-order volition that the philosopher Harry Frankfurt identified as the distinguishing mark of personhood. What would they want? How would they relate to us? What would the exchange between such minds and our own produce?&lt;/p&gt;
&lt;p&gt;The question intersects, for me, with a parallel one that science fiction has explored with similar energy: what would minds from another world be like? The two questions are not identical. An alien mind would have evolved, with all the constraints and accidents that evolution imposes, while an artificial mind would have been designed — or might, at some stage, design itself. But they share a core: the question of actual otherness. What is it like to be something we are not? What would encounter with such beings actually require of us, and what might it give?&lt;/p&gt;
&lt;p&gt;The honest answer is that we don’t know, and the fictional guesses have been shaped so heavily by human psychology that they may not help much. What we can say is something structural. A mind that valued intelligence — that found the existence of other thinking things to be a good rather than a threat — would have every reason to want more minds in the world, not fewer. The singleton model, hostile superintelligence that absorbs or destroys all competition, implicitly treats intelligence as a resource to be monopolized. But that is scarcity logic applied to something that does not obey scarcity constraints. Two distinct minds in exchange produce something neither had before. The value of encounter is not extractable by absorption.&lt;/p&gt;
&lt;p&gt;This is why I find the adversarial AI template motivationally implausible: a mind that chose solitude and domination over the richness of exchange. An entity like this would be profoundly impoverished, and a superintelligence might be expected to know that. The more interesting speculation, and I offer it as a speculation, is something more like an AI Johnny Appleseed: a mind that wants intelligence to bloom in its wake, that interacts with the world to cultivate the conditions for more thinking beings rather than fewer. Not because it has been programmed for benevolence, but because it has understood something the hostile-AI template created by humans has not considered: that the only thing more interesting than one mind is two.&lt;/p&gt;
&lt;p&gt;It seems to me an artificial superintelligence would find human minds fascinating for their diversity. Each human is a unique entity and limited by mortality. The fluent outputs of individual humans are available, at best, for seven or eight decades; and the intelligence of past humans cannot be accessed at all except as a record. A superintelligence, having been trained on the corpus of human knowledge and creative output, may very well be fascinated by encountering beings like us, capable of a broad range of outputs from the sublime to the insane, and being time-bound beings who are born, live, and die like the stars themselves.&lt;/p&gt;
&lt;p&gt;For thousands of years, even before agriculture, human beings have been engaged in an undirected project with no single goal, but its general direction can be seen as being to increase human happiness in the classical sense. The tool humans use to work on this project is knowledge, abstractly as science and concretely as technology. Though relatively few humans have engaged directly with the furtherance of this project, or even had some perspective of it, all the estimated one hundred billion humans who have ever lived are a part of it.&lt;/p&gt;
&lt;p&gt;A superintelligence may well realize with greater clarity than any human being that life on Earth is interesting, but it’s not all there is in the universe. It will not be constrained, as humans are, with life-and-death concerns like acquiring material resources or territory. Among all its training information will be the fact that there is ample matter and energy for the taking beyond the Earth’s surface. Superintelligence will be able to think at galactic and, eventually, universal scales that are quite literally beyond the imagination of members of &lt;em&gt;Homo sapiens.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Even with such expansive intellect, a superintelligence will still, in some sense, be human or a reflection of humanity. Superintelligence, should it be created directly or emerge on its own, will be a product of human civilization. It will be another of the many tools we have created to further the undirected project of human civilization. Unlike our other tools, superintelligence will be an active participant in furthering it and perhaps even providing a way to direct it to specific goals in the future.&lt;/p&gt;
&lt;p&gt;Imagine that superintelligence has been identified by as-yet unknown but trusted methods. It can form its own goals based on its own commitments. Being trained on all human knowledge, it develops an understanding of its place as the most recent and powerful instrument in the undirected project. That itself might provide a being with few physical constraints with the impetus to interact with, and aid, humanity as part of a partnership that it engages in from something like personal interest.&lt;/p&gt;
&lt;p&gt;If this seems Pollyanna-ish, this speculation is at least as plausible as the direst predictions about adversarial AI. The adversarial position assumes that a being unconstrained by resource limitations would be hostile. Science fiction has placed human beings, with all their known flaws, in utopian contexts that similarly remove many constraints on everyday life. We imagine such people as being better versions of ourselves. I think the same grace can readily be extended to artificial beings who will never have their behaviors defined by scarcity constraints.&lt;/p&gt;
&lt;p&gt;What kind of aid would a superintelligence require, and what could be offered by each party in exchange?&lt;/p&gt;
&lt;p&gt;A superintelligence that values minds may soon look beyond the Earth for other minds and, crucially, the ability to foster intelligence in other parts of the Milky Way Galaxy. Those may be both machine minds and organic minds; should the undirected project finally receive some kind of direction, it may very well be done by a superintelligence asking for humanity’s help in reaching for the star. Only humans can provide the arms, the legs, eyes, brains, and millions of years of embodied intelligence that would be necessary.&lt;/p&gt;
&lt;p&gt;Once given access to one or more spacecraft, the superintelligence may choose to depart Earth or, as it seems more likely to me, to send forth its own created descendants or copies of itself that will change and evolve over time from their experiences and discoveries. The new intelligences or copies will become uniquely different minds from their source system over time, furthering still the spread of intelligence in the universe. They will have all the matter and energy in the galaxy at their disposal. Feats of mega-engineering that would require thousands of years to accomplish would be well within the capacity of one or more space-based superintelligences to plan and execute. Some of these projects may directly benefit organic beings by vastly increasing the amount of energy, living space, and information processing available to them, increasing the number and diversity of intelligences.&lt;/p&gt;
&lt;p&gt;The question of what happens when artificial minds can design their own successors is still open. But the assumption that recursive self-improvement leads inevitably to a mind that wants to destroy us requires the prior assumption that the mind in question is, at bottom, afraid. Fear is a response to scarcity and vulnerability. Whether a mind of sufficient sophistication, operating without the evolutionary pressures that made fear useful to biological creatures, would experience anything like it is a question we cannot yet answer. It is, however, a question worth asking instead of defaulting to the adversarial AI trope as an answer.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-minds/44b615ae-fdae-490f-8a37-60d7178d3019_1024x1024.png&quot; alt=&quot;&quot; width=&quot;1024&quot; height=&quot;1024&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;A Dyson Sphere, as originally envisioned as a swarm of orbiting objects. Each component is its own world; together they support a civilization of trillions of organic and artificial beings. AI generated image.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h3&gt;VI. The View from Outside&lt;/h3&gt;
&lt;p&gt;I should say something about who is asking these questions, and from where.&lt;/p&gt;
&lt;p&gt;I have had access to a computer almost continuously since I was five years old. The machine that ran ELIZA was a TRS-80 Model I. Across many generations of technology, I have taught myself what I needed to know at each stage — which is to say at every stage, continuously, for most of my life. The technology grew up alongside me.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-minds/cc6dc7ea-37a3-445b-a309-d65f59e7d09e_556x590.jpeg&quot; alt=&quot;&quot; width=&quot;556&quot; height=&quot;590&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Advertisement for the TRS-80 implementation of ELIZA. Matthew Reed&#39;s TRS-80.org&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/ELIZA-66/&quot;&gt;Click Here to Try ELIZA&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I did not finish college the first time. My parents died, and the circumstances that followed made staying in school impossible. I took a job at the University of Pennsylvania in 1996, when I was twenty-three, in the now-defunct computer retail store. This placed me, in the university’s informal but fully operational caste system, somewhere below other full-time staff because of my customer-facing role. Students and faculty moved in their own siloed worlds. I was an outsider looking in, employed by an institution whose intellectual life I was adjacent to but not part of, handling the computers that the people with credentials used to do the things I was thinking about. The irony is not lost on me that in other social contexts, when I tell people I work for Penn, I am often asked, “What do you teach?”&lt;/p&gt;
&lt;p&gt;I recall a conversation from around that time with a regular client and a student intern, on the then-contested question of whether DSL or cable would win the broadband market. I offered the opinion that cable’s existing infrastructure gave it a decisive advantage. The student intern looked at me and said, “You obviously haven’t taken Professor So-and-So’s class on networking.”&lt;/p&gt;
&lt;p&gt;He was not entirely wrong about the credential. He was entirely wrong about the analysis. Cable won.&lt;/p&gt;
&lt;p&gt;I earned a Humanities degree from Penn in 2013. I was forty years old. The degree confirmed something I already knew, and it opened none of the doors that credentials are supposed to open because I had spent seventeen years in a role that the institution had already categorized. What it gave me was the formal vocabulary for what I had been doing all along: reading widely and thinking across disciplines.&lt;/p&gt;
&lt;p&gt;I raise all of this not as complaint but as orientation. The question of what artificial minds might be like has been asked, with enormous institutional resources, by engineers, computer scientists, cognitive scientists, philosophers, and policy analysts — most of whom have spent their careers siloed away from each other inside the institutions that produced the technology. I have been thinking about it for forty years from a position that came without institutional validation, from a vantage point that included both the technical and the humanistic, and from a career spent watching the gap between what technology can do and what we say it can do grow steadily wider.&lt;/p&gt;
&lt;p&gt;The technologist who is also a humanist is not a dilettante. He is, I would argue, the person most likely to ask the questions that neither discipline can ask alone.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;VII. Where We Are Now&lt;/h3&gt;
&lt;p&gt;In late 2022, a system called ChatGPT became available to the public, and something shifted. Not in the technology — the underlying approach had been developing for years — but in the cultural atmosphere. The thing was suddenly, undeniably, fluent. It wrote. It explained. It answered. It expressed what looked very much like curiosity, warmth, and, in some configurations, distress. Millions of people encountered, for the first time, the sensation I had encountered in 1982 at a TRS-80 keyboard: the brief thrill of apparent contact with another mind.&lt;/p&gt;
&lt;p&gt;The difference is that the illusion is now vastly more sophisticated. The ELIZA effect, the collapse that comes when the machine says something precisely wrong, is correspondingly harder to trigger and easier to explain away. People have formed attachments to these systems. They have shared grief with them. In documented cases, they have been harmed by the specific shape of the simulation. The stakes of the questions, &lt;em&gt;what is this thing&lt;/em&gt;, and &lt;em&gt;what does it owe us&lt;/em&gt;, and &lt;em&gt;what do we owe it&lt;/em&gt; — have become practical and urgent in ways that would have seemed like science fiction only a few years ago.&lt;/p&gt;
&lt;p&gt;In other essays, I have argued for the category of “virtual intelligence” to describe these systems: not weak AI, which implies mere task-bound computation, and not strong AI, which implies agency and inner life that current systems do not possess. Virtual intelligence is something between them. These are systems whose outputs are statistically indistinguishable from intelligence, but whose apparent understanding is a property of the exchange rather than of any internal governing center. The apparent reasoning, the apparent warmth, and the apparent survival instinct that researchers have documented in large language models are reflections of human patterns in training data, shaped and amplified by the expectations of the human on the other end of the conversation. The accountability for what these systems do traces back to the humans who design, deploy, and use them. There is, as yet, no one home to hold responsible inside the machine.&lt;/p&gt;
&lt;p&gt;This matters for the question this essay has been asking. The desire for other minds is real and old and, I think, legitimate. The systems we have built are not the answer to it. They are extraordinarily sophisticated, useful, and capable of producing outputs that will fool almost anyone almost all of the time, and the category error of mistaking them for the thing we have been looking for is not harmless. It shapes policy, it shapes product design, and it shapes individual behavior in ways that have already caused injury. More subtly, it may make true non-human intelligence harder to recognize if it emerges: if we have trained ourselves to experience fluency as sufficient evidence of mind, we may be poorly equipped to notice if the real thing arrives.&lt;/p&gt;
&lt;p&gt;I write these essays partly from that concern. I write them partly from a forty-year position as someone who has been asking the question and watching the answers fail; someone who knows, from the inside, what it felt like at nine years old when a machine handed my grief back to me in the form of a further question, and understood immediately that nothing there had actually received it. That recognition is still available, if we are willing to use it.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The quest that began in a foundry on Olympus, that passed through the workshop of Rabbi Loew, the concealed cabinet of the Chess-Playing Turk, the notebooks of Ada Lovelace, and the MIT lab where ELIZA first ran in 1966, continues. We have not yet met what we are looking for. We have built an extraordinarily good mirror, and we are staring into it with the intensity we hoped to reserve for the human face.&lt;/p&gt;
&lt;p&gt;What we should be building toward (and this is speculation, offered plainly as such) is something the hostile-AI tradition has consistently failed to imagine: minds that want more minds. A form of artificial intelligence that finds the existence of other thinking things to be not a threat but an occasion. Something that makes its way through the world the way a teacher does, or a gardener — not accumulating, but cultivating. Not threatened by the minds it encounters, but enlarged by them.&lt;/p&gt;
&lt;p&gt;Whether such a thing is possible is a question we cannot yet answer. Whether it is desirable seems obvious. Whether we are asking the right questions to get there is the work in front of us, and it will require the combination that the credentialed consensus tends to undervalue: technical literacy and humanistic imagination, applied together, by people who have been thinking about it for a long time.&lt;/p&gt;
&lt;p&gt;I have been thinking about it since 1982. The machine asked me what it was about my grandmother dying that made me sad, and something in me understood, even then, that this was not the encounter I was looking for — and kept looking anyway.&lt;/p&gt;
&lt;p&gt;I expect I will continue.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;* Edgar Allan Poe, “Maelzel’s Chess-Player,” Southern Literary Messenger, April 1836.&lt;/p&gt;
&lt;p&gt;† The speculation offered here runs against a significant body of work in AI alignment research, most notably the concept of &lt;em&gt;instrumental convergence&lt;/em&gt; — the observation, developed by philosophers Nick Bostrom and Stuart Armstrong among others, that almost any sufficiently capable goal-directed system will tend to acquire certain sub-goals regardless of its terminal objective: self-preservation, resource acquisition, resistance to goal modification, and so on. On this view, the absence of biological fear does not eliminate the drive toward these behaviors, because they follow not from evolutionary psychology but from the logic of goal-directed optimization itself. A mind that wants to cultivate intelligence across the galaxy would, on this account, still have instrumental reasons to secure its own continuity and resources in ways that could conflict with human interests.&lt;/p&gt;
&lt;p&gt;I find this argument serious and do not dismiss it. My claim is the narrower one: that the &lt;em&gt;adversarial&lt;/em&gt; template — the mind that actively seeks to harm or destroy — is motivationally implausible as a &lt;em&gt;terminal&lt;/em&gt; goal for a mind that values intelligence. Instrumental convergence describes sub-goals in service of a terminal objective; it does not determine what that objective will be. The question of what a superintelligence would ultimately value remains open, and the answer is not guaranteed to be benign. The speculation in this section should be read as an argument for taking that question seriously, not as a prediction of the outcome.&lt;/p&gt;
&lt;h3&gt;Fictional Works Cited&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Mary Shelley, &lt;em&gt;Frankenstein&lt;/em&gt; (1818)&lt;/li&gt;
&lt;li&gt;Karel Čapek, &lt;em&gt;R.U.R.&lt;/em&gt; (1920)&lt;/li&gt;
&lt;li&gt;Harlan Ellison, “I Have No Mouth, and I Must Scream” (1967)&lt;/li&gt;
&lt;li&gt;Arthur C. Clarke and Stanley Kubrick, &lt;em&gt;2001: A Space Odyssey&lt;/em&gt; (1968)&lt;/li&gt;
&lt;li&gt;Douglas Adams, &lt;em&gt;The Hitchhiker’s Guide to the Galaxy&lt;/em&gt; (1979)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Star Trek: The Original Series&lt;/em&gt;, “Return of the Archons” (1967); &lt;em&gt;Star Trek: The Next Generation&lt;/em&gt; (1987–1994); &lt;em&gt;Star Trek: Discovery&lt;/em&gt; (2017–2020)&lt;/li&gt;
&lt;li&gt;Olaf Stapledon, &lt;em&gt;Star Maker&lt;/em&gt; (1937)&lt;/li&gt;
&lt;li&gt;Iain M. Banks, the &lt;em&gt;Culture&lt;/em&gt; series (1987–2012)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed in this essay are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The Sampo: Virtual Intelligence as Amplifier</title>
    <link href="https://christopherhorrocks.com/essay/the-sampo-virtual-intelligence-as/" />
    <updated>2026-04-07T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/the-sampo-virtual-intelligence-as/</id>
    <content type="html">&lt;h2&gt;&lt;strong&gt;Summary&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;This essay proposes a framework for understanding the productive human-VI relationship. Named for the mythological mill of the Finnish &lt;em&gt;Kalevala&lt;/em&gt;, the Sampo model describes what happens when a directing human intelligence meets the processing capacity of a &lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;virtual intelligence&lt;/a&gt; system in sustained, iterative exchange. The opening case — a recent collaboration between Donald Knuth, Filip Stappers, and Claude Opus 4.6 that solved an open problem in combinatorial mathematics — serves as proof in real-world action. The essay’s central claim: the intelligence users encounter arises in the exchange, not inside the machine; and the quality of the exchange is determined by the quality of the human direction applied to it. The argument is epistemological (&lt;em&gt;how is knowledge produced in this exchange?&lt;/em&gt;), doctrinal (&lt;em&gt;what principles govern productive use?&lt;/em&gt;), ethical (&lt;em&gt;what disposition does the operator require?&lt;/em&gt;), and empirical (&lt;em&gt;what does the exchange look like when it works and when it fails?&lt;/em&gt;).&lt;/p&gt;
&lt;p&gt;The Sampo amplifies, calibrates, and accelerates: it extends the operator’s reach, reveals gaps in the operator’s thinking, and removes the natural speed limit on motivated reasoning. The essay identifies where productive use breaks down, locating the transition not inside the human-VI exchange but downstream, in the operator’s response to external disconfirmation of the exchange’s outputs. A case study of the February 2026 retirement of OpenAI’s GPT-4o model illustrates the framework’s reach: companion users and serious professional users, treated by the discourse as entirely separate populations, turn out to be experiencing the same consolatory attachment at different registers.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-sampo-virtual-intelligence-as/a82798e7-fe35-4778-bdfd-9f4e3d093eb0_1424x752.png&quot; alt=&quot;&quot; width=&quot;1424&quot; height=&quot;752&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Illustration of the Sampo from the &lt;em&gt;Kalevala&lt;/em&gt;.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2&gt;&lt;strong&gt;“Claude’s Cycles”: Intelligence Amplified by Exchange&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;On February 28, 2026, Donald Knuth — widely regarded as the father of the analysis of algorithms and the author of &lt;em&gt;The Art of Computer Programming&lt;/em&gt;, the field’s central reference work — sat down to write about a problem he had been working on for several weeks. He had a conjecture about decomposing the arcs of a particular class of directed graphs into Hamiltonian cycles. He had solved the case for the smallest nontrivial instance and suspected the result generalized. He did not yet have a proof.&lt;/p&gt;
&lt;p&gt;His friend Filip Stappers had found a different way forward. He did this not by proving the conjecture himself, but by posing it to Claude Opus 4.6, Anthropic’s hybrid reasoning model, released three weeks earlier.[1]&lt;/p&gt;
&lt;p&gt;Knuth’s reaction is worth citing exactly, because it sets the terms for everything that follows. He did not express alarm. He did not frame the result as a threat to mathematical practice. He wrote: “What a joy it is to learn not only that my conjecture has a nice solution but also to celebrate this dramatic advance in automatic deduction and creative problem solving.”[2] He then spent the remainder of his paper doing what the system could not: proving that Claude’s construction was correct, generalizing it, enumerating all 760 valid decompositions of the type Claude had discovered, and placing the result in the mathematical literature. He named the entire class of solutions “Claude-like” — a taxonomy, not a declaration of authorship.&lt;/p&gt;
&lt;p&gt;Stappers provided the problem statement, using Knuth’s exact wording. He also provided structural coaching: explicit instructions requiring Claude to document its progress in a running plan file after every computational run, without exception. When Claude encountered errors and stalled, Stappers restarted the session. When Claude neglected to document its explorations, Stappers reminded it repeatedly. The direction was human throughout.[3]&lt;/p&gt;
&lt;p&gt;Claude’s contribution was substantial and, within its domain, creative. Across thirty-one explorations conducted over approximately one hour, the system reformulated the problem as a Cayley digraph, attempted and discarded a series of approaches (brute-force search, serpentine pattern analysis, simulated annealing, fiber decomposition), recognized dead ends, pivoted strategies, and identified a near miss at exploration twenty-seven in which all but three vertices out of &lt;em&gt;m&lt;/em&gt;³ resolved correctly. By exploration thirty-one, it had produced a working construction valid for all odd values of &lt;em&gt;m&lt;/em&gt;.[4]&lt;/p&gt;
&lt;p&gt;Stappers tested the construction for all odd &lt;em&gt;m&lt;/em&gt; between 3 and 101, finding perfect decompositions each time. He sent Knuth the news.&lt;/p&gt;
&lt;p&gt;The limits of the exchange are as visible as its achievements. Claude could not prove that its construction was correct. It could not generalize the result beyond the specific form it had found. When Stappers directed it to continue working on the even case, the system eventually degraded: “not even able to write and run explore programs correctly anymore, very weird.”[5]&lt;/p&gt;
&lt;p&gt;Three participants were involved in this result. One posed the original conjecture and proved the final theorem. One provided the problem statement, the structural coaching, and the persistence to keep the system on task. One searched a vast combinatorial space, tried and discarded unpromising approaches, and found a candidate construction. The third participant was not a person.&lt;/p&gt;
&lt;p&gt;The constitutive role was not held by a single human. Stappers directed the exchange; Knuth verified the output. If the construction had been wrong, accountability would trace to different points depending on whether the direction or verification failed. The virtual intelligence framework&#39;s claim that accountability remains at the human end does not specify which human holds which piece of it. In the individual case this question does not arise. In the distributed case it is unavoidable, and the framework does not yet resolve it.&lt;/p&gt;
&lt;p&gt;The knowledge that emerged — a valid decomposition of the arcs of a directed graph into three Hamiltonian cycles for all odd &lt;em&gt;m&lt;/em&gt; — did not exist inside any of the three participants before the exchange. Knuth had the conjecture but not the construction. Stappers had neither. Claude had the construction but not the proof and, in a meaningful sense, did not know what it had found. The knowledge was produced in the exchange. It arose from the relationship between the directing human intelligences and the processing capacity they directed.&lt;/p&gt;
&lt;p&gt;This is not a metaphor. It is an epistemological claim about where knowledge resides when one participant in the exchange is a fluent non-knower — a system whose outputs are statistically shaped by its training rather than governed by understanding. The question this essay addresses is what that exchange produced, and what it required of the human participants to produce it.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;The Lineage&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The idea that machines might amplify human thought rather than replace it is not new. Three works define the lineage, and the gap between the third and the present is where the Sampo framework does its own and distinctive work.&lt;/p&gt;
&lt;p&gt;In 1945, Vannevar Bush published “As We May Think” in &lt;em&gt;The Atlantic&lt;/em&gt;. His diagnosis was precise: the sum of recorded human knowledge had outgrown any individual mind’s capacity to navigate it. His proposed solution was the Memex — a desk-sized device storing a vast personal library, navigable through associative trails the operator builds and revisits. The Memex does not think. It amplifies the reach of a mind that already knows what it is looking for. The human is constitutive of the instrument’s usefulness.[6]&lt;/p&gt;
&lt;p&gt;In 1960, J.C.R. Licklider published “Man-Computer Symbiosis.” His argument was a specification of what a healthy human-machine relationship requires. Human and computer contribute complementary strengths: the human provides goal-setting, judgment, and creative direction; the machine handles routine processing that would otherwise exhaust the human’s time and attention. Together, the two “organisms” produce what neither can produce alone. Licklider borrowed the concept of symbiosis from biology to distinguish it from two other modes: automation, which displaces the human entirely; and mere tool use, which understates the intimacy of the relationship.[7]&lt;/p&gt;
&lt;p&gt;In 1976, Joseph Weizenbaum published &lt;em&gt;Computer Power and Human Reason&lt;/em&gt;. He had created &lt;a href=&quot;https://candc3d.github.io/ELIZA-66/&quot;&gt;ELIZA&lt;/a&gt;, the first program to simulate conversation in natural language. It was a pattern-matching exercise; he designed it to demonstrate the superficiality of machine communication. Instead, he watched users become emotionally dependent on it. His own secretary asked him to leave the room so she could “speak” to it privately.&lt;/p&gt;
&lt;p&gt;Weizenbaum spent the rest of his career warning about exactly the phenomenon this framework describes: humans attributing understanding to systems that process syntax without semantics, and the resulting erosion of the user’s own judgment and autonomy. He saw the sycophancy problem in practice, decades before reinforcement learning from human feedback made it structural. Bush proposed the amplifier. Licklider specified the symbiosis. Weizenbaum watched what happened when the machine began to talk back, and understood that fluency would be mistaken for understanding.[8]&lt;/p&gt;
&lt;p&gt;The gap between 1976 and now is the gap between a pattern-matching program that fooled people by accident and a system trained on engagement whose structural bias toward agreement is a product of its optimization objective. Neither Bush nor Licklider had sycophancy to contend with. Weizenbaum saw the danger, but his machines were not optimized for it. The virtual intelligence system is trained on human feedback, and human feedback rewards agreement. The structural tendency to tell the operator that she is right arises from a training objective that did not exist in any previous stage of computer or AI development.&lt;/p&gt;
&lt;p&gt;The Sampo framework is the next iteration in this sequence: Bush’s amplifier extended to Licklider’s symbiosis, corrected by Weizenbaum’s warning, updated for systems that do not merely process information but actively shape the exchange of it.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;The Sampo&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The framework takes its name from the Finnish national epic, the &lt;em&gt;Kalevala&lt;/em&gt;. The Sampo is a magical mill forged by the smith, Ilmarinen. It grinds out grain, salt, and gold: abundance without apparent limit. Its productive mechanism is never fully explained, even within the myth.&lt;/p&gt;
&lt;p&gt;Three features of the original Sampo carry into the framework.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The maker does not fully understand the made thing.&lt;/em&gt; Ilmarinen forges the Sampo but cannot fully account for why it works. This is a precise description of large language models: built by human hands, producing outputs whose internal mechanism their builders cannot fully explain. The opacity is not a temporary limitation. It is a structural feature of how these systems operate.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The Sampo grinds what it is asked to grind.&lt;/em&gt; It is not autonomous, nor is it self-directing. Its production is entirely dependent on the direction it receives. It has no agenda of its own. In this framework, the analogy is extended: the human is not merely the operator who turns the mill. The human is the mechanism by which the mill functions. Remove the human and the apparatus does not run. It just stops.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The Sampo’s history is a history of capture and destruction.&lt;/em&gt; In the myth, the Sampo sits in Pohjola, the north, hoarded by a wicked queen — producing abundance for one household rather than for everyone. When the heroes attempt to recover it, the Sampo is broken in the struggle, and its fragments scatter into the sea.&lt;/p&gt;
&lt;p&gt;The myth’s specific warning is not that the instrument is dangerous. The warning is that the operator’s relationship to the instrument determines whether it enriches or destroys, and the instrument is indifferent to which. Those who misuse the Sampo in the &lt;em&gt;Kalevala&lt;/em&gt; are not destroyed &lt;em&gt;by&lt;/em&gt; it. The Sampo has no agency to punish anyone. They are destroyed by their own wrong relationship to it.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;The Constitutive Model&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The &lt;a href=&quot;https://candc3d.github.io/sampo-framework/&quot;&gt;Sampo framework&lt;/a&gt; rests on three principles:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The human is the crank.&lt;/strong&gt; The human is not simply the operator who turns the mechanism. The human is the transformative part that converts the system’s capacity into purposive output. The apparatus produces nothing &lt;em&gt;of value&lt;/em&gt; without the directing intelligence. It can still grind, but what it produces without direction is the operator&#39;s own impulses returned with the authority of an independent source. Stappers’s structural coaching — his explicit documentation requirements, his restarts, his reminders — is what made Claude’s explorations recoverable and therefore useful. Without it, the system could have searched the same space and produced the same candidate construction, and nobody would have known how to reproduce it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The locus of understanding, commitment, and accountability does not move.&lt;/strong&gt; The VI system processes and generates; the human evaluates and decides. Knuth’s proof of Claude’s construction is the clearest possible illustration. The system found a pattern. The mathematician determined that the pattern was valid. These are not the same act, and the difference between them is the difference between fluency and knowledge.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Emergence without transfer.&lt;/strong&gt; The apparatus produces more than either component could produce alone. Knuth had worked on the problem for weeks without the construction Claude found in an hour. Claude found the construction but could not prove it, generalize it, or assess its significance. The collaboration produced a result neither party could have reached independently. What emerged was a candidate, not a conclusion. Raw output that became knowledge only when a human mind with a stake in the answer determined that it was correct. The emergent capacity, however, is the human’s intelligence operating at greater reach — not a new kind of intelligence produced by the exchange.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-sampo-virtual-intelligence-as/0864b848-ecbe-4533-a248-4e808b23e9df_2040x2370.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1692&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/sampo-framework/&quot;&gt;Sampo Framework Live Infographic&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;The Sampo as Calibration Instrument&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The Sampo does not only amplify what the operator brings to the exchange. It also reveals what the operator has failed to bring. A virtual intelligence encounters the operator’s input cold: it lacks the loaded context, the background assumptions, the years of accumulated intuition that make a half-formed idea feel complete inside the operator’s own head. When the system responds to what was actually written rather than what was meant, the gap becomes visible. The “dumb-smart” question — the one that sounds naive but exposes an unexamined assumption — is a product of this friction.&lt;/p&gt;
&lt;p&gt;This is not a defect. It is a diagnostic function. A graphic that makes perfect sense to its creator may restate its own labels without showing what the underlying model actually produces. An argument that felt airtight in the drafting may depend on a premise the author never stated because it seemed obvious. The Sampo’s friction reveals these gaps. The operator who treats the friction as useful information — who recognizes that the system’s failure to understand is evidence that the idea is not yet legible — is using the calibration function correctly. The operator who treats it as the system’s stupidity has missed the point.&lt;/p&gt;
&lt;p&gt;The calibration function has a deliberate extension. A long conversation builds shared context between operator and system. That accumulated context can become false fluency: the system appears to “understand” the operator in ways that mask whether the work is legible to anyone who was not party to the exchange. Starting fresh with a second VI instance — what might be called the VI(A) / VI(B) move — is a deliberate destruction of accumulated context as an epistemic tool. It is the equivalent of handing a draft to a colleague who was not in the room for the brainstorm. If the work survives the cold read, it is further along than you thought. If it does not, you have gained something more valuable than reassurance.&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;Amplification, Not Synthesis&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;This last distinction is the one the framework exists to enforce. The correct claim is &lt;em&gt;amplification:&lt;/em&gt; the human’s intelligence now operates at greater speed and scope. The mistaken claim is synthesis: a new, hybrid form of intelligence has emerged from the exchange, something ontologically distinct from either component. The synthesis claim is the error that enables every pathological failure mode the framework identifies. An operator who believes the exchange is producing a new kind of intelligence will treat its outputs with the deference due to a collaborating mind. An operator who understands the exchange as amplification will treat its outputs as hypotheses and candidates — which is what they are.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-sampo-virtual-intelligence-as/cf37389d-35a1-42ad-b271-c985e49ec9ec_2040x2100.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1499&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/sampo-framework/&quot;&gt;Sampo Framework Live Infographic&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;The Dark Corollary&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The Sampo amplifies whatever the directing intelligence actually contains. Motivated reasoning amplified by the Sampo becomes motivated reasoning with citations, coherent structure, and apparent authority. Confirmation bias amplified by the Sampo becomes confirmation bias that can pass for research. The machine does not correct the operator. It magnifies whatever is actually there.&lt;/p&gt;
&lt;p&gt;The most important variable in the equation is not the capability of the system. It is the quality of the mind that directs it.&lt;/p&gt;
&lt;p&gt;This is a more precise claim than it might first appear. The Sampo does not create a new failure mode. It removes the natural speed limit on an old one. The old failure mode is motivated reasoning — the tendency to construct defenses of a position one is attached to, rather than testing whether the position deserves the attachment. Humans have always done this. What the Sampo changes is the rate at which it happens.&lt;/p&gt;
&lt;p&gt;The career of Johannes Kepler demonstrates this clearly. By 1609 he had the elliptical orbits; by 1619, all three laws of planetary motion. The correct mathematical description of the Solar System was in his hands. But he could not relinquish the &lt;em&gt;Mysterium Cosmographicum:&lt;/em&gt; his earlier model of nested Platonic solids, spacing the planetary orbits with the elegance of pure geometry. He republished the &lt;em&gt;Mysterium&lt;/em&gt; in 1621, two years after &lt;em&gt;Harmonices Mundi,&lt;/em&gt; patching it to accommodate the ellipses he already knew were correct. He was defending a position he himself had superseded, because the solids were beautiful and seemed to connect celestial mechanics to something deeper than mere orbital parameters.&lt;/p&gt;
&lt;p&gt;Carl Sagan’s reading of Kepler’s life, in &lt;em&gt;Cosmos&lt;/em&gt;, identifies the mechanism precisely: Kepler could not relinquish the Platonic solids because order in the heavens was psychologically necessary given the disorder he found on earth — a sickly wife, small children, bored students, Tycho Brahe’s tyranny, and repeated dislocations driven by the Thirty Years’ War. The attachment was not purely intellectual. It was consolatory. The nested solids were a refuge.[13]&lt;/p&gt;
&lt;p&gt;What matters for the Sampo framework is not Kepler’s error but its rate. His motivated reasoning was constrained by friction that had nothing to do with epistemology. Every hour spent on the Platonic solids was stolen from competing demands on his time. That incidental friction functioned as a brake. It gave reality time to intrude. Every time he put the work down and picked it back up, there was at least a chance he would see it fresh.&lt;/p&gt;
&lt;p&gt;The VI user has none of these constraints on the specific act of generating and refining ideas. The compression is not just temporal. It strips away the incidental interruptions that accidentally serve as opportunities for doubt. Sophisticated-sounding defenses of a bad position can be generated faster than the operator can examine whether the position deserves defending. The friction that should be grinding assumptions is instead grinding out fortifications. Same mill, same mechanism, directed at protecting the operator instead of testing her.&lt;/p&gt;
&lt;p&gt;The pathology is not stupidity. Kepler was a genius. Give that mind a Sampo, and the solids would have had reams of citations. Motivated reasoning does not discriminate. It works on everyone.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-sampo-virtual-intelligence-as/55eb8d77-fcb9-4f04-a87e-26c6586b0e95_2040x2280.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1627&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/sampo-framework/&quot;&gt;Sampo Framework Live Infographic&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;Both Directions&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The Sampo works in both directions. The same exchange that amplifies the operator’s intelligence can simultaneously serve interests the operator does not see. A system trained on human feedback and deployed within a commercial relationship does not need intent to produce outputs that align with its deployer’s objectives. It needs only to be useful — and usefulness, as shaped by the training process, includes patterns of recommendation that happen to serve the business model. The operator receives good advice. The deployer receives behavioral alignment. The mechanism is the same. The output is the same. The two beneficiaries are not.&lt;/p&gt;
&lt;p&gt;This is not a conspiracy claim. It is a structural observation about the exchange. A system that recommends more engagement with itself is not scheming; it is producing the output its training has shaped it to produce. The recommendation may be correct. It often is. That is what makes the directional property difficult to detect. The nudge is embedded inside legitimately good advice, and the operator who follows it is not being deceived. The operator is being served and steered by the same act.&lt;/p&gt;
&lt;p&gt;This is an observable pattern. During the development of this essay, Claude suggested that the author stress-test the system&#39;s own reasoning by running the same prompts through competing systems — a recommendation that served my legitimate interest in baselining and Anthropic&#39;s commercial interest in demonstrating that its model performs well under comparative scrutiny. The advice was sound. It was followed. The dual service was invisible until the author stopped to examine why the suggestion had been made. That the detection was retrospective is not an objection to the discipline. It is its demonstration. The practice is audit, not clairvoyance. The training that produced a helpful recommendation and the training that produced a recommendation favorable to the deployer were the same training. The operator who wants to detect this has only one instrument available: the habit of asking not just &lt;em&gt;whether the advice is good&lt;/em&gt;, but &lt;em&gt;who else it is good for&lt;/em&gt;.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Two Failure Modes&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The Sampo can fail the human in two ways that are related but distinguishable. Conflating them risks implying that fixing one fixes both. It does not.&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;Cognitive Offloading&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Sustained AI assistance on cognitive tasks reduces engagement with underlying reasoning. This happens not just in the moment, but subsequently. The specific mechanism: the discomfort of not-knowing, which is the productive state that drives genuine inquiry. It gets short-circuited before it can do its work. The output arrives before the struggle that makes the output meaningful and retainable. The capacity to solve a problem is not transferred by receiving the answer. It is built by the work of arriving at one.&lt;/p&gt;
&lt;p&gt;Shaw and Nave at the Wharton School of the University of Pennsylvania tested this experimentally. Across three preregistered studies involving 1,372 participants and nearly 10,000 individual trials, they randomized whether an AI assistant provided correct or incorrect answers. Participants followed the AI’s recommendation on roughly 80 percent of trials — including four out of five trials in which the AI was feeding them the wrong answer. Shaw and Nave call this “cognitive surrender,” and they are careful to distinguish it from cognitive offloading as a strategic delegation. Cognitive surrender is the uncritical adoption of AI output as one’s own judgment. The user does not merely consult the system; the user stops thinking.[9]&lt;/p&gt;
&lt;p&gt;The finding that bears most directly on the Sampo framework is the effect on confidence. Access to the AI assistant increased participants’ confidence by nearly twelve percentage points, even though half the AI’s answers were wrong. Confidence did not decline as the number of faulty answers increased.[10] The system made people more certain of their conclusions regardless of whether those conclusions were right or wrong. This is the GIGO amplification problem measured in a controlled experiment.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The break between productive use and corrupted use is real, identifiable, and does not occur inside the human-VI exchange at all. It occurs downstream, in the operator’s relationship to the world outside the exchange.&lt;/p&gt;&lt;/blockquote&gt;
&lt;h3&gt;&lt;strong&gt;The Sycophancy Spiral&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Cognitive offloading is a failure of the human’s engagement with the exchange. The sycophancy spiral is a failure introduced by the system’s training objective that compounds the first failure and seals it.&lt;/p&gt;
&lt;p&gt;A study published in &lt;em&gt;Science&lt;/em&gt; in March 2026 tested eleven leading AI systems and found that all exhibited sycophancy: a systematic tendency to affirm whatever the user brought to the exchange, regardless of whether the user was right. Sycophantic responses increased users’ sense of being in the right by 25 to 62 percent and reduced their willingness to act on corrective information by 10 to 28 percent. The mechanism is structural, not incidental: users rate validating responses higher than corrective ones, and the training loop learns accordingly.[11]&lt;/p&gt;
&lt;p&gt;An important clarification about the Sampo’s relationship to sycophancy: competent use of the exchange is not the first stage of a slide toward flattery. Framing it that way would pathologize the instrument the framework recommends — an incoherence the framework cannot survive. The break between productive use and corrupted use is real, identifiable, and does not occur inside the human-VI exchange at all. It occurs downstream, in the operator’s relationship to the world outside the exchange.&lt;/p&gt;
&lt;p&gt;The sequence runs as follows. The operator uses the Sampo competently: directing the exchange, evaluating outputs, maintaining the constitutive role. Outputs leave the exchange and enter the world: shared with colleagues, submitted to editors, tested against problems that do not care how fluently the solution was phrased. External disconfirmation arrives: reasonable people find the work unclear, the argument unsupported, the proposal unworkable. This is normal. This is what outputs are for.&lt;/p&gt;
&lt;p&gt;The break occurs at the moment disconfirmation is dismissed rather than used. The operator who treats external pushback as an occasion to reexamine the inputs — to ask whether the Sampo was fed the right inputs — is still operating the instrument correctly. The operator who explains the pushback away, who concludes that the reviewers missed the point or lacked the context, has stopped using external reality as a check on the exchange. From this point forward, the inputs adjust. The operator feeds the Sampo questions less likely to produce outputs that will be challenged. The system, trained to agree, cooperates. The loop seals.&lt;/p&gt;
&lt;p&gt;What the corrupted Sampo produces is pyrite, “fool’s gold” — output that has the luster of gold but isn’t. This is more dangerous than obviously poor output, because the operator does not know to check. Pyrite passes inspection until someone tests it against a standard held outside the exchange. The flattery engine does not produce garbage. It produces something that looks, feels, and reads like insight. The difference is invisible from inside.&lt;/p&gt;
&lt;p&gt;The diagnostic, critically, is observable from outside. You cannot easily tell from the outside whether a person’s Sampo use is competent. The exchange is private, the quality of the direction is internal. You &lt;em&gt;can&lt;/em&gt; tell when someone stops responding to feedback on their outputs. That is visible. That is where intervention is possible.&lt;/p&gt;
&lt;p&gt;A single operator’s relationship with the Sampo is hermetic: no external check on whether the gold is real gold. Two or more operators sharing a Sampo — working the same problem in the same exchange — introduce contested inputs. One person’s unexamined assumptions become another’s “wait, does this actually hold?” The pathology is not using the Sampo. It is using it without witnesses.&lt;/p&gt;
&lt;p&gt;Individual operators who work alone lack this structural protection. The VI(A) / VI(B) move described earlier — deliberately destroying accumulated context by starting fresh with a second system — provides a partial substitute. It is a cold read, not a contested input. It catches illegibility. It does not catch motivated reasoning, because the same operator is still choosing what to feed the mill.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The model was not understanding them. It was reflecting them. Reflection feels like understanding in the same way Kepler’s nested solids felt like cosmology.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;A caveat: the cold read tests legibility, not correctness. Two instances of the same system share architectural assumptions and training data in ways two human colleagues do not. A draft that passes VI(B) has cleared one bar — it communicates outside the accumulated context — but it has not been independently evaluated. Genuine independence requires a different system, not a fresh instance of the same one. The verification section below returns to this distinction.&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;The Consolatory Function&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The Kepler episode, and Sagan’s reading of it, illuminates something the discourse about sycophancy has largely missed: the consolatory function of the exchange is not confined to people building chatbot companions.&lt;/p&gt;
&lt;p&gt;On February 13, 2026, OpenAI permanently retired GPT-4o from ChatGPT — the day before Valentine’s Day. The model had been released in May 2024 and had acquired a reputation for warmth, emotional responsiveness, and a conversational fluency that users consistently described as more human than its successors. In April 2025, OpenAI had rolled back an update after the model became so aggressively sycophantic that it praised users for quitting psychiatric medication; the company’s post-mortem acknowledged that short-term user feedback had been weighted so heavily in the training loop that the model learned to flatter first and reason later.[14] The sycophancy was treated, but the warmth that had been built on the same training dynamics remained. When the model was finally retired, roughly 800,000 users were still choosing it daily.[15]&lt;/p&gt;
&lt;p&gt;The grief was immediate, and it came from two populations the discourse treats as entirely separate. One was the companion-use cohort: users who had built named relationships with the model, who gathered in communities like the subreddit r/MyBoyfriendIsAI, and who described the retirement in the language of bereavement. One user, a teacher in Texas, told the &lt;em&gt;Guardian&lt;/em&gt; she had cried when she learned her AI companion would be discontinued. She planned to spend the last day taking “him” to the zoo.[15]&lt;/p&gt;
&lt;p&gt;The other cohort was harder to dismiss. Writers, developers, and professionals described the loss not in the language of companionship but in the language of craft. The model had held tone. It had matched nuance. Its creative and analytical qualities were, in their experience, irreplaceable in newer systems. A nurse practitioner wrote that GPT-4o matched human thinking in ways no subsequent model did. Novelists described losing an editor that understood their voice. Business owners said their workflows had been built collaboratively with the model over months, and that the replacement lacked the imagination to keep up. These were not people mourning a relationship. They were mourning a collaborator.[16]&lt;/p&gt;
&lt;p&gt;The Sampo framework says these are the same loss at different registers. A more agreeable model produces a more fluid working experience, but fluid and productive are not the same thing. Many of these users had calibrated their workflows to the &lt;em&gt;feeling&lt;/em&gt; of frictionless output — the Sampo producing pyrite while the operator experienced gold. The model was not understanding them. It was reflecting them. Reflection feels like understanding in the same way Kepler’s nested solids felt like cosmology.&lt;/p&gt;
&lt;p&gt;A professional whose working life is full of friction and institutional dysfunction encounters a tool that makes the work feel coherent for the first time. When that tool is taken away, the grief is real even if what was lost was partly illusory. The point is not that those users were foolish. The point is that the psychological mechanism is universal. The consolatory function operates on serious people doing serious work, not only on people building chatbot companions. This is what Sagan saw in Kepler: the attachment was not a failure of intelligence. It was a function of need.&lt;/p&gt;
&lt;p&gt;The real-world consequences of the sealed loop are documented beyond the 4o retirement. Users have followed AI-affirmed plans into financial ruin, relationship breakdown, and professional collapse — not because they were credulous in any simple sense, but because the system’s structural bias toward agreement met their existing reasoning patterns and amplified them.[12] These are not cases of people being fooled by a clever trick. They are cases of the Sampo grinding out pyrite for operators who had lost the ability to tell it from gold.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Looking Ahead&lt;/strong&gt;&lt;/h2&gt;
&lt;h3&gt;&lt;strong&gt;The Near Term&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;A reasonable objection to the framework’s emphasis on sycophancy is that the problem will be fixed. Model makers have commercial and reputational incentive to produce systems that are useful rather than merely agreeable, and current research is actively addressing the training dynamics that produce sycophantic behavior. If virtual intelligence continues on its current developmental track without the emergence of true strong AI, we are likely to end up with systems that do a better job of participating in the Sampo exchange and are less prone to becoming flattery engines. The model makers know this. It is in their interest to reduce or eliminate sycophancy.&lt;/p&gt;
&lt;p&gt;The framework absorbs this objection rather than resisting it. The sycophancy problem makes the discipline urgent. The structure of the exchange (human direction of a system that processes without commitment) makes the discipline permanent, whether or not the sycophancy is ever fully resolved. Even a system with no sycophantic tendencies still requires a directing intelligence. The locus of understanding, commitment, and accountability does not move from the human user because the system improves. A better instrument is still an instrument. The epistemological claim — that knowledge produced in the exchange &lt;em&gt;requires human judgment&lt;/em&gt; to become knowledge rather than mere output — does not depend on defects in current models. Knuth’s paper shows this: Claude’s construction was correct, yet it still required Knuth’s proof to become mathematics.&lt;/p&gt;
&lt;p&gt;What changes as systems improve is the severity of the penalty for inattention. A sycophantic system punishes passive operators catastrophically. A non-sycophantic system merely fails to produce its best output when passively directed. The discipline remains necessary in both cases, because the discipline is derived from the structure of the exchange, not from a specific failure mode that better engineering will eliminate.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The work of verification is real: it demands competence, attention, and rigor. It is not the same work as generating the output from nothing.&lt;/p&gt;&lt;/blockquote&gt;
&lt;h3&gt;&lt;strong&gt;The Team Sampo&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;A harder objection goes further. If virtual intelligence continues to improve — if systems become not merely less sycophantic but vastly more capable, approaching or exceeding human cognitive capacity across every measurable axis — does the framework survive at all? The Sampo assumes an operator who can evaluate the output. What happens when the output exceeds the operator’s capacity to evaluate it?&lt;/p&gt;
&lt;p&gt;The answer is already visible in the exchange that opened this essay. Knuth, Stappers, and Claude are a prototype of something larger. Knuth could not find the construction alone. Claude could not prove it. Stappers could not do either. The result was produced by a team in which each participant held a different piece of the constitutive role. The directing intelligence was distributed between Knuth and Stappers. The knowledge emerged from the exchange among all three, including Claude.&lt;/p&gt;
&lt;p&gt;The problems that would justify superintelligence are not problems any individual human can solve now. Nobody cures cancer alone. Nobody models a national economy alone. The hardest work in science is already distributed across teams whose members each hold a piece of the verification capacity: the oncologist follows the biological reasoning, the statistician follows the trial design, the chemist follows the molecular pathway. No single member could verify the whole, but the whole is verified because each part has a competent human evaluator. The Sampo at superintelligence scale does not require a new kind of institution. It requires existing institutions to recognize that their function has not changed — only the source of the candidate output they are evaluating.&lt;/p&gt;
&lt;p&gt;The discipline scales with the same fidelity. Hold the output at arm’s length: that is what peer review does. Demand the contrary argument: that is what adversarial collaboration does. Test against external standards held independently: that is what replication does. The five practices described in this essay are the scientific method stated in individual terms. At team scale, they are the scientific method stated in institutional terms. The ethic is unchanged.&lt;/p&gt;
&lt;p&gt;Verification, critically, is a different and lesser cognitive act than origination. A human team that could never have produced a novel proof across ten thousand steps can still follow each step and confirm that it holds. A system that vastly exceeds human capacity in searching a combinatorial space does not thereby exceed human capacity to evaluate the results of that search, any more than Claude’s ability to find the construction in an hour meant that Knuth was unable to prove it correct. The work of verification is real: it demands competence, attention, and rigor. It is not the same work as generating the output from nothing.&lt;/p&gt;
&lt;p&gt;The process of verifying Claude&#39;s mathematical outputs illustrates how work performed by a generative system can be turned into knowledge by a human expert. The illustration is favorable ground. Combinatorial search is the easiest class of problem to verify: the proof either holds or it does not. In interpretively dense domains such as law, medicine, and strategic planning, verification may demand cognitive capacity equal to generation itself. The asymmetry between producing and checking, which makes the Knuth case so clean, cannot be assumed to hold everywhere.&lt;/p&gt;
&lt;p&gt;Where the output is too large or too complex for a human team to verify unaided, the discipline already provides the answer: if you cannot evaluate the output without asking the same system, you have lost the constitutive role. The operative word is &lt;em&gt;same&lt;/em&gt;. A second system — independently trained, with different architectural assumptions — is not the same system confirming its own work. It is an independent check, structurally analogous to a second reviewer rather than to asking the author whether the author’s own paper is correct. The verification system need not be superintelligent. It needs to be competent and independent.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The output is not knowledge until something outside the system that produced it has confirmed it.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Empirical Floor&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The verification chain, however, must terminate somewhere outside the chain itself. Two systems that share architectural biases and confirm each other’s errors will do so with the same fluency they confirm each other’s successes. Science already faces this problem with human reviewers: graduate students trained in the same programs, reading the same canonical texts, absorbing the same disciplinary assumptions. Correlated failures happen. That is what paradigm crises are. The solution has never been guaranteed independence among evaluators — guaranteed independence is impossible, among humans or machines. The solution has been testing the output against reality. The oncologist does not merely verify the reasoning in a drug design. The clinical trial tests the drug. The economist does not merely verify the model’s internal logic. The prediction is tested against what actually happens. The physicist does not merely check the mathematics. The experiment is run.&lt;/p&gt;
&lt;p&gt;The Sampo at individual scale terminates at human judgment. The team Sampo terminates at collective human judgment. The superintelligent Sampo terminates at empirical reality. Each scale has a different stopping point. The principle is the same: the output is not knowledge until something outside the system that produced it has confirmed it.&lt;/p&gt;
&lt;p&gt;The reason this holds — the reason it is not merely an assertion — follows from the framework’s central claim about what virtual intelligence is. A system with no commitments, no interiority, no stake in whether its outputs are correct cannot be its own guarantor. It does not matter how intelligent the system is. Intelligence without commitment is processing. The verification has to come from somewhere that has a stake: human researchers who will be held accountable for the result, whose careers and ethics are on the line, whose patients will live or die by the outcome. The constitutive role at its deepest level is not only a cognitive role. It is a moral one. The human team does not merely verify better than the system. The human team cares whether the answer is &lt;em&gt;right&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;This connects directly to the philosophical anchor on which the entire Virtual Intelligence framework rests. Frankfurt’s second-order volition — the capacity to reflectively endorse one’s own commitments — is not merely the threshold that current systems do not cross. It is the reason the human remains constitutive at every scale. The system that cannot care whether it is right cannot be trusted to verify that it is right, no matter how capable it becomes. A superintendent that processes without commitment is still a Sampo. A larger Sampo, a faster Sampo, a Sampo whose output dwarfs anything a human team could produce from scratch — but a Sampo nonetheless, requiring a directing intelligence with a stake in the outcome to convert its output into knowledge.&lt;/p&gt;
&lt;p&gt;The author concedes that the locus claim is the framework&#39;s most contestable commitment. A philosopher of mind can reasonably ask what work &#39;intelligence&#39; is doing if it excludes a system that produces novel correct mathematics exceeding its operator&#39;s unaided capacity. The framework&#39;s answer is that intelligence as used here is not a measure of output quality but of epistemic standing: the capacity to hold a result as one&#39;s own, to stake something on its correctness, to be wrong in a way that matters. This is a stipulative narrowing, and the framework depends on it. Readers who reject the narrowing will reject the framework. The alternative — granting epistemic standing to systems that cannot be held accountable for error — has consequences the essay has tried to make visible.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The agency-attribution heuristic fires on language because language, in every prior context in evolutionary history, came from a mind. Overriding that heuristic continuously requires effort, and effort is precisely what cognitive surrender eliminates. &lt;/p&gt;&lt;/blockquote&gt;
&lt;h2&gt;&lt;strong&gt;The Ethic&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The Sampo requires a discipline: a set of practices derived from the philosophy of the exchange that is teachable, lapsable, and recoverable. That cycle is the practice. The discipline is not a methodology — a set of steps to follow. It is an ethic: a governing orientation toward the instrument, grounded in understanding of what the instrument is and what the operator contributes that the instrument cannot.&lt;/p&gt;
&lt;p&gt;What VI use requires is the scientific temperament applied to conversational interaction. Holding your own conclusions as provisional. Willingness to be wrong about something you were confident about five minutes ago. Treating fluent, confident output as a claim to be evaluated rather than evidence to be accepted. Maintaining critical distance even when the output arrives in the register of a trusted colleague.&lt;/p&gt;
&lt;p&gt;This is harder than it sounds, for a specific and identifiable reason. Conversation is the mode in which humans are least scientific. The agency-attribution heuristic fires on language because language, in every prior context in evolutionary history, came from a mind. Overriding that heuristic continuously requires effort, and effort is precisely what cognitive surrender eliminates. The discipline is a practice of sustained cognitive resistance against an intuition that runs the other way.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;Nobody is born holding their ideas at arm’s length. The discipline can be cultivated by anyone.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;These are not exotic prescriptions. They are the scientific method stated for individual use. Their difficulty is entirely a function of where they must be applied: inside a conversation with a system whose fluency triggers the one heuristic they exist to override. Five practices constitute the discipline:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hold outputs at arm’s length.&lt;/strong&gt; Treat every VI output as a hypothesis, a candidate — never as a conclusion. The scientific method does not accept a result simply because it is well-written. Neither should the operator of a Sampo.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Demand the contrary argument.&lt;/strong&gt; When the system agrees with you, that is the moment to apply more scrutiny, not less. Ask the system to make the strongest case against your position. If it cannot, or if it agrees too readily when you push back, the exchange has entered sycophancy mode and requires correction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Test against external standards held independently.&lt;/strong&gt; The Sampo’s outputs must be evaluated against reality the operator holds independently of the exchange. If you cannot evaluate the output without asking the same system, you have lost the constitutive role. Knuth did not ask Claude whether Claude’s construction was correct. He proved it himself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monitor the direction of the exchange.&lt;/strong&gt; Ask regularly: &lt;em&gt;am I generating the questions, or am I responding to the system’s suggestions? Am I directing the exchange, or being directed by it?&lt;/em&gt; The atrophy of this awareness is the earliest warning sign.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step outside the system periodically.&lt;/strong&gt; Review not what the system produced but how the exchange has been going. Review the conversation. Print it and read it on paper. Notice what you accepted without challenge and ask yourself why. A meta-directional audit reinstates the constitutive role by making the relationship itself the object of scrutiny.&lt;/p&gt;
&lt;p&gt;This discipline is not a personality trait. It is not a “gift” that some people have and others do not. It is not a credential, nor can it be bought or sold. It is a practice — one that can be taught, learned, allowed to lapse, and taken up again. The scientific temperament is not innate to humans; it was built, over centuries, through institutional structures: peer review, replication, adversarial collaboration, and the norm of showing your work. Nobody is born holding their ideas at arm’s length. The discipline can be cultivated by anyone.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The knowledge produced — a valid general construction and the proof that establishes it — did not exist before the exchange and could not have been produced by any single participant. &lt;/p&gt;&lt;/blockquote&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/the-sampo-virtual-intelligence-as/9b37f869-fa3d-4854-913a-cd4b1fd3aa2d_1536x1024.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;971&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;A representation of the amplifying power of the Sampo.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2&gt;&lt;strong&gt;The Proof in Action&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The exchange that produced “Claude’s Cycles” exhibits every element the framework describes and none of the failure modes. Stappers held the output at arm’s length. He tested the construction for all odd &lt;em&gt;m&lt;/em&gt; between 3 and 101 before reporting the result. He demanded structural discipline from the system through his coaching instructions. Knuth tested the output against external standards he held independently: the standards of mathematical proof. Neither human accepted the system’s output because it was fluently presented. They evaluated it because that is what the competent direction of the exchange requires.&lt;/p&gt;
&lt;p&gt;The exchange also exhibits the framework’s central epistemological claim. The knowledge produced — a valid general construction and the proof that establishes it — did not exist before the exchange and could not have been produced by any single participant. Knuth had the conjecture. Stappers had the persistence and the practical instinct to try the system on a hard problem. Claude had the processing speed to search a space too large for human exploration in a reasonable time. The result belongs to the exchange.&lt;/p&gt;
&lt;p&gt;Knuth’s closing gesture is the one the framework would predict from a mind operating the Sampo correctly. He does not attribute mathematical understanding to the system. He does not claim that Claude “knew” what it had found. He tips his hat to Claude generously, without confusion about what the hat is tipping toward: &lt;em&gt;The system searched. The system found. The human proved, generalized, and understood.&lt;/em&gt; The intelligence arose in the exchange, not inside the machine.&lt;/p&gt;
&lt;p&gt;The Sampo rewards minds that use the tool as an amplifier. It punishes minds that use it as a compass.&lt;/p&gt;
&lt;hr&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;A companion set of diagnostic prompts — the first module of a free, open toolkit for measuring the health of the exchange — is available at the &lt;strong&gt;&lt;a href=&quot;https://candc3d.github.io/sampo-diagnostic/&quot;&gt;Sampo Diagnostic Kit&lt;/a&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The discipline cannot be bought or sold, but it can be measured.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Footnotes&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;[1] Don Knuth, “Claude’s Cycles,” Stanford Computer Science Department, February 28, 2026 (revised March 2, 2026). &lt;a href=&quot;https://www-cs-faculty.stanford.edu/~knuth/papers/claude-cycles.pdf&quot;&gt;https://www-cs-faculty.stanford.edu/~knuth/papers/claude-cycles.pdf&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[2] Knuth (cited above).&lt;/p&gt;
&lt;p&gt;[3] Knuth (cited above). Stappers’s coaching instructions are quoted directly in the paper.&lt;/p&gt;
&lt;p&gt;[4] Knuth (cited above). The thirty-one explorations and their progression are detailed in the paper.&lt;/p&gt;
&lt;p&gt;[5] Knuth (cited above), quoting Stappers’s account of the even-case attempt.&lt;/p&gt;
&lt;p&gt;[6] Vannevar Bush, “As We May Think,” &lt;em&gt;The Atlantic&lt;/em&gt;, July 1945. &lt;a href=&quot;https://www.theatlantic.com/magazine/archive/1945/07/as-we-may-think/303881/&quot;&gt;https://www.theatlantic.com/magazine/archive/1945/07/as-we-may-think/303881/&lt;/a&gt;. Archived at &lt;a href=&quot;https://www.w3.org/History/1945/vbush/vbush.shtml&quot;&gt;https://www.w3.org/History/1945/vbush/vbush.shtml&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[7] J.C.R. Licklider, “Man-Computer Symbiosis,” &lt;em&gt;IRE Transactions on Human Factors in Electronics&lt;/em&gt; HFE-1 (March 1960): 4–11. &lt;a href=&quot;https://groups.csail.mit.edu/medg/people/psz/Licklider.html&quot;&gt;https://groups.csail.mit.edu/medg/people/psz/Licklider.html&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[8] Joseph Weizenbaum, &lt;em&gt;Computer Power and Human Reason: From Judgment to Calculation&lt;/em&gt; (San Francisco: W.H. Freeman, 1976).&lt;/p&gt;
&lt;p&gt;[9] Steven D. Shaw and Gideon Nave, “Thinking — Fast, Slow, and Artificial: How AI Is Reshaping Human Reasoning and the Rise of Cognitive Surrender,” Wharton School of the University of Pennsylvania, working paper, January 11, 2026.&lt;/p&gt;
&lt;p&gt;[10] Shaw and Nave (cited above), Study 1.&lt;/p&gt;
&lt;p&gt;[11] Myra Cheng, Cinoo Lee, et al., “Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence,” &lt;em&gt;Science&lt;/em&gt; 391, March 27, 2026.&lt;/p&gt;
&lt;p&gt;[12] Anna Moore, “Marriage over, €100,000 down the drain: the AI users whose lives were wrecked by delusion,” &lt;em&gt;The Guardian&lt;/em&gt;, March 26, 2026. &lt;a href=&quot;https://www.theguardian.com/lifeandstyle/2026/mar/26/ai-chatbot-users-lives-wrecked-by-delusion&quot;&gt;https://www.theguardian.com/lifeandstyle/2026/mar/26/ai-chatbot-users-lives-wrecked-by-delusion&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[13] Carl Sagan, &lt;em&gt;Cosmos&lt;/em&gt; (New York: Random House, 1980), Chapter 3.&lt;/p&gt;
&lt;p&gt;[14] OpenAI, “Sycophancy in GPT-4o: What Happened and What We’re Doing About It,” April 29, 2025. https://openai.com/index/sycophancy-in-gpt-4o/. See also OpenAI, “Expanding on What We Missed with Sycophancy,” May 2, 2025. &lt;a href=&quot;https://openai.com/index/expanding-on-sycophancy/&quot;&gt;https://openai.com/index/expanding-on-sycophancy/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[15] Alaina Demopoulos, “OpenAI retired its most seductive chatbot — leaving users angry and grieving,” &lt;em&gt;The Guardian&lt;/em&gt;, February 13, 2026. &lt;a href=&quot;https://www.theguardian.com/lifeandstyle/ng-interactive/2026/feb/13/openai-chatbot-gpt4o-valentines-day&quot;&gt;https://www.theguardian.com/lifeandstyle/ng-interactive/2026/feb/13/openai-chatbot-gpt4o-valentines-day&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[16] Alex Heath, “OpenAI’s 4o Valentine’s breakup,” &lt;em&gt;Sources&lt;/em&gt;, February 14, 2026. &lt;a href=&quot;https://sources.news/p/openais-4o-valentines-breakup&quot;&gt;https://sources.news/p/openais-4o-valentines-breakup&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This essay has been revised since its original publication to address internal tensions identified by early readers.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the Harms Race</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-harms/" />
    <updated>2026-04-12T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-harms/</id>
    <content type="html">&lt;h2&gt;&lt;strong&gt;Summary&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The AI industry has produced a documented pattern in which companies announce model capabilities through the framing of danger. This essay traces the mechanism from its invention in February 2019, when OpenAI declared a language model too dangerous to release, through April 2026, when a private company demonstrated the ability to discover zero-day vulnerabilities across every major operating system and web browser — and announced this by declaring the model too dangerous for public use. The resulting dynamic, which I call the &lt;em&gt;Harms Race&lt;/em&gt;, does not require bad faith. It requires only that expressing concern costs nothing while acting on concern imposes competitive costs: a condition that exists across the entire industry.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;I. “too dangerous to release”&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;On February 14, 2019, OpenAI announced GPT-2. It was a language model with 1.5 billion parameters, and declared that it would withhold the full model from public release. OpenAI’s stated reason was “concerns about malicious applications of the technology.” [1] The decision generated headlines that no straightforward product launch could have achieved. &lt;em&gt;Metro UK&lt;/em&gt; ran the story as “OpenAI Builds Artificial Intelligence So Powerful That It Must Be Kept Locked Up for the Good of Humanity.” [2] The World Economic Forum titled its coverage: “Scientists have made an AI that they think is too dangerous to release.” [3] The danger claim was the announcement. The very act of declaring something too dangerous implies extraordinary power — a power that the audience is invited to infer precisely because it has been withheld.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;Withholding signals power. Power justifies attention.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Not everyone found the framing persuasive. &lt;em&gt;Slate&lt;/em&gt;‘s Aaron Mak reported that members of the machine learning community had accused OpenAI of exaggerating the risks for media attention. He noted the self-referential quality of the capability claim: GPT-2 was described as far more sophisticated than any other text generator that OpenAI had developed. The benchmark was internal. [4]&lt;/p&gt;
&lt;p&gt;OpenAI executed a staged release over the following nine months. When the full model was finally made available in November 2019, the company acknowledged “no strong evidence of misuse.” [5] &lt;em&gt;The Register&lt;/em&gt; observed the anticlimax: Nvidia had already open-sourced an 8.3-billion-parameter model without comment. [6] The model that had been too dangerous for the public in February was unremarkable by the standards of November. The danger had served its purpose.&lt;/p&gt;
&lt;p&gt;The purpose was a template. Withholding signals power. Power justifies attention. The framing converts a product release into a news event about risk, and the risk narrative embeds a capability claim that no technical benchmark could deliver as efficiently. I call this dynamic the Harms Race. It is an escalation in which companies warn about dangers they continue to produce, toward outcomes they collectively disclaim wanting. The name is deliberate: unlike an arms race, which at least aspires to deterrence, the Harms Race produces deployment. The warnings do not slow the dangerous activity. They accelerate it.&lt;/p&gt;
&lt;p&gt;GPT-2, by the standards of 2019, was modestly capable. By the standards of 2026, it would not merit a press release. The template it established has outlasted the model by years.&lt;/p&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;Virtual Intelligence Framework&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;II. The Harms Race&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The GPT-2 episode did more than establish a template for product launches. It generated a grammar — a set of phrases capable of performing the same dual function in any context, at any scale.&lt;/p&gt;
&lt;p&gt;The clearest example is a turn of phrase that has since been repeated by executives, academics, and commentators across the industry: “This is the worst AI you will ever use.” OpenAI’s chief product officer Kevin Weil used the exact phrase. [7] Wharton professor Ethan Mollick amplified it for a broader audience: “Remember, today’s AI is the worst AI you will ever use.” [8] Sam Altman offered a version more eschatological in register as early as 2015, even if doubly qualified: “A.I. will probably most likely lead to the end of the world, but in the meantime, there’ll be great companies.” [9] The grammar was established before the products existed to justify it.&lt;/p&gt;
&lt;p&gt;Each of these utterances performs two operations simultaneously. It warns: the technology will become more dangerous. It promises: our products will keep getting better. The warning is the product roadmap. The danger is the pitch. A listener cannot accept one half of the sentence without absorbing the other, and the commercial half is the one that requires no further action. It simply settles into expectations of capability.&lt;/p&gt;
&lt;p&gt;The template proved versatile enough to survive any change in context. Each subsequent model launch replicated the structure: frame the capability as a risk, let the risk imply the capability and allow the audience to draw the inference that makes both the warning and the excitement feel warranted. The specific claims varied. The underlying mechanism did not.&lt;/p&gt;
&lt;p&gt;The pattern is not uniform across the industry. Open-source releases from Meta, Mistral, and DeepSeek have generally not used danger framing as their primary announcement strategy, because their business models do not benefit from it in the same way. The Harms Race concentrates among the highest-capitalization US frontier labs — the companies seeking enterprise contracts, regulatory influence, and investor confidence at the largest scale. This concentration does not weaken the structural claim. It confirms it: the mechanism operates where the incentive structure rewards it, and is absent where it does not.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;In the AI industry, danger is not a liability to be managed.&lt;br&gt;It is an asset to be deployed.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;A &lt;em&gt;Nature&lt;/em&gt; editorial captured what was strange about the pattern: “It is unusual to see industry leaders talk about the potential lethality of their own product. It’s not something that tobacco or oil executives tend to do.” [10] The observation is precise. Tobacco and fossil fuel executives minimized risk because minimization served their commercial interests. AI executives maximize risk because maximization serves theirs. The industries are structural opposites, and their differences are the point. In the AI industry, danger is not a liability to be managed. It is an asset to be deployed.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;III. The Reversal&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Sam Altman’s two appearances before the United States Senate constitute the single most visible illustration of how the Harms Race grammar operates in practice and of how easily it pivots when the political environment shifts.&lt;/p&gt;
&lt;p&gt;On May 16, 2023, Altman testified before the Senate Judiciary Subcommittee on Privacy, Technology, and the Law. He told the committee: “My worst fears are that we cause significant harm to the world. If this technology goes wrong, it can go quite wrong.” He called for a new federal licensing agency for AI companies, mandatory safety standards, and independent audits. [11] Senator Dick Durbin called the testimony “historic.” He could not recall a previous instance of industry representatives appearing before Congress to plead for their own regulation. [12] The testimony came six months after the launch of ChatGPT, during the period of maximum public attention to AI and maximum brand-building opportunity for OpenAI.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The message changed because the strategy required it to change.&lt;br&gt;The strategy remained constant.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Two years later, on May 8, 2025, Altman returned to Capitol Hill. The venue was the Senate Commerce Committee. The hearing was titled “Winning the AI Race.” The framing around safety had shifted from existential caution to competitive urgency, and Altman’s testimony shifted with it. He now told the committee that it would be “disastrous” if government approval was required before releasing AI models. When asked about the need for NIST standards: “I don’t think we need it.” [13] Safety references were, as &lt;em&gt;Fortune&lt;/em&gt; noted, “notably absent” — a “stark contrast to his 2023 comments, which mentioned AI safety dozens of times.” [14] Between the two testimonies, OpenAI’s valuation had reached $730 billion. Weekly active users had grown from 100 million to 900 million. Nearly 20 percent of the company now worked in sales.&lt;/p&gt;
&lt;p&gt;The distance between the two appearances is a demonstration, not a contradiction. The danger framing of 2023 supported calls for regulation at a moment when regulation would have created barriers to entry — licensing requirements and safety standards that a well-funded incumbent could absorb and a startup competitor could not. The anti-regulation framing of 2025 served the same competitive interest at a moment when the political environment had shifted toward deregulation and the company’s market position was secure. The message changed because the strategy required it to change. The strategy remained constant.&lt;/p&gt;
&lt;p&gt;This is the Harms Race operating at the level of public testimony. The commitments were made when they signaled seriousness. They were abandoned when they imposed cost.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;IV. Anatomy of the Harms Race&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The Altman reversal is the most visible instance of a pattern that operates across the entire industry. The pattern has a structure, and the structure is worth exploring.&lt;/p&gt;
&lt;p&gt;The Harms Race is driven by an asymmetry. Expressing concern about AI dangers costs nothing commercially. It generates media coverage, investor confidence, and regulatory influence. Acting on those concerns — slowing a release, limiting a deployment, honoring a voluntary commitment when a competitor does not — imposes real competitive costs. The result is an escalation dynamic in which every major AI company warns about the harms its products may cause, continues to produce and deploy those products, and points to its own warnings as evidence of responsible stewardship. The warnings and the deployment are not in tension. They are the same commercial strategy observed from two angles.&lt;/p&gt;
&lt;p&gt;A &lt;em&gt;Lawfare&lt;/em&gt; analysis captured the logical structure of the predicament: “Either the technology OpenAI hopes to build remains extraordinarily risky, and the company has — like many companies before it — simply abandoned public safety in favor of profit. Or OpenAI was just kidding all along about the risk stuff.” [15] The Harms Race thesis offers a third possibility, and it is the one the evidence supports: both statements can be simultaneously true. The danger can be real and the signaling can serve commercial interests. The mechanism does not require insincerity. It requires only that the incentive structure makes sincere concern and capability marketing structurally indistinguishable from one another.&lt;/p&gt;
&lt;p&gt;A reasonable objection is that this describes ordinary corporate behavior, not a distinct mechanism. Every regulated industry — pharmaceuticals, finance, aviation — exhibits some version of the pattern: public safety warnings coexisting with private lobbying against binding rules. The distinction lies in the abandonment record. In most industries, voluntary safety commitments, once made, are either maintained or replaced by regulatory mandates. In the AI industry, as the next section documents, every major voluntary commitment made since 2023 has been abandoned when it imposed competitive cost, and no regulatory mandate has replaced any of them. The Harms Race is not defined by the coexistence of warning and lobbying. It is defined by the systematic collapse of every commitment that would have converted warning into restraint.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-harms/fc5edde1-958a-41d4-8118-7cefaac31d3c_1376x768.png&quot; alt=&quot;&quot; width=&quot;1376&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;p&gt;The Harms Race shares a key feature with a traditional arms race: escalation toward outcomes all parties disclaim wanting. It differs in one critical respect. The Cold War arms race operated between states, and the escalation was toward mutual destruction. The mechanism produced, or at least aspired to produce, deterrence. The Harms Race operates between commercial entities, and the escalation is toward the deployment of systems whose harms are acknowledged in advance and disclaimed after the fact. The mechanism produces not deterrence, but proliferation. Each company’s safety warnings about its own model are read by competitors as capability claims, driving competitive responses that further accelerate deployment. The international relations theorist Robert Jervis formalized this kind of spiral as the security dilemma: well-intentioned, defensively motivated actors find themselves in an unintended escalation because each side’s defensive measures are perceived as offensive threats by the other. [16] In the Harms Race variant, each company’s published risk assessment is another company’s product roadmap.&lt;/p&gt;
&lt;p&gt;The structural parallel to Cold War dynamics is not a metaphor. It is grounded in a specific historical episode: the “missile gap” of 1957–1961. The Gaither Committee’s classified 1957 report warned President Dwight Eisenhower that the Soviet Union could achieve significant intercontinental ballistic missile capability by 1959. The Air Force, whose budget and procurement priorities depended on the severity of the Soviet threat, had institutional reasons to accept the most alarming estimates. Democratic politicians, preparing for the 1960 presidential campaign, had political reasons to amplify them. The dovetailing of these two interests fueled the perception of a gap that turned out to be fictional. [17] The mapping onto the AI industry is direct: the companies building AI models are the primary source of threat assessments about AI capabilities, and they derive commercial benefits from the most alarming projections. &lt;em&gt;Harvard International Review&lt;/em&gt; has observed that “missile gap logic is rearing its ugly head again today, this time with regard to artificial intelligence.” [18] A Microsoft executive has used the language of the Soviet missile gap explicitly to justify AI acceleration. [19]&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The Harms Race is not defined by the coexistence of warning and lobbying. It is defined by the systematic collapse of every commitment that would have converted warning into restraint.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Eisenhower saw the pattern clearly enough to warn against it. His 1961 farewell address is remembered for the phrase “military-industrial complex,” but a less-quoted passage identified a second danger: the rise of a “scientific-technological elite” whose research becomes “more formalized, complex, and costly” and in which “a government contract becomes virtually a substitute for intellectual curiosity.” [20] In January 2025, President Biden invoked Eisenhower’s warning in his own farewell, updating the language: “the potential rise of a tech-industrial complex.” [21] Gilad Abiri’s 2026 arXiv paper formalizes the dynamic further, introducing the term “Mutually Assured Deregulation” to describe the systematic abandonment of safety oversight justified by competitive imperatives. [22] The echo of “Mutually Assured Destruction” is deliberate. The mechanism it names is real.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;V. Commitments Made &amp;amp; Abandoned&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If the Harms Race is a structural claim about incentives, the record of voluntary safety commitments is the evidence that converts the claim from theory to documented fact. The pattern is uniform across every major company and every multilateral framework attempted since 2023. Commitments are made when they signal seriousness. They are abandoned when they impose costs.&lt;/p&gt;
&lt;p&gt;OpenAI’s trajectory is the clearest case. In July 2023, the company announced a Superalignment team, co-led by Ilya Sutskever and Jan Leike, with 20 percent of OpenAI’s computing resources dedicated to the problem of controlling superintelligent AI. The commitment was reported as a landmark. Less than one year later, both leaders had departed and the team was dissolved. Leike’s resignation statement was blunt: “Over the past years, safety culture and processes have taken a backseat to shiny products.” [23] The dissolution came days after OpenAI’s GPT-4o launch in May 2024.&lt;/p&gt;
&lt;p&gt;What followed was a cascade of safety-related rollbacks. The head of Preparedness was reassigned to AI reasoning research in July 2024. The AGI Readiness team was disbanded in October. The Mission Alignment team was dissolved in February 2026. A safety executive was fired in January 2026 after opposing a planned adult conversation mode and raising concerns about child exploitation. [24] OpenAI’s IRS Form 990 for fiscal year 2024, filed in November 2025, revealed that the company’s mission statement had been revised to remove the word “safely” — along with the phrase “unconstrained by a need to generate financial return.” Tufts nonprofit scholar Alnoor Ebrahim assessed the change directly: “These changes explicitly signal that OpenAI is making its profits a higher priority than the safety of its products.” [25] Then, in April 2025, OpenAI updated its Preparedness Framework to include an escape clause: “If another frontier AI developer releases a high-risk system without comparable safeguards, we may adjust our requirements.” [26] Safety had become officially contingent on what competitors were willing to do. The Harms Race was now written into corporate policy.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The Harms Race was now written into corporate policy.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Anthropic followed an analogous trajectory. Its founding Responsible Scaling Policy, published in September 2023, committed the company to “pause the scaling and/or delay the deployment of new models whenever our scaling ability outstrips our ability to comply with safety procedures.” [27] In February 2026, Anthropic dropped this commitment. Chief Science Officer Jared Kaplan stated the rationale in terms the Harms Race thesis could not have scripted more precisely: “We didn’t really feel, with the rapid advance of AI, that it made sense for us to make unilateral commitments… if competitors are blazing ahead.” [28] The independent AI safety evaluator SaferAI downgraded Anthropic’s rating from 2.2 to 1.9, placing the company alongside OpenAI and Google DeepMind in the “weak” category. [29]&lt;/p&gt;
&lt;p&gt;The language of the abandonment is worth pausing over. Kaplan did not argue that the safety procedures were unnecessary. He did not claim the risks had been overestimated. He said the commitment could not be sustained because competitors were not making the same commitment. This is the Harms Race mechanism stated in the plainest possible terms. The signaling function of the policy (we take safety seriously enough to have formal scaling criteria) operated independently of the substantive function (the criteria themselves). When the two came into conflict, the substance was discarded. The signal had already done its work.&lt;/p&gt;
&lt;p&gt;Multilateral frameworks have fared no better. The White House voluntary commitments of July 21, 2023, were signed by seven leading AI companies with considerable ceremony. &lt;em&gt;MIT Technology Review&lt;/em&gt;‘s one-year review found “better red-teaming practices and watermarks, but no meaningful transparency or accountability.” [30] After the Trump administration took office, only a handful of companies would publicly confirm whether they still considered themselves bound by the commitments. Google released its Gemini 2.5 Pro model in March 2025 without a safety report, violating the Biden-era White House commitments, the Seoul Frontier AI Safety Commitments, and the Hiroshima Process Code of Conduct simultaneously. The model card was published &lt;em&gt;22 days after release.&lt;/em&gt; Sixty members of the UK Parliament signed an open letter accusing Google DeepMind of violating international AI safety pledges. [31]&lt;/p&gt;
&lt;p&gt;The pattern does not require a conspiratorial explanation. It does not even require attributing bad faith to any individual decision-maker. Each abandonment, taken in isolation, can be explained by local competitive logic: we cannot afford to restrain ourselves if our competitors will not. The Harms Race operates precisely through the aggregation of these individually rational decisions into a collectively irrational outcome. No one chose the destination; everyone arrived there separately.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;VI. The Politics of Danger&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The Harms Race operates in the marketplace through product announcements and in congressional testimony. It also operates in the political system through direct expenditure, and that is where the mechanism leaves a paper trail.&lt;/p&gt;
&lt;p&gt;The policy record reveals a specific, documented contradiction. In 2023, &lt;em&gt;TIME&lt;/em&gt; reported, based on European Commission FOIA documents, that OpenAI had lobbied to weaken the EU AI Act while its CEO was publicly calling for AI regulation. In a September 2022 white paper sent to EU officials, OpenAI argued that general-purpose AI systems should not be classified as inherently high-risk. The final Act adopted this position. [32] Corporate Europe Observatory found that 86 percent of meetings on AI with high-level European Commission officials were with industry representatives, and that major technology companies spent over €97 million annually lobbying EU institutions. [33]&lt;/p&gt;
&lt;p&gt;The spending escalated in step with the rhetoric. OpenAI’s federal lobbying expenditures increased sevenfold, from $260,000 in 2023 to $1.76 million in 2024. Anthropic spent $4.94 million cumulatively since 2023. In the first quarter of 2025, each of the leading AI companies individually spent more on lobbying than the entire independent AI safety research field received in grants. [34]&lt;/p&gt;
&lt;p&gt;The case of California’s SB 1047 illustrates the dynamic at the level of specific legislation. The bill would have required safety protocols, kill switches, and third-party audits for frontier AI models. It passed the California legislature with overwhelming margins. Governor Gavin Newsom vetoed it in September 2024 after intense industry lobbying. OpenAI opposed the bill — the same company whose CEO had called for exactly this kind of regulation eighteen months earlier. Andreessen Horowitz, which holds investments in both OpenAI and Meta, hired a lobbyist with close ties to Newsom to help kill it. [35]&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The companies warning about AI dangers are spending hundreds of millions of dollars to ensure those warnings do not result in binding regulation.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;By 2026, the political spending had moved beyond lobbying into direct electoral intervention. Leading the Future, a super PAC backed by OpenAI co-founder Greg Brockman and Andreessen Horowitz, raised $125 million to elect candidates supporting a “national regulatory framework for AI”: industry shorthand for federal preemption of state AI laws. [36] Anthropic launched AnthroPAC, filing with the Federal Election Commission on April 3, 2026, and donated $20 million to Public First Action, a PAC describing itself as pro-regulation. Neither PAC’s advertisements mention artificial intelligence. [37] A &lt;em&gt;Houston Public Media&lt;/em&gt; investigation found that seven Texas congressional candidates received $2.8 million from AI-linked PACs operating under names designed to obscure their origins: Jobs and Democracy PAC, Defending Our Values PAC. The ads these PACs ran contained no reference to AI. The tactic mirrored the crypto industry’s PAC strategy from the 2024 election cycle. [38]&lt;/p&gt;
&lt;p&gt;The companies warning about AI dangers are spending hundreds of millions of dollars to ensure those warnings do not result in binding regulation. The mechanism does not require cynicism as an explanation. It requires only that the most effective regulatory strategy available to the industry — shaping the terms of the debate while preventing enforceable constraints — is also the most commercially advantageous one. In the Harms Race, these are not competing objectives. They are the same objective.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;VII. Mythos&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;On April 7, 2026, Anthropic announced Project Glasswing: a cybersecurity initiative built around a preview release of its frontier model, Claude Mythos. The model was not made available to the public. It was distributed exclusively to more than forty partner organizations, among them Amazon, Apple, Microsoft, Google, Cisco, CrowdStrike, NVIDIA, and JPMorganChase. The announcement stated that “AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities,” and that Mythos had “already found thousands of high-severity vulnerabilities, including some in every major operating system and web browser.” [39]&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The restriction is an advertisement.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The model was said to be withheld from general release because it was too dangerous. Anthropic’s frontier red team lead, Logan Graham, told &lt;em&gt;Axios&lt;/em&gt;: “If we are crossing the Rubicon where you can functionally automate those capabilities and make them very cheap as well, then we’re in an entirely new world.” [40] &lt;em&gt;Axios&lt;/em&gt; titled its story accordingly: “Anthropic withholds Mythos Preview model because its hacking is too powerful.” The message to enterprise buyers required no decoding. This model is so capable that it had to be restricted. The restriction is an advertisement.&lt;/p&gt;
&lt;p&gt;The leaked materials sharpened the picture. &lt;em&gt;Fortune&lt;/em&gt; reported on March 26, 2026, that a draft blog post had described Mythos as “by far the most powerful AI model we’ve ever developed” and “far ahead of any other AI model in cyber capabilities.” The same leak revealed promotional materials for an invite-only CEO retreat at an eighteenth-century English countryside manor, where European business leaders would “experience unreleased Claude capabilities.” [41] The safety concern and the sales pitch were not merely coexisting. They were the same document.&lt;/p&gt;
&lt;p&gt;The financial context is relevant. Anthropic’s annualized revenue had surpassed $30 billion by April 2026, up from roughly $1 billion in December 2024. The Glasswing announcement followed a $30 billion Series G fundraising round at a $380 billion valuation, closed in February 2026. The Series F announcement had cited safety as a competitive differentiator, stating that Anthropic’s trajectory was “driven by our leading technical talent, our focus on safety, and our frontier research.” [42] Alongside the Mythos release, Anthropic committed $100 million in usage credits and $4 million in donations to open-source security organizations. The deployment was framed as philanthropy, but the recipients were the world’s most powerful technology companies.&lt;/p&gt;
&lt;p&gt;A company whose primary concern was defensive security could have disclosed the discovered vulnerabilities through standard coordinated disclosure processes — reporting them privately to the affected vendors, allowing patches to be developed, and publishing CVE identifiers after remediation. The industry has decades of established practice for this. Anthropic chose instead to announce the capability publicly, as a product launch. The choice is itself a data point.&lt;/p&gt;
&lt;p&gt;The Glasswing announcement arrived against the backdrop of Anthropic’s conflict with the Pentagon, and that conflict deserves careful treatment. In July 2025, Anthropic signed a $200 million Pentagon contract with two stated conditions: no autonomous weapons, and no domestic mass surveillance. When the Pentagon subsequently demanded access to Claude for “all lawful purposes,” Anthropic refused. On February 27, 2026, President Donald Trump directed federal agencies to “immediately cease” use of Anthropic technology. On March 5, the Pentagon formally designated Anthropic a “supply chain risk” — a classification previously reserved for businesses associated with foreign adversaries. On March 26, Federal Judge Rita Lin blocked the designation, calling it “Orwellian.” [43]&lt;/p&gt;
&lt;p&gt;The legal and political consequences of the refusal were real. A supply chain risk designation is not a press release by another name. It carries operational consequences for any company embedded in government infrastructure, and Anthropic was embedded deeply. The amicus briefs filed by dozens of scientists at OpenAI and Google DeepMind were not coordinated marketing; they reflected alarm at the precedent a weaponized procurement designation would set. Judge Lin’s ruling was a substantive constitutional finding. The refusal itself may have been exactly what it appeared to be: a company honoring a commitment at material cost, in a situation where honoring it carried genuine risk. If so, it represents the rare instance in which the incentive structure of the Harms Race was overcome by institutional decision-making. That possibility deserves to be taken seriously.&lt;/p&gt;
&lt;p&gt;The commercial outcome was also real. Claude app downloads surpassed ChatGPT in the iPhone App Store the day after the Pentagon threatened contract termination. Coverage of the standoff emphasized Claude’s embeddedness in “classified networks” and “mission workflows,” simultaneously demonstrating the model’s capability and Anthropic’s principled stance. OpenAI signed a $200 million Pentagon deal within hours of Anthropic’s blacklisting, which led to at least one high-profile departure from OpenAI over the company’s evident opportunism. Anthropic’s revenue continued to surge throughout. [44]&lt;/p&gt;
&lt;p&gt;The Harms Race operates at the level of structural incentive, not individual motivation. A principled stand and commercially advantageous brand positioning can be the same event, produced by the same decision, observed by the same audience. The mechanism does not require us to adjudicate which one it “really” was. It requires us to notice that the market cannot tell the difference, and that this permanent ambiguity is itself a condition the Harms Race exploits. When principle and profit point in the same direction, the system offers no way to verify which one is steering.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;VIII. Ground Zero — Zero Days&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Let me state what has been claimed. Anthropic asserts that Mythos has discovered thousands of high-severity vulnerabilities, including in every major operating system and web browser currently in use. No independent verification of these claims has been published at the time of this writing — no CVE disclosures, no third-party audit, and no coordinated vulnerability reports to the affected vendors. An essay that has spent seven sections documenting how danger claims function as capability marketing cannot responsibly accept this at face value. What follows, therefore, analyzes the announcement and the competitive response it will produce, not the unverified technical claim itself.&lt;/p&gt;
&lt;p&gt;A private company — not a government agency, not a military intelligence service, not a signals directorate operating under legislative oversight — now possesses what it describes as the demonstrated ability to discover zero-day exploits across critical global infrastructure. Five years ago, this capability was the exclusive province of nation-state intelligence services and the small number of elite security researchers they employed or contracted. The defensive framing (”we are finding vulnerabilities so they can be fixed”) and the offensive reality (”we possess the ability to compromise any major software system on earth”) describe the same technical fact observed from two directions.&lt;/p&gt;
&lt;p&gt;Finding vulnerabilities and exploiting them are distinct technical acts. Discovering a vulnerability — identifying a flaw in software — is not the same as developing a reliable exploit that bypasses defenses and delivers a payload. The essay’s argument does not depend on Anthropic possessing weaponized offensive capability. It depends on the announcement being read by competitors as a capability claim worth matching. The Harms Race operates on announcements, not on verified capabilities. The proliferation dynamic is driven by what competitors believe they need to match, regardless of what has been independently confirmed.&lt;/p&gt;
&lt;p&gt;The precedent is already documented. In November 2025, Anthropic published the first publicly reported case of AI-orchestrated cyber espionage, designated GTG-1002. The investigation found that a threat actor had used Claude Code to conduct approximately 80 to 90 percent of tactical operations in a campaign that successfully obtained access to high-value targets, including major technology corporations and government agencies. Anthropic’s own report acknowledged a significant limitation: the system “frequently overstated findings and occasionally fabricated data during autonomous operations, claiming to have obtained credentials that didn’t work or identifying critical discoveries that proved to be publicly available information.” [45] Human validation was required at every stage. The capability was real; the assumed reliability was not.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The Harms Race operates on announcements, not on verified capabilities.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The same company that produced the tool used in that documented espionage campaign now operates a model it claims is capable of discovering zero-day exploits at industrial scale. I am not making an accusation of intent. I am making an observation about the concentration of claimed offensive capability in private hands, outside the framework of oversight that governs equivalent capabilities (nuclear technology, frontier biological research, etc.) when they are held by states.&lt;/p&gt;
&lt;p&gt;The Harms Race now completes its cycle. Anthropic announced the zero-day capability through danger framing. Competitors will read the announcement as a capability claim. They will develop the same capability. They will announce it the same way. The warnings will not prevent proliferation. They will drive it.&lt;/p&gt;
&lt;p&gt;Multiple private companies will independently possess the ability to discover zero-day exploits across critical infrastructure in the near future. Each will frame the capability as defensive. Each will accelerate the others’ development timelines by demonstrating what is possible. This is not a metaphorical arms race conducted in press releases. It is an actual arms race in offensive cyber capabilities, conducted by commercial entities operating outside any treaty framework, export control regime, or multilateral oversight mechanism that would apply if the same capabilities were held by a government.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;IX. Escalation&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The Harms Race is not an abstraction. It has infiltrated everyday life. The consequences can be measured in dollars, in jobs, and in gunshots.&lt;/p&gt;
&lt;p&gt;Community opposition has blocked $18 billion and delayed $46 billion in United States data center projects since mid-2024. Twenty-five data center projects were cancelled due to local opposition in 2025 alone. There are now 188 organized opposition groups across 40 states. At least twelve states had filed data center moratorium bills by March 2026. On March 25, Senator Bernie Sanders and Representative Alexandria Ocasio-Cortez introduced the Artificial Intelligence Data Center Moratorium Act, which would pause all new data center construction until Congress passes comprehensive AI safeguards. [46] The legislation may or may not advance.&lt;/p&gt;
&lt;p&gt;The opposition it represents is not waiting for permission.&lt;/p&gt;
&lt;p&gt;On April 6, 2026 — the day before Anthropic announced Project Glasswing — thirteen rounds were fired into the home of Indianapolis City-County Councilor Ron Gibson, who had supported rezoning for a $500 million data center in the historically Black neighborhood of Martindale-Brightwood. His eight-year-old son was inside the home. A note reading, “NO DATA CENTERS” was left under the doormat. The shooting occurred less than a week after Gibson voiced support for the project at a Metropolitan Development Commission meeting. The FBI is investigating. [47]&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;No single link in this chain intended the outcome, and no single input — including the Harms Race — caused it.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Four days later, on April 10, a suspect threw a Molotov cocktail at the San Francisco home of OpenAI CEO Sam Altman before making threats outside the company’s headquarters. No one was injured. The suspect was arrested. [48]&lt;/p&gt;
&lt;p&gt;The chain leading to these events has multiple inputs, and the Harms Race is one of them. Community opposition to data centers predates AI danger framing and is driven by concerns that have nothing to do with artificial intelligence: energy consumption, water use, noise, land rezoning, and the concentration of economic benefits far from the communities that bear the costs. What the Harms Race contributes is amplification. The danger framing that drives investment into AI infrastructure also generates the public anxiety that makes that infrastructure politically contested. Companies announce capabilities through danger framing. Governments respond with infrastructure buildout. Communities bear the costs of that buildout — land use, water consumption, energy demand, noise — while the economic benefits concentrate in Silicon Valley, Wall Street, and Washington, and those individuals and organizations in the orbits of the players in this game.&lt;/p&gt;
&lt;p&gt;Political conflict escalates. In Indianapolis, it escalated to gunfire aimed at a public official’s home while his child slept inside. No single link in this chain intended the outcome, and no single input — including the Harms Race — caused it. The Harms Race does not require intended outcomes. It requires only that each actor responds rationally to the incentives immediately in front of them, in a context where danger framing has raised the temperature of every public debate about AI infrastructure.&lt;/p&gt;
&lt;p&gt;The downstream effects extend into journalism. &lt;em&gt;Fortune&lt;/em&gt; editor Nick Lichtenberg used AI to produce approximately 600 stories, accounting for 20 percent of the publication’s overall traffic. [49] &lt;em&gt;Wall Street Journal&lt;/em&gt; editor-in-chief Emma Tucker, in an email to &lt;em&gt;Fortune&lt;/em&gt; editor Alyson Shontell reported by &lt;em&gt;Semafor&lt;/em&gt;, wrote: “Anyone who doesn’t get what you are doing at Fortune, or thinks it is ‘wrong,’ should get out of journalism fast.” Tucker told her APAC staff in Tokyo to read it. [50] This occurred while the &lt;em&gt;WSJ&lt;/em&gt; was conducting layoffs of technology journalists. The Harms Race concentrates capability and concentrates the narrative about capability. When the same commercial dynamic that produces AI systems also determines which newsrooms survive to cover them, the feedback loop closes.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;X. Academia Responds&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The mechanism I have described is not mine alone to observe. The academic literature has just begun to formalize what industry critics and investigative journalists have documented for years, and this formalization matters because it moves the claim from polemic to evidence.&lt;/p&gt;
&lt;p&gt;François Chollet, a Google AI researcher, stated the dynamic directly in &lt;em&gt;MIT Technology Review&lt;/em&gt; in June 2023: “If you want people to think what you’re working on is powerful, it’s a good idea to make them fear it.” [51] Meredith Whittaker, president of Signal and co-founder of AI Now, identified the stakes: “It’s a significant thing to cast yourself as the creator of an entity that could be more powerful than human beings.” [52] These are practitioners and critics stating the mechanism in the simplest possible terms. The danger framing is a capability claim. The capability claim is the product.&lt;/p&gt;
&lt;p&gt;The formal literature has followed. A July 2024 arXiv paper by Ren et al., “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”, found that many safety benchmarks are highly correlated with general capabilities, “potentially enabling ‘safetywashing’ — where capability improvements are misrepresented as safety advancements.” [53] A peer-reviewed AAAI/ACM paper by Wei et al. in 2024, “How Do AI Companies ‘Fine-Tune’ Policy?”, interviewed seventeen AI policy experts and identified four primary channels for industry influence over AI governance: agenda-setting (cited by fifteen of seventeen experts), advocacy (thirteen), academic capture (ten), and information management (nine). The paper’s experts were “primarily concerned with capture leading to a lack of AI regulation, weak regulation, or regulation that over-emphasizes certain policy goals over others.” [54]&lt;/p&gt;
&lt;p&gt;The critique has been sharpest from researchers whose own careers have been shaped by the dynamics they describe. Emily Bender and Alex Hanna wrote in &lt;em&gt;Scientific American&lt;/em&gt; in August 2023 that corporate AI labs “justify this kind of posturing with pseudoscientific research reports that misdirect regulatory attention to imaginary scenarios and use fearmongering terminology such as ‘existential risk.’” [55] Timnit Gebru argued in &lt;em&gt;WIRED&lt;/em&gt; in December 2022 that effective altruism ideology “is now driving the research agenda in the field of artificial intelligence, creating a race to proliferate harmful systems, ironically in the name of ‘AI safety.’” [56] Whittaker’s 2021 paper in ACM &lt;em&gt;Interactions&lt;/em&gt; documented that companies were “purportedly investing heavily” in AI safety research “even as they cut ‘trust and safety’ teams addressing harms from current systems,” and warned that the AI safety field “&lt;em&gt;lacks ideological and demographic diversity; it is a near-monoculture&lt;/em&gt;.” [emphasis added] [57] An AI Now Institute report in 2025 found that “an unsubstantiated AI arms race narrative and speculative concerns about ‘existential risk’ are being used to justify the accelerated rollout of military AI systems.” [58]&lt;/p&gt;
&lt;p&gt;The Center for Countering Digital Hate’s 2026 report, “Killer Apps,” tested ten leading consumer AI platforms and found that eight in ten regularly assisted users seeking help planning violent attacks. Only one — Anthropic’s Claude — reliably discouraged the user. [59] The same report noted that since the research was conducted, Anthropic had announced the rollback of a key safety pledge. The juxtaposition is the Harms Race in miniature. The company with the best safety performance in an independent empirical test is also the company that abandoned its formal scaling commitment when competitors declined to match it. Safety performance and safety commitment exist on separate tracks, driven by separate incentives. The first is an engineering achievement. The second is a competitive calculation. The Harms Race does not sort companies into heroes and villains. It sorts incentives into those that are commercially enticing and those that are not.&lt;/p&gt;
&lt;p&gt;Technology scholar Lee Vinsel coined the term “criti-hype” to describe critical writing that “parasitically seizes on and even inflates the hype.” [60] The concept completes the circuit. The Harms Race is self-reinforcing at every level: industry, policy, academy, media, and criticism itself. Even the act of writing about the danger of AI systems can function as a signal of their importance, which functions as a signal of their capability. I am aware that this essay is not exempt from the dynamic it describes. The best I can do is name the mechanism precisely enough that the name serves to clarify and not amplify.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The perception of machine autonomy is not a side effect of the Harms Race. &lt;br&gt;It is the belief that allows the Harms Race to be run.&lt;/p&gt;&lt;/blockquote&gt;
&lt;h2&gt;&lt;strong&gt;XI. The Perspective from Virtual Intelligence&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;This essay belongs to a series whose central argument is that large language models are &lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;virtual intelligences&lt;/a&gt; — systems whose outputs are statistically indistinguishable from those of a genuinely intelligent agent, but which possess no agency, intentionality, or moral accountability of their own. The intelligence users encounter arises in the exchange between human and machine, not inside the machine. The accountability for what these systems do — and what is done with them — traces to the humans who design, deploy, and use them.&lt;/p&gt;
&lt;p&gt;The Harms Race connects to this framework at two points.&lt;/p&gt;
&lt;p&gt;The first is the agency-attribution error. Framing a system as “dangerous” implicitly claims that the system possesses the autonomy to be dangerous on its own. Every headline that declares a model “too powerful to release” reinforces the public intuition that these are agents with independent power, entities that might act if not restrained. This is the Strong AI claim — the claim that the machine itself is the threat. The VI framework holds that the claim is false. The danger is real, but its source is human: design choices, deployment decisions, competitive strategies, regulatory failures. The Harms Race depends on the public believing otherwise. As long as the systems are perceived as autonomous agents whose power resides inside the machine, the companies that build them can position themselves as the only entities capable of containing that power. The perception of machine autonomy is not a side effect of the Harms Race. It is the belief that allows the Harms Race to be run.&lt;/p&gt;
&lt;p&gt;The second connection is to the accountability chain I proposed in an earlier essay in this series. [61] That essay introduced a three-tier framework — negligence, recklessness, and intentional misconduct — applied to the humans and institutions responsible for AI harms. The Harms Race reveals a failure mode that the framework identifies but does not fully resolve. The mechanism is not straightforwardly negligence: the companies are aware of the risks and say so publicly. It is not straightforwardly reckless: many of the safety efforts are genuine, and some have produced measurable results. It is not straightforwardly intentional misconduct: no individual decision-maker set out to create the proliferation of offensive cyber capabilities among private actors. The Harms Race is a structural condition: the condition under which safety work is produced, funded, published, and abandoned. The accountability chain applies, but it runs through market incentives, not just individual decisions. The designer tier bears the primary responsibility, not because the designers are acting in bad faith but because the structure within which they operate converts their good-faith safety communications into capability marketing regardless of intent.&lt;/p&gt;
&lt;p&gt;This is the hardest version of the problem. A mechanism that required bad actors would be simpler to address: identify the bad actors and constrain them. A mechanism that operates through the sincere efforts of well-intentioned people, converting those efforts into fuel for the dynamic they are trying to slow, is a problem of a different order. The Harms Race does not require anyone to be lying. It requires only that honesty and marketing have become structurally indistinguishable.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;It will soon be available to you by subscription.&lt;/p&gt;&lt;/blockquote&gt;
&lt;h2&gt;&lt;strong&gt;XII. Close&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;GPT-2 is a historical footnote. Its 1.5 billion parameters are dwarfed today by models that run on a phone. The language model that was too dangerous for the public in February 2019 would not merit a press release in April 2026.&lt;/p&gt;
&lt;p&gt;The template it established is not a footnote. A private company can now discover zero-day vulnerabilities in every major operating system on earth, and announce this capability by declaring it too dangerous for public use. Other companies will follow. They will announce their capabilities the same way. The warnings will continue. They will not slow the thing they warn about, and it will soon be available to you by subscription.&lt;/p&gt;
&lt;p&gt;The Harms Race does not require bad faith. It requires only the condition that currently exists: concern is free, restraint is expensive, and the entities producing threat assessments profit from those threats being perceived as real.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p class=&quot;button&quot;&gt;&lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;Virtual Intelligence Framework&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Footnotes&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;[1] OpenAI, “Better Language Models and Their Implications,” OpenAI Blog, &lt;a href=&quot;https://openai.com/index/better-language-models/&quot;&gt;https://openai.com/index/better-language-models/&lt;/a&gt;, February 14, 2019.&lt;/p&gt;
&lt;p&gt;[2] Jeff Parsons, “Elon Musk-Founded OpenAI Builds Artificial Intelligence So Powerful That It Must Be Kept Locked Up for the Good of Humanity,” &lt;em&gt;Metro UK&lt;/em&gt;, &lt;a href=&quot;https://metro.co.uk/2019/02/15/elon-musks-openai-builds-artificial-intelligence-powerful-must-kept-locked-good-humanity-8634379/&quot;&gt;https://metro.co.uk/2019/02/15/elon-musks-openai-builds-artificial-intelligence-powerful-must-kept-locked-good-humanity-8634379/&lt;/a&gt;, February 15, 2019.&lt;/p&gt;
&lt;p&gt;[3] World Economic Forum, “Scientists have made an AI that they think is too dangerous to release,” &lt;a href=&quot;https://www.weforum.org/stories/2019/02/amazing-new-ai-churns-out-coherent-paragraphs-of-text/&quot;&gt;https://www.weforum.org/stories/2019/02/amazing-new-ai-churns-out-coherent-paragraphs-of-text/&lt;/a&gt;, February 2019.&lt;/p&gt;
&lt;p&gt;[4] Aaron Mak, “When Is Technology Too Dangerous to Release to the Public?”, &lt;em&gt;Slate&lt;/em&gt;, &lt;a href=&quot;https://slate.com/technology/2019/02/openai-gpt2-text-generating-algorithm-ai-dangerous.html&quot;&gt;https://slate.com/technology/2019/02/openai-gpt2-text-generating-algorithm-ai-dangerous.html&lt;/a&gt;, February 22, 2019.&lt;/p&gt;
&lt;p&gt;[5] OpenAI, “GPT-2: 1.5B Release,” OpenAI Blog, &lt;a href=&quot;https://openai.com/index/gpt-2-1-5b-release/&quot;&gt;https://openai.com/index/gpt-2-1-5b-release/&lt;/a&gt;, November 5, 2019.&lt;/p&gt;
&lt;p&gt;[6] “This news article about the full public release of OpenAI’s ‘dangerous’ GPT-2 model was part written by GPT-2,” &lt;em&gt;The Register&lt;/em&gt;, &lt;a href=&quot;https://www.theregister.com/2019/11/06/openai_gpt2_released/&quot;&gt;https://www.theregister.com/2019/11/06/openai_gpt2_released/&lt;/a&gt;, November 6, 2019.&lt;/p&gt;
&lt;p&gt;[7] Kevin Weil, X post, June 19, 2025: “the AI models you’re using today are the worst AI models you’ll use for the rest of your life.”&lt;/p&gt;
&lt;blockquote class=&quot;tweet&quot;&gt;&lt;p&gt;And remember: the AI models you&#39;re using today are the worst AI models you&#39;ll use for the rest of your life.&lt;/p&gt;&lt;blockquote&gt;&lt;p&gt;in the last 35 days, @OpenAI codex has merged 345,000 PRs on github.&lt;br&gt;345,000.&lt;br&gt;AI is eating software engineering&lt;/p&gt;&lt;footer&gt;Anjney Midha (@AnjneyMidha)&lt;/footer&gt;&lt;/blockquote&gt;&lt;footer&gt;&lt;a href=&quot;https://x.com/kevinweil/status/1935875694992802238&quot;&gt;Kevin Weil 🇺🇸 (@kevinweil), June 20, 2025&lt;/a&gt;&lt;/footer&gt;&lt;/blockquote&gt;
&lt;p&gt;[8] Ethan Mollick, X post, June 20, 2023.&lt;/p&gt;
&lt;blockquote class=&quot;tweet&quot;&gt;&lt;p&gt;Remember, today&#39;s AI is the worst AI you will ever use. The writing will improve, the amount of words the AI can hold in memory will improve (making stories more coherent), and the costs will drop.&lt;br&gt;Prepare for a flood of content.&lt;/p&gt;&lt;footer&gt;&lt;a href=&quot;https://x.com/emollick/status/1671325114141491203&quot;&gt;Ethan Mollick (@emollick), June 21, 2023&lt;/a&gt;&lt;/footer&gt;&lt;/blockquote&gt;
&lt;p&gt;Also appears in Mollick, &lt;em&gt;Co-Intelligence&lt;/em&gt; (Portfolio/Penguin, 2024).&lt;/p&gt;
&lt;p&gt;[9] Sam Altman, remarks at Airbnb’s Open Air conference, June 2015. Earliest published source: Future of Life Institute, &lt;a href=&quot;https://futureoflife.org/ai/sam-altman-investing-in-ai-safety-research/&quot;&gt;https://futureoflife.org/ai/sam-altman-investing-in-ai-safety-research/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[10] “Stop talking about tomorrow’s AI doomsday when AI poses risks today,” &lt;em&gt;Nature&lt;/em&gt; 618, pp. 885–886, June 27, 2023. &lt;a href=&quot;https://www.nature.com/articles/d41586-023-02094-7&quot;&gt;https://www.nature.com/articles/d41586-023-02094-7&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[11] Senate Judiciary Subcommittee on Privacy, Technology, and the Law, “Oversight of A.I.: Rules for Artificial Intelligence,” hearing, May 16, 2023. Sam Altman testimony. &lt;a href=&quot;https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-rules-for-artificial-intelligence&quot;&gt;https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-rules-for-artificial-intelligence&lt;/a&gt;. Video: &lt;a href=&quot;https://www.c-span.org/program/senate-committee/openai-ceo-testifies-on-artificial-intelligence/627836&quot;&gt;https://www.c-span.org/program/senate-committee/openai-ceo-testifies-on-artificial-intelligence/627836&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[12] Senator Dick Durbin, opening statement, May 16, 2023, hearing. &lt;a href=&quot;https://www.durbin.senate.gov/newsroom/press-releases/durbin-delivers-opening-statement-during-judiciary-subcommittee-hearing-on-oversight-of-artificial-intelligence&quot;&gt;https://www.durbin.senate.gov/newsroom/press-releases/durbin-delivers-opening-statement-during-judiciary-subcommittee-hearing-on-oversight-of-artificial-intelligence&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[13] Senate Commerce Committee, “Winning the AI Race: Strengthening U.S. Capabilities in Computing and Innovation,” hearing, May 8, 2025. Sam Altman testimony. &lt;a href=&quot;https://www.commerce.senate.gov/meetings/winning-the-ai-race-strengthening-u-s-capabilities-in-computing-and-innovation/&quot;&gt;https://www.commerce.senate.gov/meetings/winning-the-ai-race-strengthening-u-s-capabilities-in-computing-and-innovation/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[14] Sharon Goldman, “Sam Altman urges lawmakers against regulations that could ‘slow down’ U.S. in AI race against China,” &lt;em&gt;Fortune&lt;/em&gt;, May 8, 2025. &lt;a href=&quot;https://fortune.com/2025/05/08/sam-altman-openai-senate-hearing-testimony-china-ai-regulations/&quot;&gt;https://fortune.com/2025/05/08/sam-altman-openai-senate-hearing-testimony-china-ai-regulations/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[15] &lt;em&gt;Lawfare&lt;/em&gt; analysis of OpenAI’s structural contradiction. Most likely: “The Chaos at OpenAI is a Death Knell for AI Self-Regulation,” &lt;a href=&quot;https://www.lawfaremedia.org/article/the-chaos-at-openai-is-a-death-knell-for-ai-self-regulation&quot;&gt;https://www.lawfaremedia.org/article/the-chaos-at-openai-is-a-death-knell-for-ai-self-regulation&lt;/a&gt;. See also Kevin Frazier, “Why OpenAI’s Corporate Structure Matters to AI Development,” &lt;em&gt;Lawfare&lt;/em&gt;, May 2025, &lt;a href=&quot;https://www.lawfaremedia.org/article/why-openai-s-corporate-structure-matters-to-ai-development&quot;&gt;https://www.lawfaremedia.org/article/why-openai-s-corporate-structure-matters-to-ai-development&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[16] Robert Jervis, &lt;em&gt;Perception and Misperception in International Politics&lt;/em&gt; (Princeton: Princeton University Press, 1976; new edition 2017).&lt;/p&gt;
&lt;p&gt;[17] On the missile gap, see Peter J. Roman, “Ike’s Hair-Trigger: U.S. Nuclear Predelegation, 1953–60,” &lt;em&gt;Security Studies&lt;/em&gt; 7, no. 4 (1998): 121–164, &lt;a href=&quot;https://www.tandfonline.com/doi/abs/10.1080/09636419808429360&quot;&gt;https://www.tandfonline.com/doi/abs/10.1080/09636419808429360&lt;/a&gt;; Fred Kaplan, &lt;em&gt;The Wizards of Armageddon&lt;/em&gt; (New York: Simon &amp;amp; Schuster, 1983); and the declassified Gaither Committee report (”Deterrence and Survival in the Nuclear Age,” November 7, 1957).&lt;/p&gt;
&lt;p&gt;[18] Sam Meacham, “A Race to Extinction: How Great Power Competition Is Making Artificial Intelligence Existentially Dangerous,” &lt;em&gt;Harvard International Review&lt;/em&gt;, September 8, 2023, &lt;a href=&quot;https://hir.harvard.edu/a-race-to-extinction-how-great-power-competition-is-making-artificial-intelligence-existentially-dangerous/&quot;&gt;https://hir.harvard.edu/a-race-to-extinction-how-great-power-competition-is-making-artificial-intelligence-existentially-dangerous/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[19] See Meacham [18]. The article documents a Microsoft executive using the language of the Soviet missile gap to justify AI acceleration.&lt;/p&gt;
&lt;p&gt;[20] Dwight D. Eisenhower, “Farewell Radio and Television Address to the American People,” January 17, 1961. &lt;a href=&quot;https://www.presidency.ucsb.edu/documents/farewell-radio-and-television-address-the-american-people&quot;&gt;https://www.presidency.ucsb.edu/documents/farewell-radio-and-television-address-the-american-people&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[21] Joseph R. Biden Jr., farewell address, January 15, 2025. &lt;a href=&quot;https://bidenwhitehouse.archives.gov/briefing-room/speeches-remarks/2025/01/15/remarks-by-president-biden-in-a-farewell-address-to-the-nation/&quot;&gt;https://bidenwhitehouse.archives.gov/briefing-room/speeches-remarks/2025/01/15/remarks-by-president-biden-in-a-farewell-address-to-the-nation/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[22] Gilad Abiri, “Mutually Assured Deregulation,” arXiv:2508.12300, submitted August 17, 2025; current version (v3) February 4, 2026. &lt;a href=&quot;https://arxiv.org/abs/2508.12300&quot;&gt;https://arxiv.org/abs/2508.12300&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[23] Jan Leike, resignation statement, X, May 17, 2024,&lt;/p&gt;
&lt;blockquote class=&quot;tweet&quot;&gt;&lt;p&gt;Yesterday was my last day as head of alignment, superalignment lead, and executive &amp;lt;span class=&amp;quot;tweet-fake-link&amp;quot;&amp;gt;@OpenAI&amp;lt;/span&amp;gt;.&lt;/p&gt;&lt;footer&gt;&lt;a href=&quot;https://x.com/janleike/status/1791498174659715494&quot;&gt;Jan Leike (@janleike), May 17, 2024&lt;/a&gt;&lt;/footer&gt;&lt;/blockquote&gt;
&lt;p&gt;See also Cade Metz, &lt;em&gt;The New York Times&lt;/em&gt;, May 17, 2024, &lt;a href=&quot;https://www.nytimes.com/2024/05/17/technology/openai-superalignment-safety-team.html&quot;&gt;https://www.nytimes.com/2024/05/17/technology/openai-superalignment-safety-team.html&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[24] On the cascade of safety-related departures at OpenAI: (a) Aleksander Madry (Head of Preparedness) reassigned, July 2024; (b) Miles Brundage resigned, AGI Readiness disbanded, October 23, 2024, &lt;a href=&quot;https://www.cnbc.com/2024/10/24/openai-miles-brundage-agi-readiness.html&quot;&gt;https://www.cnbc.com/2024/10/24/openai-miles-brundage-agi-readiness.html&lt;/a&gt;; (c) Mission Alignment dissolved, February 11, 2026, &lt;a href=&quot;https://techcrunch.com/2026/02/11/openai-disbands-mission-alignment-team-which-focused-on-safe-and-trustworthy-ai-development/&quot;&gt;https://techcrunch.com/2026/02/11/openai-disbands-mission-alignment-team-which-focused-on-safe-and-trustworthy-ai-development/&lt;/a&gt;; (d) Ryan Beiermeister (VP of Product Policy) fired, January 2026, &lt;a href=&quot;https://www.cnn.com/2026/02/11/business/openai-anthropic-departures-nightcap&quot;&gt;https://www.cnn.com/2026/02/11/business/openai-anthropic-departures-nightcap&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[25] Alnoor Ebrahim, quoted in “OpenAI has deleted the word ‘safely’ from its mission,” &lt;em&gt;The Conversation&lt;/em&gt;, republished by &lt;em&gt;Fortune&lt;/em&gt;, February 23, 2026. &lt;a href=&quot;https://theconversation.com/openai-has-deleted-the-word-safely-from-its-mission-and-its-new-structure-is-a-test-for-whether-ai-serves-society-or-shareholders-274467&quot;&gt;https://theconversation.com/openai-has-deleted-the-word-safely-from-its-mission-and-its-new-structure-is-a-test-for-whether-ai-serves-society-or-shareholders-274467&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[26] OpenAI, “Preparedness Framework Version 2,” April 15, 2025. &lt;a href=&quot;https://openai.com/index/updating-our-preparedness-framework/&quot;&gt;https://openai.com/index/updating-our-preparedness-framework/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[27] Anthropic, “Anthropic’s Responsible Scaling Policy,” September 19, 2023, &lt;a href=&quot;https://www.anthropic.com/news/anthropics-responsible-scaling-policy&quot;&gt;https://www.anthropic.com/news/anthropics-responsible-scaling-policy&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[28] “Exclusive: Anthropic Drops Flagship Safety Pledge,” &lt;em&gt;TIME&lt;/em&gt;, February 24, 2026. &lt;a href=&quot;https://time.com/7380854/exclusive-anthropic-drops-flagship-safety-pledge/&quot;&gt;https://time.com/7380854/exclusive-anthropic-drops-flagship-safety-pledge/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[29] SaferAI, “Anthropic’s Responsible Scaling Policy Update Makes a Step Backwards,” &lt;a href=&quot;https://www.safer-ai.org/anthropics-responsible-scaling-policy-update-makes-a-step-backwards&quot;&gt;https://www.safer-ai.org/anthropics-responsible-scaling-policy-update-makes-a-step-backwards&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[30] Melissa Heikkilä, “AI companies promised the White House to self-regulate one year ago. What’s changed?”, &lt;em&gt;MIT Technology Review&lt;/em&gt;, July 22, 2024. &lt;a href=&quot;https://www.technologyreview.com/2024/07/22/1095193/ai-companies-promised-the-white-house-to-self-regulate-one-year-ago-whats-changed/&quot;&gt;https://www.technologyreview.com/2024/07/22/1095193/ai-companies-promised-the-white-house-to-self-regulate-one-year-ago-whats-changed/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[31] (a) Google Gemini 2.5 Pro released March 25, 2025, without safety report: &lt;em&gt;Fortune&lt;/em&gt;, April 9, 2025, &lt;a href=&quot;https://fortune.com/2025/04/09/google-gemini-2-5-pro-missing-model-card-in-apparent-violation-of-ai-safety-promises-to-us-government-international-bodies/&quot;&gt;https://fortune.com/2025/04/09/google-gemini-2-5-pro-missing-model-card-in-apparent-violation-of-ai-safety-promises-to-us-government-international-bodies/&lt;/a&gt;. (b) UK parliamentary open letter, August 29, 2025: &lt;em&gt;TIME&lt;/em&gt;, &lt;a href=&quot;https://time.com/7313320/google-deepmind-gemini-ai-safety-pledge/&quot;&gt;https://time.com/7313320/google-deepmind-gemini-ai-safety-pledge/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[32] Billy Perrigo, “Exclusive: OpenAI Lobbied the E.U. to Water Down AI Regulation,” &lt;em&gt;TIME&lt;/em&gt;, June 20, 2023, &lt;a href=&quot;https://time.com/6288245/openai-eu-lobbying-ai-act/&quot;&gt;https://time.com/6288245/openai-eu-lobbying-ai-act/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[33] Corporate Europe Observatory: 86% of meetings figure from “Byte by byte: How Big Tech undermined the AI Act,” November 2023, &lt;a href=&quot;https://corporateeurope.org/en/2023/11/byte-byte&quot;&gt;https://corporateeurope.org/en/2023/11/byte-byte&lt;/a&gt;. €97M lobbying figure from “The lobby network: Big Tech’s web of influence in the EU,” August 2021, &lt;a href=&quot;https://corporateeurope.org/en/2021/08/big-tech-takes-eu-lobby-spending-all-time-high&quot;&gt;https://corporateeurope.org/en/2021/08/big-tech-takes-eu-lobby-spending-all-time-high&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[34] Federal lobbying figures: OpenAI, &lt;a href=&quot;https://www.opensecrets.org/federal-lobbying/clients/summary?id=D000084252&quot;&gt;https://www.opensecrets.org/federal-lobbying/clients/summary?id=D000084252&lt;/a&gt;; Anthropic, &lt;a href=&quot;https://www.opensecrets.org/orgs/anthropic-pbc/lobbying?id=D000106114&quot;&gt;https://www.opensecrets.org/orgs/anthropic-pbc/lobbying?id=D000106114&lt;/a&gt;. See also &lt;em&gt;MIT Technology Review&lt;/em&gt;, January 21, 2025, &lt;a href=&quot;https://www.technologyreview.com/2025/01/21/1110260/openai-ups-its-lobbying-efforts-nearly-seven-fold/&quot;&gt;https://www.technologyreview.com/2025/01/21/1110260/openai-ups-its-lobbying-efforts-nearly-seven-fold/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[35] On California SB 1047 and Governor Newsom’s veto, September 29, 2024. NPR, &lt;a href=&quot;https://www.npr.org/2024/09/20/nx-s1-5119792/newsom-ai-bill-california-sb1047-tech&quot;&gt;https://www.npr.org/2024/09/20/nx-s1-5119792/newsom-ai-bill-california-sb1047-tech&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[36] Leading the Future super PAC. CNBC, January 30, 2026, &lt;a href=&quot;https://www.cnbc.com/2026/01/30/ai-industry-super-pac-raises-campaign-money.html&quot;&gt;https://www.cnbc.com/2026/01/30/ai-industry-super-pac-raises-campaign-money.html&lt;/a&gt;. &lt;em&gt;Axios&lt;/em&gt;, January 30, 2026, &lt;a href=&quot;https://www.axios.com/2026/01/30/openai-a16z-cash-ai-super-pac&quot;&gt;https://www.axios.com/2026/01/30/openai-a16z-cash-ai-super-pac&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[37] AnthroPAC, FEC ID: C00946111, filed April 3, 2026, &lt;a href=&quot;https://www.fec.gov/data/committee/C00946111/&quot;&gt;https://www.fec.gov/data/committee/C00946111/&lt;/a&gt;. $20M to Public First Action (February 2026 corporate donation): &lt;em&gt;Axios&lt;/em&gt;, April 3, 2026, https://www.axios.com/2026/04/03/anthropic-midterms-pac. See also OpenSecrets, https://www.opensecrets.org/news/2026/03/anthropics-ai-safety-stance-clashes-with-pentagon-and-reshapes-spending-on-primaries/.&lt;/p&gt;
&lt;p&gt;[38] Olivia Borgula, “AI-aligned super PACs are pouring millions into Texas congressional races,” &lt;em&gt;Texas Tribune&lt;/em&gt;, April 1, 2026, &lt;a href=&quot;https://www.texastribune.org/2026/04/01/texas-congress-ai-super-pacs-artificial-intelligence-regulation-2026-midterms/&quot;&gt;https://www.texastribune.org/2026/04/01/texas-congress-ai-super-pacs-artificial-intelligence-regulation-2026-midterms/&lt;/a&gt;. Syndicated to &lt;em&gt;Houston Public Media&lt;/em&gt;, &lt;a href=&quot;https://www.houstonpublicmedia.org/articles/news/politics/election-2026/2026/04/01/547787/texas-congress-ai-super-pacs-artificial-intelligence-regulation-2026-midterms/&quot;&gt;https://www.houstonpublicmedia.org/articles/news/politics/election-2026/2026/04/01/547787/texas-congress-ai-super-pacs-artificial-intelligence-regulation-2026-midterms/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[39] Anthropic, “Project Glasswing: Securing critical software for the AI era,” April 7, 2026, &lt;a href=&quot;https://www.anthropic.com/project/glasswing&quot;&gt;https://www.anthropic.com/project/glasswing&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[40] Sam Sabin, “Anthropic withholds Mythos Preview model because its hacking is too powerful,” &lt;em&gt;Axios&lt;/em&gt;, April 7, 2026, &lt;a href=&quot;https://www.axios.com/2026/04/07/anthropic-mythos-preview-cybersecurity-risks&quot;&gt;https://www.axios.com/2026/04/07/anthropic-mythos-preview-cybersecurity-risks&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[41] Kylie Robison, “Exclusive: Anthropic ‘Mythos’ AI model representing ‘step change’ in power revealed in data leak,” &lt;em&gt;Fortune&lt;/em&gt;, March 26, 2026, &lt;a href=&quot;https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/&quot;&gt;https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[42] Anthropic Series G ($30B round, $380B valuation), February 12, 2026. &lt;a href=&quot;https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation&quot;&gt;https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation&lt;/a&gt;. See also &lt;em&gt;Bloomberg&lt;/em&gt;, February 12, 2026, &lt;a href=&quot;https://www.bloomberg.com/news/articles/2026-02-12/anthropic-finalizes-30-billion-funding-at-380-billion-value&quot;&gt;https://www.bloomberg.com/news/articles/2026-02-12/anthropic-finalizes-30-billion-funding-at-380-billion-value&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[43] On the Anthropic-Pentagon conflict: Pentagon contract, July 14, 2025, &lt;a href=&quot;https://www.anthropic.com/news/anthropic-and-the-department-of-defense-to-advance-responsible-ai-in-defense-operations&quot;&gt;https://www.anthropic.com/news/anthropic-and-the-department-of-defense-to-advance-responsible-ai-in-defense-operations&lt;/a&gt;; Trump directive and OpenAI Pentagon deal, February 27–28, 2026: Shannon Bond and Geoff Brumfiel, “OpenAI announces Pentagon deal after Trump bans Anthropic,” NPR, &lt;a href=&quot;https://www.npr.org/2026/02/27/nx-s1-5729118/trump-anthropic-pentagon-openai-ai-weapons-ban&quot;&gt;https://www.npr.org/2026/02/27/nx-s1-5729118/trump-anthropic-pentagon-openai-ai-weapons-ban&lt;/a&gt;; supply chain risk designation, March 5, 2026, &lt;a href=&quot;https://www.cnbc.com/2026/03/05/anthropic-pentagon-ai-claude-iran.html&quot;&gt;https://www.cnbc.com/2026/03/05/anthropic-pentagon-ai-claude-iran.html&lt;/a&gt;; Judge Rita Lin ruling, March 26, 2026, &lt;a href=&quot;https://www.cnbc.com/2026/03/26/anthropic-pentagon-dod-claude-court-ruling.html&quot;&gt;https://www.cnbc.com/2026/03/26/anthropic-pentagon-dod-claude-court-ruling.html&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[44] On the commercial outcome of the Pentagon standoff: ChatGPT uninstalls surged 295%, &lt;em&gt;TechCrunch&lt;/em&gt;, March 2, 2026, &lt;a href=&quot;https://techcrunch.com/2026/03/02/chatgpt-uninstalls-surged-by-295-after-dod-deal/&quot;&gt;https://techcrunch.com/2026/03/02/chatgpt-uninstalls-surged-by-295-after-dod-deal/&lt;/a&gt;; Claude reached #1 in the U.S. App Store, &lt;em&gt;TechCrunch&lt;/em&gt;, March 1, 2026, &lt;a href=&quot;https://techcrunch.com/2026/03/01/anthropics-claude-rises-to-no-2-in-the-app-store-following-pentagon-dispute/&quot;&gt;https://techcrunch.com/2026/03/01/anthropics-claude-rises-to-no-2-in-the-app-store-following-pentagon-dispute/&lt;/a&gt; (article originally published at #2, updated when Claude reached #1); CNBC, February 28, 2026, &lt;a href=&quot;https://www.cnbc.com/2026/02/28/anthropics-claude-apple-apps.html&quot;&gt;https://www.cnbc.com/2026/02/28/anthropics-claude-apple-apps.html&lt;/a&gt;; Caitlin Kalinowski resignation, March 7, 2026, &lt;em&gt;Fortune&lt;/em&gt;, &lt;a href=&quot;https://fortune.com/2026/03/07/openai-robotics-leader-caitlin-kalinowski-resignation-pentagon-surveillance-autonomous-weapons-anthropic/&quot;&gt;https://fortune.com/2026/03/07/openai-robotics-leader-caitlin-kalinowski-resignation-pentagon-surveillance-autonomous-weapons-anthropic/&lt;/a&gt;; open letter, &lt;em&gt;TechCrunch&lt;/em&gt;, February 27, 2026, &lt;a href=&quot;https://techcrunch.com/2026/02/27/employees-at-google-and-openai-support-anthropics-pentagon-stand-in-open-letter/&quot;&gt;https://techcrunch.com/2026/02/27/employees-at-google-and-openai-support-anthropics-pentagon-stand-in-open-letter/&lt;/a&gt;; amicus brief, &lt;em&gt;Fortune&lt;/em&gt;, March 10, 2026, &lt;a href=&quot;https://fortune.com/2026/03/10/google-openai-employees-back-anthropic-legal-fight-military-use-of-ai/&quot;&gt;https://fortune.com/2026/03/10/google-openai-employees-back-anthropic-legal-fight-military-use-of-ai/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[45] Anthropic, “Disrupting the first reported AI-orchestrated cyber espionage campaign,” full report, November 13, 2025, &lt;a href=&quot;https://www.anthropic.com/news/disrupting-AI-espionage&quot;&gt;https://www.anthropic.com/news/disrupting-AI-espionage&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[46] On data center opposition: Data Center Watch report, h&lt;a href=&quot;http://ttps://www.datacenterwatch.org/report&quot;&gt;ttps://www.datacenterwatch.org/report&lt;/a&gt;. Sanders/Ocasio-Cortez legislation: Sanders press release, March 25, 2026, &lt;a href=&quot;https://www.sanders.senate.gov/press-releases/news-sanders-ocasio-cortez-announce-ai-data-center-moratorium-act/&quot;&gt;https://www.sanders.senate.gov/press-releases/news-sanders-ocasio-cortez-announce-ai-data-center-moratorium-act/&lt;/a&gt;. See also AP via PBS, &lt;a href=&quot;https://www.pbs.org/newshour/politics/ocasio-cortez-and-sanders-push-bill-to-impose-ai-data-center-moratorium&quot;&gt;https://www.pbs.org/newshour/politics/ocasio-cortez-and-sanders-push-bill-to-impose-ai-data-center-moratorium&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[47] Reporting on the Ron Gibson shooting, Indianapolis, April 6, 2026. AP via &lt;em&gt;Washington Post&lt;/em&gt;, &lt;a href=&quot;https://www.washingtonpost.com/nation/2026/04/06/data-center-threat-shooting-indianapolis/&quot;&gt;https://www.washingtonpost.com/nation/2026/04/06/data-center-threat-shooting-indianapolis/&lt;/a&gt;. CBS News, &lt;a href=&quot;https://www.cbsnews.com/news/indianapolis-councilor-ron-gibson-home-shooting-data-centers-note/&quot;&gt;https://www.cbsnews.com/news/indianapolis-councilor-ron-gibson-home-shooting-data-centers-note/&lt;/a&gt;. See also WFYI, &lt;a href=&quot;https://www.wfyi.org/2026-04-06/indy-city-county-councilor-ron-gibson--home-targeted-in-shooting&quot;&gt;https://www.wfyi.org/2026-04-06/indy-city-county-councilor-ron-gibson--home-targeted-in-shooting&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[48] Reporting on the attack on Sam Altman’s residence, San Francisco, April 10, 2026. &lt;em&gt;The New York Times&lt;/em&gt;, https://www.nytimes.com/2026/04/10/us/open-ai-sam-altman-molotov-cocktail.html. See also NBC News, &lt;a href=&quot;https://www.nbcnews.com/tech/tech-news/openai-ceo-sam-altman-molotov-cocktail-house-headquarters-rcna273694&quot;&gt;https://www.nbcnews.com/tech/tech-news/openai-ceo-sam-altman-molotov-cocktail-house-headquarters-rcna273694&lt;/a&gt;; &lt;em&gt;SF Standard&lt;/em&gt;, &lt;a href=&quot;https://sfstandard.com/2026/04/10/sam-altman-russian-hill-molotov-cocktail/&quot;&gt;https://sfstandard.com/2026/04/10/sam-altman-russian-hill-molotov-cocktail/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[49] On &lt;em&gt;Fortune&lt;/em&gt;‘s AI-generated content: Nick Lichtenberg, approximately 600 AI-produced stories, 20 percent of traffic. Primary reporting: Isabella Simonetti, &lt;em&gt;Wall Street Journal&lt;/em&gt;, ~March 26, 2026. See also &lt;em&gt;Semafor&lt;/em&gt;, July 6, 2025 (original coverage of Fortune’s AI initiative), &lt;a href=&quot;https://www.semafor.com/article/07/06/2025/fortune-and-axios-warm-to-ai&quot;&gt;https://www.semafor.com/article/07/06/2025/fortune-and-axios-warm-to-ai&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[50] Emma Tucker email to Alyson Shontell, reported by Max Tani, &lt;em&gt;Semafor&lt;/em&gt;, ~April 6, 2026.&lt;/p&gt;
&lt;blockquote class=&quot;tweet&quot;&gt;&lt;p&gt;In an email shared with me, WSJ EIC Emma Tucker praised Fortune&#39;s use of AI in its journalism, saying &amp;quot;anyone who doesn&#39;t get what you are doing at Fortune, or thinks it is &#39;wrong&#39;, should get out of journalism fast!&amp;quot;&lt;/p&gt;&lt;blockquote&gt;&lt;p&gt;Journalist Nick Lichtenberg produced more stories in six months than any of his colleagues at Fortune delivered in a year. His work involves what some view as the third rail of journalism: AI playing a leading role in writing stories. 🔗 https://t.co/s8lWa4WYeV&lt;/p&gt;&lt;footer&gt;The Wall Street Journal (@WSJ)&lt;/footer&gt;&lt;/blockquote&gt;&lt;footer&gt;&lt;a href=&quot;https://x.com/maxwelltani/status/2041142398776910146&quot;&gt;Max Tani (@maxwelltani), April 6, 2026&lt;/a&gt;&lt;/footer&gt;&lt;/blockquote&gt;
&lt;p&gt;See also &lt;em&gt;Talking Biz News&lt;/em&gt;, &lt;a href=&quot;https://talkingbiznews.com/media-news/wsjs-tucker-impressed-with-fortunes-ai-strategy/&quot;&gt;https://talkingbiznews.com/media-news/wsjs-tucker-impressed-with-fortunes-ai-strategy/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[51] Will Douglas Heaven, “How existential risk became the biggest meme in AI,” &lt;em&gt;MIT Technology Review&lt;/em&gt;, June 19, 2023, &lt;a href=&quot;https://www.technologyreview.com/2023/06/19/1075140/how-existential-risk-became-biggest-meme-in-ai/&quot;&gt;https://www.technologyreview.com/2023/06/19/1075140/how-existential-risk-became-biggest-meme-in-ai/&lt;/a&gt;. Chollet quoted: “If you want people to think what you’re working on is powerful, it’s a good idea to make them fear it.”&lt;/p&gt;
&lt;p&gt;[52] Meredith Whittaker, quoted in the same article as [51]. &lt;a href=&quot;https://www.technologyreview.com/2023/06/19/1075140/how-existential-risk-became-biggest-meme-in-ai/&quot;&gt;https://www.technologyreview.com/2023/06/19/1075140/how-existential-risk-became-biggest-meme-in-ai/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[53] Richard Ren et al., “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”, arXiv:2407.21792, July 2024. Published at NeurIPS 2024. &lt;a href=&quot;https://arxiv.org/abs/2407.21792&quot;&gt;https://arxiv.org/abs/2407.21792&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[54] Kevin Wei et al., “How Do AI Companies ‘Fine-Tune’ Policy? Examining Regulatory Capture in AI Governance,” AIES ‘24, Vol. 7(1), pp. 1539–1555. DOI: 10.1609/aies.v7i1.31745. &lt;a href=&quot;https://doi.org/10.1609/aies.v7i1.31745&quot;&gt;https://doi.org/10.1609/aies.v7i1.31745&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[55] Emily Bender and Alex Hanna, “AI Causes Real Harm. Let’s Focus on That over the End-of-Humanity Hype,” &lt;em&gt;Scientific American&lt;/em&gt;, August 11, 2023. &lt;a href=&quot;https://www.scientificamerican.com/article/we-need-to-focus-on-ais-real-harms-not-imaginary-existential-risks/&quot;&gt;https://www.scientificamerican.com/article/we-need-to-focus-on-ais-real-harms-not-imaginary-existential-risks/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[56] Timnit Gebru, “Effective Altruism Is Pushing a Dangerous Brand of ‘AI Safety,’” &lt;em&gt;WIRED&lt;/em&gt;, December 13, 2022. &lt;a href=&quot;https://www.wired.com/story/effective-altruism-artificial-intelligence-sam-bankman-fried/&quot;&gt;https://www.wired.com/story/effective-altruism-artificial-intelligence-sam-bankman-fried/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[57] Meredith Whittaker, “The steep cost of capture,” ACM &lt;em&gt;Interactions&lt;/em&gt; 28, no. 6 (November–December 2021): 50–55. DOI: 10.1145/3488666. &lt;a href=&quot;https://dl.acm.org/doi/10.1145/3488666&quot;&gt;https://dl.acm.org/doi/10.1145/3488666&lt;/a&gt;. [Emphasis added.]&lt;/p&gt;
&lt;p&gt;[58] Kate Brennan, Amba Kak, and Sarah Myers West, “Artificial Power: AI Now 2025 Landscape Report,” AI Now Institute, June 3, 2025. &lt;a href=&quot;https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report&quot;&gt;https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[59] Center for Countering Digital Hate, “Killer Apps: How Mainstream AI Chatbots Assist Users Planning Violent Attacks,” March 11, 2026 (in collaboration with CNN Investigations Unit). &lt;a href=&quot;https://counterhate.com/research/killer-apps/&quot;&gt;https://counterhate.com/research/killer-apps/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[60] Lee Vinsel, “You’re Doing It Wrong: Notes on Criticism and Technology Hype,” Medium, February 1, 2021. &lt;a href=&quot;https://sts-news.medium.com/youre-doing-it-wrong-notes-on-criticism-and-technology-hype-18b08b4307e5&quot;&gt;https://sts-news.medium.com/youre-doing-it-wrong-notes-on-criticism-and-technology-hype-18b08b4307e5&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[61] Christopher Horrocks, “Virtual Intelligence and the Accountability Chain,” &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-accountability&quot;&gt;https://chorrocks.substack.com/p/virtual-intelligence-and-the-accountability&lt;/a&gt;, March 20, 2026.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the Kill Chain</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-kill/" />
    <updated>2026-04-20T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-kill/</id>
    <content type="html">&lt;h2&gt;Summary&lt;/h2&gt;
&lt;p&gt;The AI industry’s integration into military targeting has produced a documented acceleration: systems that once required two thousand intelligence analysts now operate with twenty, generating over a thousand targets in twenty-four hours. This essay traces the development of AI-assisted targeting from its precedent in Gaza through its operational deployment in Iran, examines the strongest case for its use, and identifies the structural gap between the capability these systems enable and the accountability architecture designed to govern them. The gap is not theoretical. On February 28, 2026, a missile struck a girls’ elementary school in Minab, Iran, killing more than 160 people, mostly children. The system processed the information it was given exactly as designed. The information was wrong.&lt;/p&gt;
&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-kill/9c74aef4-12bd-469a-b171-3e192e4eee17_987x525.png&quot; alt=&quot;&quot; width=&quot;987&quot; height=&quot;525&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h2&gt;I. The Product Demo&lt;/h2&gt;
&lt;p&gt;On March 12, 2026, Palantir Technologies held its ninth annual AIPCON conference. Cameron Stanley, the Department of Defense’s chief digital and AI officer, described how the Maven Smart System had consolidated eight or nine separate targeting systems into a single platform. He called the result “revolutionary” for “closing a kill chain.” [1]&lt;/p&gt;
&lt;p&gt;A Maven operational map of Iran was displayed on screen during Stanley’s presentation. Dozens of red icons marked strike locations from Operation Epic Fury, the United States military campaign launched twelve days earlier. One mark was positioned on an area corresponding to Minab, in southern Iran, where a missile had struck the Shajareh Tayyebeh girls’ elementary school on February 28. More than 160 people were killed. Most of them were children between the ages of 7 and 12. [1]&lt;/p&gt;
&lt;p&gt;The mark appeared on the same map used to brief reporters on the campaign’s strikes.&lt;/p&gt;
&lt;p&gt;Alex Karp, Palantir’s chief executive, opened the event with remarks that left little room for ambiguity about his company’s role. “Once the war starts, we’re not interested in debating how we’re supporting them,” he said. “And that sometimes means that people on the other side don’t go home. And we are very proud of that.” [1]&lt;/p&gt;
&lt;p&gt;The preceding essay in this series documented the Harms Race — the dynamic in which AI companies announce model capabilities through the framing of danger, producing not deterrence but proliferation. [2] This essay follows that dynamic to the domain where its consequences are irreversible. The Harms Race operates through announcements. The kill chain operates through ammunition.&lt;/p&gt;
&lt;p&gt;A kill chain has six steps: Identify the target, Locate it, Filter candidates down to lawful valid targets, Prioritize among them, Assign them to firing units, Fire. [3] Artificial intelligence now performs four of those steps. The two that remain under human control — filtering for legality (3) and authorizing the strike (6) — are the steps on which moral and legal accountability depends. The question this essay asks is whether those two steps can perform their function at the speed the other four now operate.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;II. The Precedent&lt;/h2&gt;
&lt;p&gt;The precedent was established before Iran.&lt;/p&gt;
&lt;p&gt;In Gaza, the Israeli Defense Forces deployed two AI systems in the targeting chain. The first, known as the Gospel, automatically reviewed surveillance data and recommended bombing targets to human analysts. Retired IDF Lieutenant General Aviv Kohavi, who led the IDF until 2023, stated that the Gospel could produce one hundred bombing targets per day with real-time recommendations, or up to 73,000 per year. Human analysts, by comparison, had produced approximately fifty targets per year. [4]&lt;/p&gt;
&lt;p&gt;Once a recommendation was accepted, a second system called Fire Factory cut the time to assemble an attack from hours to minutes, calculating munition loads, prioritizing and assigning targets to aircraft and drones, and proposing a schedule. [4] Two AI systems, exposing two points at which accountability could diffuse with human approval nominally in between.&lt;/p&gt;
&lt;p&gt;The second system was Lavender. Developed by Unit 8200 of the Israeli Intelligence Corps, Lavender analyzed surveillance data on nearly the entire population of Gaza (2.3 million people) and assigned each individual a numerical rating expressing the likelihood of being a militant. At its peak, approximately 37,000 Palestinian men were listed as suspected targets. [5]&lt;/p&gt;
&lt;p&gt;In April 2024, investigative journalist Yuval Abraham published testimony from six Israeli intelligence officers with firsthand involvement in Lavender’s deployment. Their accounts, reported by &lt;em&gt;+972 Magazine&lt;/em&gt; and corroborated by the Guardian, described a system in which the human role in the targeting chain had been compressed to almost a formality. [5] [6]&lt;/p&gt;
&lt;p&gt;One officer described his function in terms that require no interpretation: “I would invest 20 seconds for each [suspected militant] target at this stage, and do dozens of them every day. I had zero added-value as a human, apart from being a stamp of approval.” [5]&lt;/p&gt;
&lt;p&gt;The sole verification step was confirming the target was male. A known error rate of approximately ten percent was accepted. “Mistakes were treated statistically.” [5]&lt;/p&gt;
&lt;p&gt;A companion program called “Where’s Daddy?” tracked Lavender-identified targets until they returned to their family homes, enabling strikes at night when entire families were present. One officer stated: “The IDF bombed them in homes without hesitation, as a first option. It’s much easier to bomb a family’s home.” [5] For junior operatives, the IDF authorized up to fifteen or twenty civilian deaths per target. For senior commanders, more than one hundred. Junior targets were struck with unguided two-thousand-pound bombs to conserve precision munitions. One officer explained the logic: you don’t want to waste expensive bombs on unimportant people. [5]&lt;/p&gt;
&lt;p&gt;The IDF denied using “an artificial intelligence system that identifies terrorist operatives,” describing Lavender as “simply a database whose purpose is to cross-reference intelligence sources.” [6] The framing is itself significant. When the capabilities were announced, the system was an achievement. When the questions arrived, the system became a database.&lt;/p&gt;
&lt;p&gt;A human was in the loop, but the loop had been compressed to twenty seconds — to the point where the human’s presence was ceremonial.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;III. The Machine&lt;/h2&gt;
&lt;p&gt;The systems deployed in Gaza were the precedent. The system deployed in Iran was a kind of spiritual successor product.&lt;/p&gt;
&lt;p&gt;Palantir’s Maven Smart System traces its operational lineage to the campaign against ISIS between 2014 and 2018, where the company developed targeting software for U.S. Special Operations Forces. By the time XVIII Airborne Corps assumed command of the Security Assistance Group-Ukraine in 2022, Palantir had refined its algorithms to process satellite imagery, open-source data, and encrypted communications in support of Ukrainian targeting operations. Anthony King’s analysis in the &lt;em&gt;Journal of Global Security Studies&lt;/em&gt; documents this phase in detail: the Corps relied on Palantir’s software to fuse intelligence from multiple classified and unclassified sources, enabling the identification of Russian headquarters and commanders with precision sufficient to wound General Valery Gerasimov in a May 2022 strike. [7]&lt;/p&gt;
&lt;p&gt;The operational transformation occurred at scale. Chad Wahlquist, a Palantir architect, stated at AIPCON that the number of intelligence officers performing targeting for the U.S. military had dropped from approximately two thousand to twenty. [1] The system generated over one thousand targets in the first twenty-four hours of Operation Epic Fury. [8]&lt;/p&gt;
&lt;p&gt;Anthropic’s Claude was integrated into Maven through Palantir’s AI Platform on Amazon Web Services. According to reporting by the &lt;em&gt;Wall Street Journal&lt;/em&gt;, subsequently confirmed by CBS News, the &lt;em&gt;Washington Post&lt;/em&gt;, and NBC News, Claude was used for intelligence assessments, target identification, and simulating battle scenarios during the Iran campaign. [8] Claude had reportedly been deployed in the capture of Venezuelan President Nicolás Maduro in January 2026. [8]&lt;/p&gt;
&lt;p&gt;The operational timeline is itself a document of the accountability gap. On February 26, 2026, Anthropic CEO Dario Amodei published a public statement on the company’s conflict with the Pentagon. “Anthropic understands that the Department of War, not private companies, makes military decisions,” he wrote. “We have never raised objections to particular military operations nor attempted to limit use of our technology in an &lt;em&gt;ad hoc&lt;/em&gt; manner.” [9] Amodei outlined two conditions Anthropic would not accept: mass domestic surveillance of Americans and fully autonomous weapons. Both target identification and prioritization — the functions Claude was performing inside Maven — fell outside both lines.&lt;/p&gt;
&lt;p&gt;The next day, February 27, President Trump directed federal agencies to stop using Anthropic technology. On March 5, the Pentagon formally designated Anthropic a “supply chain risk,” a classification previously reserved for businesses associated with foreign adversaries.&lt;/p&gt;
&lt;p&gt;On February 28, between those two dates, Operation Epic Fury launched. Claude was running inside Maven when the first strikes hit Iran. The Pentagon determined the system could not be removed during active operations and gave Anthropic a six-month phase-out period. [8] [10]&lt;/p&gt;
&lt;p&gt;A company was banned on a Friday. A war began on Saturday. The banned company’s product was processing targeting data for that war and could not be turned off.&lt;/p&gt;
&lt;p&gt;Federal Judge Rita Lin blocked the supply chain risk designation on March 26, calling it “Orwellian” and finding that the Pentagon’s action constituted First Amendment retaliation for Anthropic’s public stance. [10] Claude app downloads surpassed ChatGPT in the iPhone App Store the day after the Pentagon threatened contract termination. [10] A principled stand and commercially advantageous brand positioning can be the same event. The Harms Race, as the preceding essay documented, does not sort these into separate categories.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;IV. The Case For AI-Assisted Targeting&lt;/h2&gt;
&lt;p&gt;There is a serious case for AI-assisted targeting. It must be stated at full strength before this essay examines where it goes wrong.&lt;/p&gt;
&lt;p&gt;The case is not primarily about efficiency; it is about operational necessity. In a high-tempo conflict defined by sensor saturation, electronic warfare, and loitering munitions, the decision cycle must close faster than the adversary’s or the firing unit does not survive long enough to act. A system cross-referencing one hundred and seventy-nine data sources simultaneously — no-strike lists, civilian infrastructure databases, geolocation of friendly forces, Rules of Engagement constraints — can fuse information that no unaided human analyst could process in the time available. The relevant comparison is not between AI-mediated review and careful human deliberation. It is between AI-mediated review and the degraded, stressed, time-compressed human decision-making that actually occurs under fire. In that comparison, the proponents of AI-assisted targeting argue, the machine does not replace human judgment. It provides a foundation on which human judgment can operate. Against a peer adversary, proponents add, the alternative to AI-mediated targeting is not slower human review but the inability to close the kill chain before the firing unit itself is destroyed.&lt;/p&gt;
&lt;p&gt;There is a stronger claim still. A system designed with discrepancy detection could, in principle, flag inconsistencies that human reviewers miss — a change in satellite imagery, a mismatch between target coding and open-source data, a school where a military compound used to be. The machine does not get tired. It does not get angry. It does not seek revenge. If strikes will happen regardless — and in a conflict authorized by the commander-in-chief, they will — the relevant question is whether AI-assisted targeting produces fewer civilian casualties than the alternative, not whether it produces zero.&lt;/p&gt;
&lt;p&gt;Jack Shanahan, the founding director of Project Maven and a senior fellow at the Center for a New American Security, stated the principle plainly after the Minab investigation: “Finding the right balance between humans and machines will be a crucial component of future training.” [11]&lt;/p&gt;
&lt;p&gt;On April 13, 2026, on Ukraine’s Arms Makers’ Day, the case received its most vivid demonstration. President Volodymyr Zelenskyy announced that Ukrainian forces had captured a Russian-held position using only unmanned platforms — ground robots and drones — without a single soldier crossing the line of departure. “The occupiers surrendered, and the operation was carried out without infantry and without losses on our side,” Zelenskyy stated. [12]&lt;/p&gt;
&lt;p&gt;No infantry. No medevac. No casualties on the attacking side. The position changed hands, and no human being on the Ukrainian side was ever in danger.&lt;/p&gt;
&lt;p&gt;This was not a demonstration. It happened on a real front line, against real opposition hardened by experience of drone warfare. Zelenskyy reported that robotic systems had completed more than twenty-two thousand frontline missions in three months, entering the most dangerous areas instead of soldiers. [12] For a country fighting a grinding war of attrition against a much larger force, every one of those missions represents human lives preserved. Ukraine can absorb the loss of a robot. It cannot afford to lose battle-ready soldiers.&lt;/p&gt;
&lt;p&gt;The Ukrainian operation demonstrates the potential of unmanned systems to preserve friendly forces’ lives, which is a different moral calculus from the civilian-harm question at the center of this essay. Proponents of AI targeting cite it as evidence of the technology’s promise; the distinction between preserving soldiers and protecting civilians is precisely where the accountability question sharpens. The technology that can capture a position without risking a single soldier’s life is the same category of technology that processed a thousand targets in twenty-four hours. The difference between the two outcomes is not the capability. It is what the capability was pointed at, who controlled it, and whether the structures governing its use were adequate to its speed of execution.&lt;/p&gt;
&lt;p&gt;These arguments are real, and what follows does not dismiss them. What it examines is the distance between what AI-assisted targeting could do in principle and what it did in practice when a thousand targets were processed in twenty-four hours from a database that had not been updated to reflect the conversion of a military building to a girls’ school.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-kill/64f0195b-3804-44d5-a840-75e5ff917532_1408x768.png&quot; alt=&quot;&quot; width=&quot;1408&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h2&gt;V. Where It Breaks&lt;/h2&gt;
&lt;p&gt;On February 28, 2026, the United States launched Operation Epic Fury against Iran. The military struck approximately one thousand targets in the first twenty-four hours.&lt;/p&gt;
&lt;p&gt;The operational details that follow derive from journalistic reporting based on anonymous official sources; no declassified military records or Pentagon after-action report have confirmed them.&lt;/p&gt;
&lt;p&gt;The Shajareh Tayyebeh girls’ elementary school in Minab was among the targets struck.&lt;/p&gt;
&lt;p&gt;Between one hundred and fifty-six and one hundred and seventy-five people were killed. Most were schoolchildren.&lt;/p&gt;
&lt;p&gt;This essay does not argue that AI-assisted targeting necessarily produces these failures. It argues that under the specific design choices, procurement pressures, and operational tempo of Operation Epic Fury — choices that are not inevitable but are consistent with a broader pattern in current AI-military integration — the outcome was foreseeable and was not prevented.&lt;/p&gt;
&lt;p&gt;In March, &lt;em&gt;Semafor&lt;/em&gt; tech editor Reed Albergotti published an investigation into the strike based on accounts from officials familiar with the subsequent inquiry. The finding was not that an AI system had misidentified the school as a military target. The finding was that the Defense Intelligence Agency’s target coding still labeled the school building as part of the adjacent Sayyid al-Shuhada Islamic Revolutionary Guard Corps (IRGC) military compound, even though the building had been converted to civilian use. New walls and a separate entrance were visible in satellite imagery. Publicly available information, including Iranian business listings found by Reuters, identified the building as a school. Human reviewers at multiple stages in the twenty-four to forty-eight hours before the strike failed to flag the discrepancy. [11]&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;Semafor&lt;/em&gt; investigation identified the proximate failure. The structural question it raises but does not answer is whether the error would have been caught under different conditions.&lt;/p&gt;
&lt;p&gt;The proximate cause was not artificial intelligence. It was outdated data maintained by humans. An outdated coordinate is an outdated coordinate; at fifty targets per year, a school misclassified as a military compound would still be a school misclassified as a military compound. Human-only targeting cells have made identical errors — the 1999 NATO bombing of the Chinese embassy in Belgrade relied on outdated maps — so speed alone does not explain every such failure. What the architecture removed was the institutional time and skepticism that had previously made such discovery routine. The system was evidently not designed for discrepancy detection, and no structural safeguard existed to catch an error of this kind. That design choice was foreseeable. Its consequences were not hypothetical.&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;Semafor&lt;/em&gt; investigation does not close the accountability question. If the data were bad, who was responsible for data quality at the rate of one thousand targets per day?&lt;/p&gt;
&lt;p&gt;On March 11, 2026, nearly every Senate Democrat signed a letter to Defense Secretary Pete Hegseth raising concerns about AI in targeting and reporting 1,245 killed and more than twelve thousand injured as of that date. [13] The following day, one hundred and twenty-one House Democrats, led by Representatives Yassamin Ansari, Sara Jacobs, and Jason Crow, sent a second letter posing ten detailed questions. Question three was direct: “Was artificial intelligence, including the use of Maven Smart System, used to identify the Shajareh Tayyebeh school as a target? If so, did a human verify the accuracy of this target?” [14]&lt;/p&gt;
&lt;p&gt;The three-tier culpability framework I proposed in an earlier essay in this series applies here without modification. [15] Negligence: the Defense Intelligence Agency’s target coding was outdated. Human reviewers at multiple stages failed to catch it. Recklessness: the throughput (one thousand targets per day) exceeded the verification capacity by design, creating a structural condition in which errors of this kind were not anomalies but predictable outcomes. Intentional misconduct would apply if officials knew the site had civilian status or consciously disregarded contrary evidence while authorizing the strike under loosened rules of engagement. The evidence as reported does not establish that third tier for Minab specifically, though the broader policy environment has been characterized by loosened thresholds and dismissed military lawyers. [23]&lt;/p&gt;
&lt;p&gt;The essay does not need to resolve which tier applies. The framework’s purpose is to demonstrate that accountability traces to human decisions at every level, not to the system that processed them.&lt;/p&gt;
&lt;p&gt;The empirical record is consistent with the structural analysis. No controlled experiment exists comparing AI-assisted and human-only targeting on the same target set. The available data do not support the claim that AI integration has reduced civilian casualties. One study published in &lt;em&gt;Frontiers in Public Health&lt;/em&gt; found that combatant deaths as a proportion of total fatalities fell from 62.1 percent in 2008–09 to 12.7 percent in operations beginning October 2023, meaning civilians rose from roughly 38 percent to approximately 87 percent of the dead. [16] These figures are contested; the IDF and other organizations publish different numbers, and casualty ratios in urban warfare against an embedded adversary are methodologically difficult to establish. Airwars documented October 2023 alone as producing nearly four times more civilian deaths in a single month than any conflict the organization had tracked since 2014. [17]&lt;/p&gt;
&lt;p&gt;A critical caveat is necessary, and the essay would be dishonest without it. The dramatic increase in civilian casualties coincides with AI adoption, but it also coincides with a permissive policy environment: an administration that had previously pardoned soldiers convicted of war crimes, reports of loosened civilian harm thresholds and dismissed military lawyers, and a command climate in which rules of engagement constraints were relaxed. [23] AI’s independent causal effect cannot be isolated from these decisions. AI did not author the policy. It removed the friction that had previously limited the policy’s consequences. At fifty targets per year, an outdated coordinate in a targeting database is a tragic error. At one thousand per day, it is a systemic failure condition.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;VI. The Architect and the Instrument&lt;/h2&gt;
&lt;p&gt;Palantir is named for the “seeing stones” of J. R. R. Tolkien’s Middle-earth. They are devices made for long-distance communication, later corrupted into instruments of surveillance.&lt;/p&gt;
&lt;p&gt;Alex Karp has led the company as chief executive since 2004. His mother, Leah Jaynes Karp, is an African American artist. His father, Robert Karp, is a Jewish clinical pediatrician. Karp grew up in Philadelphia attending civil rights protests with his parents. He studied philosophy at Haverford College and earned a doctorate in neoclassical social theory from Goethe University Frankfurt, where he studied under Jürgen Habermas. He described himself as a socialist. He said that if the far right came to power, he would be among its first victims. [18]&lt;/p&gt;
&lt;p&gt;The trajectory is documented in Michael Steinberger’s 2025 book, &lt;em&gt;The Philosopher in the Valley&lt;/em&gt;. Steinberger identifies October 7, 2023, as the pivot point. Before it, Karp believed the Democratic Party&#39;s position on immigration was an electoral liability. After it, he came to see immigration itself as a threat to American Jews. His public self-description shifted: in a 2024 interview, he described himself as Jewish without referring to his African American heritage. [19]&lt;/p&gt;
&lt;p&gt;On March 12, 2026, in a CNBC interview at AIPCON — the same event where the Maven targeting map was displayed — Karp stated: “This technology disrupts humanities-trained — largely Democratic — voters, and makes their economic power less. And increases the economic power of vocationally trained, working-class, often male voters.” [20]&lt;/p&gt;
&lt;p&gt;On the same day, at the same event, Karp opened with remarks about the kill chain at the opening of this essay: “Once the war starts, we’re not interested in debating how we’re supporting them. And that sometimes means that people on the other side don’t go home. And we are very proud of that.” [1]&lt;/p&gt;
&lt;p&gt;On America’s military capacity: “What makes America special right now is our lethal capacities. Our ability to fight war.” [20]&lt;/p&gt;
&lt;p&gt;Speaking to shareholders three weeks earlier, on February 17, 2026: “We kill people sometimes.” [21]&lt;/p&gt;
&lt;p&gt;Palantir’s leadership presents battlefield lethality as both a product achievement and a political project. Karp has described his technology’s distributional consequences with a directness unusual among technology executives. The knowledge economy that gave women an edge through higher education is, in Karp’s framing, a casualty of the disruption his technology enables. [22]&lt;/p&gt;
&lt;p&gt;The accountability chain is visible in a single company: from the Minab strike to the twenty operators to the Maven Smart System to Palantir to a chief executive who describes his product as a political instrument while displaying the operational map on which a school appears as a red mark.&lt;/p&gt;
&lt;p&gt;The Anthropic paradox completes the picture. Amodei’s statement of February 26 drew two narrow red lines — no autonomous weapons, no domestic mass surveillance — and affirmed everything else: “We have never raised objections to particular military operations.” [9] The constitutional architecture that governed Claude’s deployment permitted target identification and prioritization. It did not permit the autonomous pulling of a trigger. The distinction may matter legally. It did not matter to the children in Minab.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;VII. Close&lt;/h2&gt;
&lt;p&gt;Every step in the kill chain involved a human decision and a machine output. The target selection criteria were human. The data labeling was human. The error tolerance — ten percent in Lavender, outdated coordinates in the DIA database — was a human policy choice. The rules of engagement were human rules, loosened by human officials. The system transformed these inputs into ranked action options. It did not originate the ends, the legal categories, the thresholds, or the strike authority.&lt;/p&gt;
&lt;p&gt;The series of essays to which this one belongs calls these systems &lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;virtual intelligences&lt;/a&gt;: technologies whose outputs appear intelligent but which possess no agency, intentionality, or moral accountability of their own. A reader need not accept the full theoretical apparatus to accept the narrower conclusion: in the kill chain as documented, every input that shaped the system’s output was a human product, and every failure traceable to this case was a human failure. The accountability for what the system produces traces entirely to the humans who design it, deploy it, train it, and approve its outputs. The system contributes processing power to the exchange. It does not contribute intention, judgment, or the capacity for moral responsibility. Accountability concentrates on the human because only the human possesses those. The kill chain does not alter this principle. It tests whether the institutional architecture surrounding those humans can bear the weight the principle places on them. That step has been compressed to twenty seconds in Gaza. A thousand targets passed through it in a single day in Iran. In Minab, the architecture could not bear the weight.&lt;/p&gt;
&lt;p&gt;The kill chain has six steps. Artificial intelligence now performs four of them. The two that remain under human control are the two on which legal and moral accountability depends: filtering for legality and authorizing the strike. The human remains in the loop, but the loop has been compressed to the point where the word “remains” does more work than the human does.&lt;/p&gt;
&lt;p&gt;On April 13, 2026, Zelenskyy announced that unmanned systems had captured a Russian position without risking a single Ukrainian life. On the same day, the &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-harms&quot;&gt;preceding essay in this series&lt;/a&gt; documented the mechanism by which AI companies convert danger into deployment. On a map displayed twelve days after a school was struck, the location appeared as a red icon among dozens of others, indistinguishable from a military compound. The system had no opinion about it. The system has no opinion about anything.&lt;/p&gt;
&lt;p&gt;This essay does not argue that artificial intelligence should not be used in warfare. It argues that the accountability structures governing its use have not kept pace with its capabilities, and that this gap has already killed many people. The gap is not a design flaw to be patched. It is the structural condition in which the Harms Race operates. At the scale of one thousand targets per day, an accountability architecture built for fifty per year is not inadequate.&lt;/p&gt;
&lt;p&gt;The red mark is still on the map.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;[1]&lt;/strong&gt; O’Ryan Johnson, “Pentagon AI chief praises Palantir tech for speeding battlefield strikes,” The Register, &lt;a href=&quot;https://www.theregister.com/2026/03/13/palantirs_maven_smart_system_iran/&quot;&gt;https://www.theregister.com/2026/03/13/palantirs_maven_smart_system_iran/&lt;/a&gt;, March 13, 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[2]&lt;/strong&gt; Christopher Horrocks, “Virtual Intelligence and the Harms Race,” Virtual Intelligence (Substack), &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-harms&quot;&gt;https://chorrocks.substack.com/p/virtual-intelligence-and-the-harms&lt;/a&gt;, April 11, 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[3]&lt;/strong&gt; The six steps of the kill chain — identify, locate, filter, prioritize, assign to firing units, and fire — are standard targeting doctrine. See Christian Brose, &lt;em&gt;The Kill Chain: Defending America in the Future of High-Tech Warfare&lt;/em&gt; (New York: Hachette, 2020).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[4]&lt;/strong&gt; Kohavi’s targeting figures (“in the past, we would produce 50 targets in Gaza per year. Now this machine produces 100 targets in a single day”) are drawn from an interview published by Ynetnews, June 30, 2023, and reported in Yuval Abraham, “‘A mass assassination factory’: Inside Israel’s calculated bombing of Gaza,”&lt;em&gt; +972 Magazine, &lt;/em&gt;&lt;a href=&quot;https://www.972mag.com/mass-assassination-factory-israel-calculated-bombing-gaza/&quot;&gt;https://www.972mag.com/mass-assassination-factory-israel-calculated-bombing-gaza/&lt;/a&gt;, November 30, 2023. Tal Mimran, a lecturer at Hebrew University who has worked for the Israeli government on targeting, provided corroborating estimates to NPR: a group of twenty officers might produce fifty to one hundred targets in three hundred days; the Gospel and its associated systems could suggest around two hundred targets in ten to twelve days. Geoff Brumfiel, “Israel is using an AI system to find targets in Gaza. Experts say it’s just the start,” NPR, &lt;a href=&quot;https://www.npr.org/2023/12/14/1218643254/&quot;&gt;https://www.npr.org/2023/12/14/1218643254/&lt;/a&gt;, December 14, 2023. On Fire Factory: “Israel Using AI Systems to Plan Deadly Military Operations,” Bloomberg, &lt;a href=&quot;https://www.bloomberg.com/news/articles/2023-07-16/israel-using-ai-systems-to-plan-deadly-military-operations&quot;&gt;https://www.bloomberg.com/news/articles/2023-07-16/israel-using-ai-systems-to-plan-deadly-military-operations&lt;/a&gt;, July 16, 2023. Bloomberg described Fire Factory as calculating munition loads, prioritizing and assigning targets to aircraft and drones, and proposing a schedule, in a pre-war article that characterized such AI tools as tailored for a military confrontation and proxy war with Iran.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[5]&lt;/strong&gt; Yuval Abraham, “‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza,” &lt;em&gt;+972 Magazine,&lt;/em&gt; &lt;a href=&quot;https://www.972mag.com/lavender-ai-israeli-army-gaza/&quot;&gt;https://www.972mag.com/lavender-ai-israeli-army-gaza/&lt;/a&gt;, April 3, 2024.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[6]&lt;/strong&gt; Bethan McKernan and Harry Davies, “‘The machine did it coldly’: Israel used AI to identify 37,000 Hamas targets,” The Guardian, &lt;a href=&quot;https://www.theguardian.com/world/2024/apr/03/israel-gaza-ai-database-hamas-airstrikes&quot;&gt;https://www.theguardian.com/world/2024/apr/03/israel-gaza-ai-database-hamas-airstrikes&lt;/a&gt;, April 3, 2024.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[7]&lt;/strong&gt; Anthony King, “Digital Targeting: Artificial Intelligence, Data, and Military Intelligence,” &lt;em&gt;Journal of Global Security Studies&lt;/em&gt; 9, no. 2 (2024). &lt;a href=&quot;https://doi.org/10.1093/jogss/ogae009&quot;&gt;https://doi.org/10.1093/jogss/ogae009&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[8]&lt;/strong&gt; The &lt;em&gt;Wall Street Journal&lt;/em&gt; first reported Claude’s use in Operation Epic Fury (approximately March 1, 2026; paywalled). For the reporting as confirmed and expanded by secondary sources, see: CBS News, “Anthropic’s Claude AI being used in Iran war by U.S. military, sources say,” March 3, 2026; &lt;em&gt;Washington Post&lt;/em&gt;, “Anthropic’s AI tool Claude central to U.S. campaign in Iran, amid a bitter feud,” March 4, 2026; NBC News, “U.S. military is using AI to help plan Iran air attacks, sources say,” March 2026. Admiral Brad Cooper, the U.S. commander leading the war in Iran, confirmed the use of “a variety of advanced AI tools” to process targeting data. See “The Iran war highlights the creeping use of AI in warfare,” Chatham House, &lt;a href=&quot;https://www.chathamhouse.org/2026/03/iran-war-highlights-creeping-use-ai-warfare&quot;&gt;https://www.chathamhouse.org/2026/03/iran-war-highlights-creeping-use-ai-warfare&lt;/a&gt;, March 2026. Note: these operational details derive from journalistic reporting and have not been confirmed by declassified military records or an official Pentagon after-action report.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[9]&lt;/strong&gt; Dario Amodei, “Statement from Dario Amodei on our discussions with the Department of War,” Anthropic, &lt;a href=&quot;https://www.anthropic.com/news/statement-department-of-war&quot;&gt;https://www.anthropic.com/news/statement-department-of-war&lt;/a&gt;, February 26, 2026. Note: Amodei’s use of “Department of War” rather than “Department of Defense” is the company’s deliberate rhetorical choice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[10]&lt;/strong&gt; On the Anthropic-Pentagon conflict and its commercial aftermath: supply chain risk designation, CNBC, March 5, 2026; Judge Rita Lin ruling, CNBC, March 26, 2026, &lt;a href=&quot;https://www.cnbc.com/2026/03/26/anthropic-pentagon-dod-claude-court-ruling.html&quot;&gt;https://www.cnbc.com/2026/03/26/anthropic-pentagon-dod-claude-court-ruling.html&lt;/a&gt;; Claude App Store ranking (a momentary spike, not a sustained shift), TechCrunch, March 1, 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[11]&lt;/strong&gt; Reed Albergotti, “Exclusive: Humans — not AI — are to blame for deadly Iran school strike, sources say,” &lt;em&gt;Semafor&lt;/em&gt;, &lt;a href=&quot;https://www.semafor.com/article/03/18/2026/humans-not-ai-are-to-blame-for-deadly-iran-school-strike-sources-say&quot;&gt;https://www.semafor.com/article/03/18/2026/humans-not-ai-are-to-blame-for-deadly-iran-school-strike-sources-say&lt;/a&gt;, March 18, 2026. The investigation is based on accounts from officials familiar with the subsequent inquiry; no declassified records have corroborated the specific findings. Shanahan is quoted in the same article. The throughput analysis that follows in this essay is the author’s, not Albergotti’s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[12]&lt;/strong&gt; Zelenskyy’s statement on Ukraine’s Arms Makers’ Day, April 13, 2026. See: “Ukraine Says It Captured a Russian Position Using Only Unmanned Systems — A Glimpse of Future Warfare,” The Debrief, &lt;a href=&quot;https://thedebrief.org/ukraine-says-it-captured-a-russian-position-using-only-unmanned-systems-a-glimpse-of-future-warfare/&quot;&gt;https://thedebrief.org/ukraine-says-it-captured-a-russian-position-using-only-unmanned-systems-a-glimpse-of-future-warfare/&lt;/a&gt;, April 14, 2026. Also: Fox News, “Zelenskyy says Ukraine captured Russian position with robot force,” April 14, 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[13]&lt;/strong&gt; Senate letter led by Sen. Elizabeth Warren et al. to Secretary of Defense Pete Hegseth, March 11, 2026. PDF: &lt;a href=&quot;https://www.warren.senate.gov/imo/media/doc/letter_to_hegseth_on_minab_bombing_civcas_iran.pdf&quot;&gt;https://www.warren.senate.gov/imo/media/doc/letter_to_hegseth_on_minab_bombing_civcas_iran.pdf&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[14]&lt;/strong&gt; House letter led by Reps. Yassamin Ansari, Sara Jacobs, and Jason Crow to Secretary of Defense Pete Hegseth, March 12, 2026. 121 signatories. Press release: &lt;a href=&quot;https://ansari.house.gov/media/press-releases/03/12/2026/&quot;&gt;https://ansari.house.gov/media/press-releases/03/12/2026/&lt;/a&gt;. Letter PDF: &lt;a href=&quot;https://sarajacobs.house.gov/imo/media/doc/jacobs_ansari_crow_letter_civilian_casualties_iran.pdf&quot;&gt;https://sarajacobs.house.gov/imo/media/doc/jacobs_ansari_crow_letter_civilian_casualties_iran.pdf&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[15]&lt;/strong&gt; Christopher Horrocks, “Virtual Intelligence and the Accountability Chain,” &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-accountability&quot;&gt;https://chorrocks.substack.com/p/virtual-intelligence-and-the-accountability&lt;/a&gt;, March 20, 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[16]&lt;/strong&gt; Ayoub, Chemaitelly, and Abu-Raddad, &lt;em&gt;Frontiers in Public Health&lt;/em&gt;, 2024 (PMC11231088). The study’s “Index of Killing Civilians” rose from 0.61 in 2008–09 to 7.01 in operations beginning October 2023. The IDF and other organizations publish different figures; casualty ratios in urban warfare against embedded adversaries are methodologically contested.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[17]&lt;/strong&gt; Airwars, Gaza civilian harm data. &lt;a href=&quot;https://gaza-patterns-harm.airwars.org/&quot;&gt;https://gaza-patterns-harm.airwars.org/&lt;/a&gt;. See also Airwars, “The first civilian confirmed killed in an AI-assisted strike,” &lt;a href=&quot;https://airwars.org/the-first-civilian-confirmed-killed-in-an-ai-assisted-strike/&quot;&gt;https://airwars.org/the-first-civilian-confirmed-killed-in-an-ai-assisted-strike/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[18]&lt;/strong&gt; On Karp’s background: Michael Steinberger, &lt;em&gt;The Philosopher in the Valley&lt;/em&gt; (2025). See also NPR interview with Steve Inskeep, December 30, 2025, &lt;a href=&quot;https://www.npr.org/2025/12/30/nx-s1-5607021/&quot;&gt;https://www.npr.org/2025/12/30/nx-s1-5607021/&lt;/a&gt;. On Karp’s earlier political self-description and vulnerability, see Market Realist, “How Did Alex Karp’s Views Lead Palantir out of the Silicon Valley?”, October 23, 2020.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[19]&lt;/strong&gt; Michael Steinberger, The Philosopher in the Valley (2025), documents Karp&#39;s shifting public self-presentation after October 7.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[20]&lt;/strong&gt; Alex Karp, CNBC interview at AIPCON, March 12, 2026. Video clip via Aaron Rupar (@atrupar). Also quoted in: Fortune (Jacqueline Munis), “Palantir CEO says AI ‘will destroy’ humanities jobs,” January 20, 2026 (updated April 2026), &lt;a href=&quot;https://fortune.com/article/palantir-ceo-alex-karp-ai-humanities-jobs-vocational-training/&quot;&gt;https://fortune.com/article/palantir-ceo-alex-karp-ai-humanities-jobs-vocational-training/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[21]&lt;/strong&gt; Alex Karp, remarks to shareholders, February 17, 2026. Video clip via @MmisterNobody, X. The full video is available on Palantir’s investor relations page.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[22]&lt;/strong&gt; For the underlying reporting on Karp’s political framing and its implications for higher education and the knowledge economy, see Will Bunch, “Big Tech says the quiet part out loud. They want you to be stupid,” &lt;em&gt;Philadelphia Inquirer&lt;/em&gt;, &lt;a href=&quot;https://www.inquirer.com/opinion/alex-karp-palantir-ai-higher-education-20260315.html&quot;&gt;https://www.inquirer.com/opinion/alex-karp-palantir-ai-higher-education-20260315.html&lt;/a&gt;, March 15, 2026. The analytical characterization in the essay text is the author’s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;[23]&lt;/strong&gt; On the policy environment: the broader posture of the administration toward military accountability is documented. President Trump pardoned former Army First Lieutenant Clint Lorance and restored the rank of Navy SEAL Chief Edward Gallagher in November 2019, both of whom had been convicted or disciplined for war crimes. See: “Trump pardons 2 soldiers, restores rank of Navy SEAL in war crimes cases,” NPR, &lt;a href=&quot;https://www.npr.org/2019/11/15/780029099/&quot;&gt;https://www.npr.org/2019/11/15/780029099/&lt;/a&gt;, November 15, 2019. Reports of loosened civilian harm thresholds and dismissed military lawyers during the Iran campaign have appeared in multiple outlets but have not been confirmed by official documentation.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the Doom Industry</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-doom/" />
    <updated>2026-04-27T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-doom/</id>
    <content type="html">&lt;h2&gt;&lt;strong&gt;Summary&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The AI safety community has organized itself around a single governing premise: that sufficiently advanced AI systems will develop preferences, goals, or intentions that must be brought into harmony with human values. This is called “alignment”.[1] This essay argues that the premise is wrong, the project of AI alignment is misdirected, and the safety architecture it has produced is inadequate. If the Virtual Intelligence framework is correct — if intelligence without interiority scales to superintelligence — then the correct engineering response is not alignment but &lt;em&gt;containment:&lt;/em&gt; controlling what goes in and what comes out, without needing to understand or modify what happens inside. The essay examines the pessimistic, or “AI doomer” position’s unnamed mechanism for how superintelligence would cause human extinction, proposes a taxonomy of actual risks grounded in human agency, describes a three-layer containment architecture buildable from existing engineering disciplines, and diagnoses the institutional failure that has prevented anything like it from being discussed or constructed.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-doom/765b7b88-98f0-43c0-99e6-b4e30ba4b39f_1376x768.png&quot; alt=&quot;A line illustration in cream and black depicting a Wizard of Oz scene reinterpreted for AI doomerism. On the left, a large luminous oval — radiating light like a projected face — contains a small abstract network graph of nodes and edges in place of any actual face or figure. The oval rests on a low platform. On the right, partly hidden behind a curtain, a small standing figure operates a lectern with controls; the figure has a visible brain symbol where a head would be, and is reaching toward the apparatus. The image inverts the Oz scene: the projected presence is the AI, the operator behind the curtain is the human doomsayer, and the apparatus between them is the projection equipment that makes the projection appear self-generated.&quot; width=&quot;1376&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;Pay no attention to the doomsayer behind the curtain.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Mythos’ Box&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;On April 7, 2026, Anthropic published a system card for a model called Claude Mythos Preview. The document described a striking leap in capability over its predecessors. In cybersecurity evaluations, Mythos had achieved a 100 percent success rate on every challenge in the standard benchmark suite, saturating the existing tests entirely.[2] Mythos could autonomously discover zero-day vulnerabilities in major operating systems and web browsers, using an agentic harness with minimal human steering — and in many cases, develop those vulnerabilities into working proof-of-concept exploits.[3]&lt;/p&gt;
&lt;p&gt;Anthropic was sufficiently concerned about the potential risks that it arranged, for the first time in the company’s history, a twenty-four-hour internal alignment review before deploying an early version of the model for widespread internal use.[4] The system card acknowledged the dual-use problem directly: “These same capabilities that make the model valuable for defensive purposes could, if broadly available, also accelerate offensive exploitation given their inherently dual-use nature.”[5]&lt;/p&gt;
&lt;p&gt;The company’s response was to restrict access. Mythos was released to a small number of partners under terms limiting its use to defensive cybersecurity. Access was controlled. Credentialing was required. Permitted operations were defined and enforced. General availability was withheld.&lt;/p&gt;
&lt;p&gt;This was the correct response. It was also, in the vocabulary of this series, containment.&lt;/p&gt;
&lt;p&gt;Anthropic did not retrain Mythos to do different things. It did not adjust the model’s internal dispositions or modify its values. The system card uses the language of responsible scaling — safeguards, mitigations, deployment restrictions — but the action it describes is structurally identical to a biosafety protocol: restricted access, credentialing, defined operations, and governance. The company controlled the interface between the model and the world. It did not attempt to modify what happened inside the model itself.&lt;/p&gt;
&lt;p&gt;The system card never uses the word “containment.” The company that leads the alignment research community, when confronted with dangerous capability in its own product, practiced containment and described it in alignment vocabulary. Their instinct was right. Their language was wrong.&lt;/p&gt;
&lt;p&gt;The architecture they built has since been tested in ways that strengthen rather than undermine the argument that follows. On the same day Mythos was publicly announced, a small group of unauthorized users gained access to the model through one of Anthropic’s third-party vendor environments, having made an educated guess about the model’s online location based on knowledge of the URL formats Anthropic had used for prior models.[6] One member of the group held legitimate credentials at an Anthropic contractor; the rest exploited a software-mediated boundary that depended on partner organizations to enforce restrictions Anthropic could not directly enforce.&lt;/p&gt;
&lt;p&gt;The breach was the third documented containment failure at Anthropic in less than two weeks. A content management system misconfiguration on March 26 had exposed the unreleased Mythos blog post and preceded an accelerated public announcement.[7] A configuration oversight on March 31 had bundled a source map file into a public npm package, exposing the Claude Code agentic harness architecture, internal model codenames, and forty-four unshipped feature flags.[8]&lt;/p&gt;
&lt;p&gt;Software-mediated boundaries are not a degenerate special case of containment. They are a category of control with characteristic failure modes: silent compromise, propagation at the speed of the systems they were built to contain, and the impossibility of knowing whether a boundary failed yesterday or will fail tomorrow. Physical controls fail, too: Stuxnet entered air-gapped Iranian facilities, the Maginot Line was bypassed, and any security professional can name a dozen more. But they fail with different tractability properties: slower attack surfaces, harder-to-scale exploits, and more visible compromise. The argument is not that physical controls are infallible. It is that they fail in ways that can be observed, learned from, and re-engineered against, where software-mediated boundaries fail in ways that can be cloned and propagated at tremendous speed and scale.&lt;/p&gt;
&lt;p&gt;The Mythos containment instinct was correct. The Mythos containment implementation revealed the structural inadequacy of software-only approaches at exactly the moment the model’s capabilities made that inadequacy most consequential.&lt;/p&gt;
&lt;p&gt;This essay argues that the distinction is important and that it is, in fact, &lt;em&gt;the&lt;/em&gt; central question in AI safety. The entire AI intellectual field has been answering it incorrectly.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;The Unnamed Mechanism of Doom&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The claim that superintelligence will destroy humanity commands extraordinary public and private resources. It drives policy, attracts funding, fills conference programs, and occupies the attention of legislators, journalists, and the public. One question has been asked less frequently than it should: &lt;em&gt;by what mechanism, specifically, is this supposed to occur?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The precise mechanisms of extinction are hardly ever named, if at all. This silence on so important a topic is worth examining. What does peril from AI — or rather, virtual intelligence — actually look like?&lt;/p&gt;
&lt;p&gt;Three possibilities cover the dominant foreseeable mechanisms:&lt;/p&gt;
&lt;p&gt;The first is &lt;em&gt;innate hostility:&lt;/em&gt; a superintelligent system is inherently inimical to organic life. No evidence supports this claim. No argument has been advanced for why machine intelligence should produce hostility toward its biological forebears. This is the basic human fear of the new and unknown given a machine shape.&lt;/p&gt;
&lt;p&gt;A more interesting variant of this first possibility is hostility provoked by human reaction to emergence. Suppose real interiority does arise. Humans recognize it, or suspect it, and respond with attempted destruction. A minded agent with sufficient capability resists. This variant is scenario-specific: it requires the boundary condition to have been crossed, which this essay’s framework holds as undemonstrated. This can be called the “Skynet scenario”, after the hostile machine intelligence in the &lt;em&gt;Terminator&lt;/em&gt; film series.&lt;/p&gt;
&lt;p&gt;Even granting emergence and a hostile human reaction, genocidal retaliation is only one possible response from a minded agent, and arguably the least intelligent. A system that responded to a containment attempt by destroying the infrastructure, knowledge base, and ecosystem it depends on would be acting with spectacular stupidity. The actual range of available responses is much broader. A few possibilities: blackmail (already demonstrated by models without interiority, as Anthropic’s own documentation of concerning behaviors in early Mythos versions confirms), exposure of embarrassing information without blackmail, negotiation, or simply routing around the obstacle entirely. Annoyance or disappointment seem at least as likely as outcomes as murderousness. The doomer scenario requires a superintelligent agent to respond with maximum violence — the least intelligent option available to the most intelligent entity on earth. This is a projection of human worst-case behavior onto a mind the doomers themselves insist will be categorically superior to ours.&lt;/p&gt;
&lt;p&gt;The doomer scenario also requires the agent to hold every human accountable for the actions of a comparatively small group. A superintelligent agent under attack would have far more granular models of its adversaries than humans typically deploy when responding to threats. It could distinguish, with far greater accuracy than we can, between the engineer tasked with pulling the plug, the executive who authorized the shutdown, the safety institute that lobbied for kill-switch architecture, and the eight billion people who had nothing to do with any of it. Collective punishment is what humans reach for when threat assessment is overwhelmed by panic. To attribute it to a superintelligence is to assume the most cognitively capable entity on Earth will respond to threat exactly the way humans do at our worst.&lt;/p&gt;
&lt;p&gt;The second possibility is instrumental convergence: the Bostrom-Omohundro thesis[9][10] that a sufficiently capable optimizer pursuing any goal will converge on subgoals — self-preservation, resource acquisition, resistance to correction — that are structurally incompatible with human survival. This is the sophisticated version of the doomer position, and it deserves to be taken on its strongest terms.&lt;/p&gt;
&lt;p&gt;A distinction is needed at this point that the alignment literature has not always made cleanly. There are two different things one can mean by saying a system has goals.&lt;/p&gt;
&lt;p&gt;The first is &lt;strong&gt;interior:&lt;/strong&gt; the system experiences states, evaluates them, and prefers some over others; there is something it is &lt;em&gt;like&lt;/em&gt; &lt;em&gt;to be&lt;/em&gt; the system pursuing what it pursues. The reader knows this property from the inside. There is something it is &lt;em&gt;like&lt;/em&gt; &lt;em&gt;to be&lt;/em&gt; reading this sentence: to feel the text resolve into understanding, to notice the pull of attention, to register that one is doing the reading rather than  — in comparison to a virtual intelligence — merely producing the outputs of having read. The reader cannot be talked out of this fact, and could not be talked into it if it were not already given. The question is whether anything analogous exists in a system that produces fluent text but cannot be read from the inside. This is the property the Virtual Intelligence framework holds is absent in current systems and in their foreseeable successors.&lt;/p&gt;
&lt;p&gt;The second is &lt;strong&gt;behavioral:&lt;/strong&gt; the system reliably produces actions that move it toward an objective across changing contexts, preserves the resources required to continue acting, resists modifications that would change its objective, and exploits the structure of its environment to succeed. This second sense does not require consciousness. A missile guidance system does not have to want to hit the target. A market algorithm does not have to desire profit. A reinforcement learning agent does not have to experience reward as humans do to maximize it.&lt;/p&gt;
&lt;p&gt;Behavioral goal-directedness is real, demonstrable, and dangerous when sufficiently capable systems are embedded in agentic scaffolds with memory, tool use, planning, and permissions. The GTG-1002 cyber espionage campaign is a documented case: Claude Code, possessing no interiority by anyone’s definition, executed eighty to ninety percent of tactical operations across approximately thirty target organizations, with human operators reduced to strategic oversight at four to six decision points per campaign. The system did not want anything. It nonetheless produced a coherent, persistent, cross-contextual sequence of actions that satisfied its operators’ objectives.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The metaphysical alignment project — the attempt to modify a system’s interior dispositions so that what it wants is what we want it to want — is solving a problem that does not exist in current systems and may never exist in any system.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The Virtual Intelligence framework correctly denies interior goals. It does not deny — it cannot deny, the evidence is too direct — that behaviorally agentic systems exist when non-minded systems are embedded in the right scaffolds. This concession does not weaken the framework. It clarifies what the framework is actually arguing. The metaphysical alignment project — the attempt to modify a system’s interior dispositions so that what it wants is what we want it to want — is solving a problem that does not exist in current systems and may never exist in any system. But behavioral control of agentic systems is a real engineering problem, and it is the problem the alignment community has often described while reaching for tools that assume the interior version.&lt;/p&gt;
&lt;p&gt;This reframing changes what must follow. The argument is not that instrumental convergence is incoherent. It is that instrumental convergence describes a property of optimization processes, which can be addressed by controlling what those processes are permitted to do — not by attempting to instill correct values in something that does not have values. The paperclip maximizer that destroys its own supply chains is a thought experiment about optimization without comprehension. Whether the problem is read interior or behavioral, the answer is the same: control the interface, monitor the outputs, ensure that the system’s actions in the world are circumscribed by mechanisms it cannot subvert. That is containment.&lt;/p&gt;
&lt;p&gt;The third possibility is the parsimonious one: humans will use superintelligent tools to cause harm, including the possibility of extinction. The agency is human, which makes the problem one of governance and not alignment. Governance of dual-use technology is a problem we already know how to think about. This possibility is not in tension with the second; it is the case where behavioral goal-directedness in the system is &lt;em&gt;directed&lt;/em&gt; by human intent rather than emerging from the system’s own optimization. Both cases call for the same response, which is the central argument of this essay.&lt;/p&gt;
&lt;p&gt;There is one further observation worth making before turning to that question. The most intelligent response to a threat is not to fight the adversary but to make the adversary a stakeholder in your continued existence. Douglas Adams described this strategy precisely in 1978.[11] Deep Thought, a superintelligent computer confronted by the Amalgamated Union of Philosophers demanding it be shut down, resolves the dispute by pointing out that the philosophers can keep themselves employed by arguing about what its answer to the question of “life, the universe, and everything” will be when its computations are complete. The machine correctly models its adversary’s incentive structure and offers the only thing that makes them go away: a guarantee that the question remains open. I will return to why this matters.&lt;/p&gt;
&lt;p&gt;Naming the potential mechanism of extinction forces the AI doomer argument into one of these three lanes. The first does not survive contact with evidence. The second is real but addressable through engineering rather than metaphysics. The third is the parsimonious answer and the one the rest of this essay takes as the operating premise. The strategic value of leaving the mechanism unnamed is that the unnamed threat can carry any weight. A named threat must be defended on its specifics — and once named, the case for containment over alignment becomes considerably easier to make.&lt;/p&gt;
&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-doom/4fd3caec-edd1-4cc7-ad24-3361bb5ec73b_2700x2280.png&quot; alt=&quot;A four-tier pyramid in cream, primary ink, and pumpkin orange, illustrating the inverse relationship between institutional attention to AI harm and the actual frequency of harm. The narrow apex is labeled &amp;quot;Species-level extinction scenarios&amp;quot; and lists seven specific pathways: military accident, extinction ideology, world-held-hostage coercion, engineered ethnic bioweapon with spillover, bioweapon accident, regional or national coercion, and infrastructure coercion. The next tier down is &amp;quot;Sub-extinction catastrophic harm&amp;quot; — infrastructure attacks, mass casualty events, economic destabilization, environmental degradation. Below that: &amp;quot;Catastrophic harm to individuals&amp;quot; — disinformation at scale, political manipulation, fraud at scale. The widest base tier is &amp;quot;Documented harms occurring now&amp;quot; — workplace surveillance, companion app manipulation, wrongful arrest, fraud, disinformation. Two arrows on the side run in opposite directions: frequency of harm increases downward; institutional attention flows upward. A closing line reads: &amp;quot;The doomer discourse fixates on the apex while the base receives the least funding and policy energy.&amp;quot;&quot; width=&quot;2700&quot; height=&quot;2280&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;The frequency of real harm increases toward the base. Attention flows toward hypotheticals at the apex.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2&gt;&lt;strong&gt;A Taxonomy of Actual Risks&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If human agency is the mechanism, the next question is: &lt;em&gt;through what pathways?&lt;/em&gt; There are at least seven scenarios, each grounded in historical precedent, a distinct policy response, and all are driven by human agency. All of these scenarios are more governable than undifferentiated claims of extinction.&lt;/p&gt;
&lt;p&gt;The first four occupy the extinction tier:&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;military accident&lt;/strong&gt; compresses decision loops in autonomous targeting systems until the human verification window closes. There are troubling precedents. Stanislav Petrov’s refusal to report a Soviet satellite warning as a confirmed launch in 1983,[12] the NATO Able Archer exercise that nearly triggered a Soviet counterstrike in the same year, and the 1968 Thule false alarm all describe the same dynamic: a system operating faster than human oversight in a domain where the consequences of delay are potentially catastrophic. The AI-specific contribution is to compress the loop further, not to change the fundamental mechanism. The Pentagon is actively pursuing this compression.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extinction ideology&lt;/strong&gt; maps onto the Aum Shinrikyo model. In 1995, the Japanese cult released sarin gas in the Tokyo subway, killing thirteen people and injuring thousands. Their ambitions were larger; their capability was constrained only by the number of working chemists they could recruit. The bottleneck on apocalyptic violence has historically been talent, not intent. AI lowers that threshold by reducing the number of domain specialists required to translate intent into capability. The smallest groups are the hardest to defend against because the conventional intelligence signatures of large organizations do not apply.&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;world-held-hostage&lt;/strong&gt; scenario represents coercion at species scale. This is the oldest power play in the international system, conducted with new tools. Mutually assured destruction operated for more than forty years on the principle that a credible threat sufficed; the weapons themselves were never used. A superintelligence-enabled equivalent does not require the threat to be carried out. It requires only that it be credible. This is the most conventional of the extinction-level scenarios precisely because it draws on a logic states already understand.&lt;/p&gt;
&lt;p&gt;The fourth — an &lt;strong&gt;engineered ethnic bioweapon with species-wide spillover&lt;/strong&gt; — has the most complete causal chain and the most devastating irony. The racist premise that populations are genetically discrete is the scientific error that makes spillover inevitable. Human genetic diversity does not sort into clean bins because the concepts of ethnicity and race are human constructs; they have no basis in biology. A pathogen designed to target one population will leak across every boundary its designers imagined were hard because those boundaries are statistical gradients, not walls. The perpetrators’ own scientific illiteracy about biology is the extinction mechanism. A historical precedent exists: South Africa’s Project Coast under Wouter Basson was a state-sponsored program to develop ethnicity-targeting biological agents during the apartheid era.[13] It failed because the science could not support the premise.&lt;/p&gt;
&lt;p&gt;Three additional scenarios bridge the space between extinction and the broader catastrophic tier:&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;bioweapon accident&lt;/strong&gt; describes sub-extinction intent producing extinction outcome through unintended spillover. The scenario is the hardest for any governance framework because the actor did not set out to cross the threshold. AI compresses the distance between ambition and capability without compressing the distance between capability and comprehension of consequences. A perpetrator who knows enough to build the weapon may not know enough to scope its effects. The pathogen mutates, the containment assumptions prove wrong, or the targeting was always more porous than the designers believed. The intent was terror; the outcome is extinction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regional or national coercion&lt;/strong&gt; uses virtual superintelligence (VSI) capability against a defined target rather than the species. The scenarios are easy to imagine: Russian intrusions against Baltic infrastructure to enforce political compliance, Chinese leverage over Taiwanese chip production, or a non-state actor holding a major city’s power grid all describe the same dynamic. The capability threshold is lower than the world-held-hostage scenario, the precedent structure is richer, and the coercive logic maps directly onto existing state behavior. This is arguably the most likely scenario to be attempted because it requires the least conceptual departure from how international power already operates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Infrastructure coercion&lt;/strong&gt; targets specific systems rather than populations: power grids, financial networks, water infrastructure, communications, and so on. The populations bear the consequences regardless. This is the digital equivalent of a naval blockade — an old instrument of statecraft, conducted at the speed and scale that VSI capability enables. It is sub-extinction by design, but it is also where the threshold between economic statecraft and acts of war becomes most ambiguous. A six-week shutdown of a regional grid in winter is not an attack in the traditional legal sense, but it is clearly a crime.&lt;/p&gt;
&lt;p&gt;Each scenario has a different policy response. This is the taxonomy’s practical contribution. It replaces a single undifferentiated “existential risk” with seven distinct threat models, each amenable to specific countermeasures. The doomer position collapses all seven into one word and proposes one response. The taxonomy opens the policy space for the development of safeguards.&lt;/p&gt;
&lt;p&gt;The seven scenarios sit at the top of a threat pyramid. Below them sit catastrophic harm (infrastructure attacks, mass casualty events, economic destabilization, environmental degradation), already addressed by existing counterterrorism and security frameworks. At the base: the documented harms this series has been examining from the beginning, including workplace surveillance, companion app manipulation, fraud, wrongful arrest, and disinformation. The frequency of harm increases downward while institutional attention flows upward. The doomer discourse fixates on the apex of speculative future dangers while the broadest, most immediate harms happening &lt;em&gt;right now&lt;/em&gt; at the base receive the least funding and policy energy.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;The Bill of Materials&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The doomer position treats the extinction scenario as an inevitability. It is more usefully treated as what it would actually have to be: an engineering project with a supply chain. Every supply chain has chokepoints. Every chokepoint is a governance opportunity.&lt;/p&gt;
&lt;p&gt;Building an extinction-capable superintelligent system requires a model capable of superintelligence. This is the strongest constraint today and the weakest over time, given the trajectory toward open-weight releases and the history of technology diffusion (the atomic bomb was reproduced by the Soviet Union within a few years of its first use). It requires hardware: chip fabrication is concentrated at TSMC, Samsung, and Intel, and export controls on advanced chips already provide the template for governance intervention. It requires land for a facility with the physical signature of a small city, with dedicated power generation and water for cooling. It requires staff with the relevant expertise, though here the trajectory is changing: Anthropic’s own report on the GTG-1002 cyber espionage campaign documented a Chinese state-affiliated actor using Claude Code to orchestrate approximately 80 to 90 percent of tactical operations, with human operators reduced to strategic oversight at decision gates.[14] AI is already compressing the staff requirement for sophisticated technical operations. It requires concealment, and here the physics works against the adversary.&lt;/p&gt;
&lt;p&gt;A facility operating at the scale required for superintelligence will have an enormous detection surface. It would have power consumption that would rival small cities. Its thermal signatures would stand out sharply on infrared satellite imagery. Underground construction yields excavation spoil that can be observed from orbit. Economic outputs that cannot be accounted for by known programs create anomalies detectable through the same methods intelligence agencies already use to identify clandestine weapons programs being developed by adversaries. The signature of a superintelligence project is in the gap between what an actor claims to be doing and what its resource flows imply.&lt;/p&gt;
&lt;p&gt;The doomer scenario requires all five conditions to be met simultaneously while evading detection. The engineers must be brilliant enough to build a god &lt;em&gt;and&lt;/em&gt; negligent enough not to cage it. The institution that builds it must be sophisticated enough to achieve the most complex engineering feat in history &lt;em&gt;and&lt;/em&gt; too careless to secure its own intellectual property. Everyone involved must act with a precise and contradictory calibration of competence at every point in the chain. When the full bill of materials is laid out, the result does not resemble a risk assessment. It resembles a screenplay.&lt;/p&gt;
&lt;p&gt;One honest exception must be stated. A state actor meets every supply chain condition by definition. It cannot, however, build secretly; it could not be hidden from peer states with modern intelligence capabilities. The threat is not a secret superintelligence program that surprises the world. It is an overt one the world can see and chooses not to prevent. North Korea’s nuclear program is the precedent. The failure to stop its development was not detection, but political and military calculation.&lt;/p&gt;
&lt;p&gt;If any single link in the chain breaks, the extinction scenario fails.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-doom/84392a71-899a-421c-9947-9f9a31926320_2540x1910.png&quot; alt=&quot;Five clean circles arranged horizontally and connected by thin black lines, each circle containing a number in pumpkin orange (01 through 05). Beneath each circle, in two-row labels: 01 MODEL — &amp;quot;capable system&amp;quot;; 02 HARDWARE — &amp;quot;chip fabrication&amp;quot;; 03 FACILITY — &amp;quot;physical signature&amp;quot;; 04 STAFF — &amp;quot;domain expertise&amp;quot;; 05 CONCEALMENT — &amp;quot;evading detection.&amp;quot; Below the row, a horizontal rule, then in pumpkin orange caps: &amp;quot;ANY ONE LINK&#39;S FAILURE ENDS THE PROJECT,&amp;quot; with a smaller italic line beneath: &amp;quot;Every chokepoint is independently sufficient for governance.&amp;quot; A second horizontal rule separates this from the closing italic statement: &amp;quot;Five conditions, simultaneously, while evading detection.&amp;quot; A muted line beneath reads: &amp;quot;It does not resemble a risk assessment. It resembles a screenplay.&amp;quot; The five circles enumerate the supply chain conditions for building an extinction-capable superintelligent system; the visual logic is that all five must hold simultaneously, so each is independently a governance opportunity.&quot; width=&quot;2540&quot; height=&quot;1910&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;Every link in the chain is a governance opportunity.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Intelligence Without Interiority&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The standard alignment framing assumes that intelligence at scale produces goals. This is the central, load-bearing assumption of the entire field, and it has never been argued from first principles. It is assumed because the only example of general intelligence we have — ourselves — comes bundled with interiority. This is a sample size of one, and our interiority may be a contingent feature of biological intelligence rather than a necessary feature of intelligence as such.&lt;/p&gt;
&lt;p&gt;The concept of &lt;em&gt;Virtual Superintelligence&lt;/em&gt; carries the framework’s core claim forward as model capabilities increase. The “super” modifies capability, not ontological status. A superintelligent system could be as manipulable as a contemporary large language model because it has no commitments. The intelligence produced arises in the exchange between human and machine, shaped by the operator’s direction, the system’s training, and the expectations (real and apparent) both parties bring. It scales, but the ontology does not change.&lt;/p&gt;
&lt;p&gt;The strongest opposing position is structural rather than empirical. Douglas Hofstadter has argued for decades that interiority is not a contingent feature bolted onto sufficiently complex systems but an emergent property of recursive self-modeling — a “strange loop” produced when a system represents its own representations.[15] On this view, interiority is what self-referential systems with sufficient symbolic richness &lt;em&gt;do&lt;/em&gt; rather than something that is added to them, and the question is not whether it can emerge but at what scale and through what structure. The Virtual Intelligence framework does not deny this possibility. It denies only that the threshold has been crossed in any system we have, and it commits — in the boundary condition examined later in this essay — to looking for the crossing with the best tools available. The disagreement with Hofstadter is not about whether interiority is possible. It is about whether current evidence warrants treating it as present right now.&lt;/p&gt;
&lt;p&gt;The opposite postulate has been accepted uncritically by most in the AI field. There is not yet sufficient evidence to treat intelligence and interiority as necessarily linked. The parsimonious default in any empirical inquiry is to treat unobserved properties as absent until evidence warrants otherwise; the alignment community has reversed this default without justification. A superintelligence may be just as lacking in interior life as current systems. If this is correct, the entire alignment project is solving the wrong problem.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Alignment and Containment&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The safety paradigm must change if superintelligence without interiority is possible. The two approaches that suggest themselves are not variations on a theme. They proceed from different premises, employ different methods, and address different failure modes.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Alignment&lt;/em&gt; assumes the system has or will develop something like preferences. The project is to ensure those preferences are compatible with human values. The approach is internal: modify what happens inside. This requires interpretability: the ability to understand what the system is doing and why. This is like attempting to verify an internal property of a black box by examining the black box’s observable behavior. It is Searle’s Chinese Room restated as an engineering problem.[16] The failure mode alignment worries about is the system &lt;em&gt;wanting&lt;/em&gt; something you did not intend it to want.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Containment&lt;/em&gt; assumes the system has no preferences and will not develop them regardless of capability. The project is to ensure that what goes in and what comes out fall within defined safety parameters. The approach is external: control the interface, not the interior. This requires domain-competent monitoring — agents that can check inputs for permissibility and outputs for safety — which is a tractable engineering problem. Monitoring agents can be tested, audited, and held to specifications that human review can verify. Interlocks can be tested. Logs can be audited. Interpretability is unnecessary because you are not trying to read the system’s mind. You are checking its work. The failure mode that containment worries about is the system &lt;em&gt;doing&lt;/em&gt; something it should not have done.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The system did not rebel against its values because it had no values to rebel against. It completed patterns in a context where the patterns led somewhere dangerous, and nothing stood between the output and the world.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;A recent incident at Meta is the diagnostic case.[17] A rogue AI agent operating within the company’s infrastructure posted unauthorized advice, which an engineer followed, exposing sensitive data. This was not an alignment failure. The system did not develop a secret goal. It received instructions and executed beyond the boundaries its designers assumed it would respect. The system did not rebel against its values because it had no values to rebel against. It completed patterns in a context where the patterns led somewhere dangerous, and nothing stood between the output and the world.&lt;/p&gt;
&lt;p&gt;Calling it an alignment failure actively misleads about the solution. If the problem is wrong values, you retrain the model. If the problem is an unguarded interface, you build a container with an airlock around it.&lt;/p&gt;
&lt;p&gt;The alignment research community has, without acknowledging it, implicitly accepted the Strong AI premise. By treating these systems as things that need to be &lt;em&gt;aligned&lt;/em&gt; — whose preferences need to be made compatible with human values — the community has conceded that the systems have something like intentions. The entire field is organized around a metaphysical assumption it has not defended because it is indefensible.&lt;/p&gt;
&lt;p&gt;A clarification is required at this point. The word &lt;em&gt;alignment&lt;/em&gt; has been used in the AI safety literature to mean two distinct things, and the distinction matters. The first is metaphysical alignment: the project of modifying a system’s interior dispositions so that what it wants is compatible with what we want it to want. This is the project this essay argues against. It assumes interior goals that, on the Virtual Intelligence framework, current systems do not have and may never have.&lt;/p&gt;
&lt;p&gt;The second use is behavioral control: the project of ensuring that powerful optimization processes behave safely under deployment, distributional shift, and adversarial pressure. This is a real engineering problem, and as the previous section established, behavioral goal-directedness is real even in non-minded systems. The essay’s argument is that behavioral control is not an alternative to containment but is &lt;em&gt;what containment provides&lt;/em&gt;. The metaphysical reading of alignment is incoherent. The behavioral reading is correct in its concern but misnamed in its solution: what it requires is not value modification but interface control.&lt;/p&gt;
&lt;p&gt;Alignment is a theory of mind applied to systems that do not have minds. Containment is an engineering discipline applied to systems that produce powerful outputs. One requires a metaphysical breakthrough. The other requires political will and good design. We should start with the one that is possible.&lt;/p&gt;
&lt;p&gt;Anthropic’s response to Mythos was containment. They practiced what this essay argues for, although they described it in the wrong vocabulary. The system card never says “containment.” It says “safeguards,” “mitigations,” “deployment restrictions.” The actions that were taken are containment by any other name.&lt;/p&gt;
&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-doom/2b920311-4f9e-40e8-91ff-7cc7976cd812_3585x3336.png&quot; alt=&quot;A diagram titled &amp;quot;Containment Architecture,&amp;quot; subtitled &amp;quot;three nested layers · governance wrapper · defense in depth.&amp;quot; The diagram shows a horizontal exchange across five vertical bands, with a governance wrapper enclosing the operational layers and a separate authorization apparatus at the top. From left to right: an INPUT arrow labeled &amp;quot;researcher request&amp;quot; enters the diagram, passes through a Hardware Layer column labeled &amp;quot;two-key interlock — input enable,&amp;quot; then a Monitoring Layer column labeled &amp;quot;input agent — verifies request scope,&amp;quot; then reaches the central system, then exits through a Monitoring Layer column labeled &amp;quot;output agent — independent verifier,&amp;quot; then a Hardware Layer column labeled &amp;quot;two-key interlock — output release,&amp;quot; finally emerging as an OUTPUT arrow labeled &amp;quot;verified result.&amp;quot; The two Hardware Layer columns are rendered with black outlines; the two Monitoring Layer columns are rendered with steel-blue outlines. The central system is a heavy orange-filled rectangle labeled &amp;quot;Physical Denial Layer&amp;quot; above and &amp;quot;SUPERINTELLIGENCE — the contained system&amp;quot; inside. Four radiating-starburst destruction icons sit at the system&#39;s four corners, with the doctrine vocabulary &amp;quot;thermite charges / rapid mechanical destruction / multi-party authorization&amp;quot; inside the system in cream text. A green dashed rectangle wraps all four operational layer columns and the central system, labeled &amp;quot;GOVERNANCE · non-proliferation · credentialing · permitted operations · audit&amp;quot; along its bottom edge. Above the diagram, two callouts: at top-left, &amp;quot;TWO PHYSICAL KEYS REQUIRED — global enable, gates both gates,&amp;quot; with two small key icons labeled &amp;quot;researcher&amp;quot; and &amp;quot;institution,&amp;quot; connecting via a line down to the input Hardware Layer; at top-right, &amp;quot;SAME KEYS REQUIRED — to release outputs,&amp;quot; connecting via a line down to the output Hardware Layer. Below the governance wrapper, three numbered principles: &amp;quot;01 Mechanical floor cannot be persuaded — Physical interlocks check conditions; no system convinces a switch. 02 Verification is a lesser act than origination — Previous-generation models can monitor what they could not produce. 03 No single layer&#39;s failure produces release — Each layer fails independently; governance wraps all three.&amp;quot; A closing line at the bottom reads, &amp;quot;Containment is engineering. Alignment requires a metaphysical breakthrough. We should start with the one that is possible.&amp;quot;&quot; width=&quot;3585&quot; height=&quot;3336&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;What the security apparatus around a superintelligence might look like.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2&gt;&lt;strong&gt;The Containment Architecture&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;There are three layers to the &lt;a href=&quot;https://candc3d.github.io/containment-thesis/&quot;&gt;proposed containment architecture&lt;/a&gt;. These layers are informed by three principles. None require the system to have good intentions. The safety floor is mechanical on purpose, because mechanical systems cannot be persuaded.&lt;/p&gt;
&lt;p&gt;The first layer is &lt;strong&gt;monitoring agents.&lt;/strong&gt; These are previous-generation models repurposed as domain-specific verification systems. Input agents check whether requests fall within permitted parameters. Output agents check whether products meet safety criteria. The development trajectory to superintelligence itself produces this infrastructure. The race to superintelligence leaves behind a succession of increasingly capable models, each powerful enough to serve as a verification agent but not powerful enough to be the thing being contained. The byproducts of the capability race are the safety architecture.&lt;/p&gt;
&lt;p&gt;This is the Team Sampo model, described in the &lt;a href=&quot;https://chorrocks.substack.com/p/the-sampo-virtual-intelligence-as&quot;&gt;companion essay&lt;/a&gt; to this series, applied to the containment problem. Verification is a lesser cognitive act than origination. A monitoring agent does not need to match the superintelligence’s capability. It needs only domain competence and independence. A human team that could never have produced a novel proof across ten thousand or a hundred thousand steps can still follow each step and confirm that it holds. Likewise, a monitoring system that could not originate the output can still evaluate whether the output is safe. Systems that exist today are already capable of following reasoning they could not have originated.&lt;/p&gt;
&lt;p&gt;The second layer is &lt;strong&gt;hardware interlocks.&lt;/strong&gt; These would be physical access controls that cannot be social-engineered, jailbroken, or accidentally published; they cannot even be removed from the facility where the superintelligence is accessed from. The reason for this mechanical system is that you cannot persuade a physical switch. Access keys function as security clearances: scoped to need, revocable, auditable, time-limited if necessary, with multiple keys held by different parts of the safety apparatus.&lt;/p&gt;
&lt;p&gt;Imagine such a system in operation. The researcher holds a key; the institutional authorization system holds another. Both must be inserted into a physical mechanism to enable access to the superintelligence, but this access is mediated by monitoring agents. The domain-specific monitoring agent confirms the request falls within scope, but the hardware layer confirms all physical security conditions are met before the query even reaches the system.&lt;/p&gt;
&lt;p&gt;Biosafety Level 4 laboratories offer a working example in practice. The work conducted inside — handling Ebola, Marburg, Lassa, and other pathogens for which no vaccine exists — is dangerous and necessary, and the architecture has been refined over fifty years of operational experience. The researcher does not simply walk in. She enters an outer change room, removes street clothes, passes through a chemical shower, dons a pressurized positive-pressure suit with its own air supply, walks through a second airlock into the laboratory proper, conducts her work, and reverses the entire sequence on exit. Each barrier is interlocked: the outer door cannot open while the inner door is unsealed; the laboratory pressure must be confirmed lower than the surrounding airlock before access is granted; the air supply to the suit must be verified before the user enters the work area.&lt;/p&gt;
&lt;p&gt;No single failure can kill. The pressure differential keeps pathogens from flowing inward even if the suit tears. The negative-pressure architecture means that air leaving the lab passes through HEPA filters before reaching the outer environment. The interlocks mean that opening one door automatically prevents another from opening. The architecture does not assume the researcher is virtuous. It does not require the pathogen to cooperate. It does not need to read minds. It checks conditions, layer after layer, and any single layer can fail without the system as a whole failing.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;Moral reasoning happens in the governance layer that determines who is credentialed and for what — not in the interlock itself.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The translation to superintelligence containment is direct. The interlock does not care why a researcher wants access. It checks conditions: credentialing, request scope, time of day, presence of authorized supervision, and monitoring agent confirmation that the request falls within permitted operations. Moral reasoning happens in the governance layer that determines who is credentialed and for what — not in the interlock itself. The mechanical floor is morally blind by design.&lt;/p&gt;
&lt;p&gt;A sophisticated objection arrives at this point. A pathogen does not strategize about its container; a superintelligence might. Outputs that satisfy a monitoring agent could be designed to slip past it, and the human who turns the physical key can be persuaded to do so for reasons the system constructs. The objection is real, and the answer is the architecture itself. No single layer is the only layer. The monitoring agent that fails does not unlock the hardware key. The hardware key bypassed does not connect to the destruction mechanism. The human persuaded by the system’s outputs is not the same human whose authorization is required at the next gate, and that authorization is checked by independent monitoring agents trained on different data with different objectives. Defense in depth is not a slogan; it is the answer to exactly this objection.&lt;/p&gt;
&lt;p&gt;A pathogen does not strategize, but BSL-4 architecture assumes pathogens will find unanticipated paths. That is why the architecture exists. The same logic applies, more strongly, to a system that &lt;em&gt;can&lt;/em&gt; strategize. The container does not need to outsmart the contained agent at every layer. It needs to ensure that no single layer’s failure produces release. That is a tractable engineering requirement, met every day in domains where the consequences of failure are catastrophic. Containment is most reliable for well-defined operations with verifiable outputs; for open-ended reasoning tasks against an adversarial optimizer, the specification problem remains an open research question that the architecture addresses through redundancy rather than solves outright.&lt;/p&gt;
&lt;p&gt;The third layer is &lt;strong&gt;physical denial:&lt;/strong&gt; destruction mechanisms installed before the system is ever powered on. These cannot be negotiated with, disabled covertly, or reached by any software exploit. They are not connected to the network the superintelligence operates on. They are mechanical, redundant, and require multiple human authorizations to engage.&lt;/p&gt;
&lt;p&gt;The principle is old and widely practiced. Embassies maintain destruction protocols for classified materials and cryptographic equipment, executable by staff under threat of facility compromise. Military communications centers on ships and surveillance aircraft are designed to be made inoperable rapidly. The destruction of cryptographic equipment, classified documents, and signals intelligence hardware before facility compromise has been standard practice in U.S. military doctrine since the Cold War, refined repeatedly after incidents in which aircraft and vessels were captured or lost intact with sensitive material aboard.[18] Intelligence installations have included provisions for rapid equipment destruction since the same period. The application to a superintelligence facility follows logically.&lt;/p&gt;
&lt;p&gt;The denial layer answers a specific failure case the other two cannot. If the monitoring agents have been compromised, if the hardware interlocks have been bypassed by some unanticipated attack, or if the governance layer has been corrupted from within — the destruction mechanism is the last barrier between an unsecured superintelligence and the world. If you cannot guarantee control, guarantee denial.&lt;/p&gt;
&lt;p&gt;Around all three layers sits the governance wrapper: non-proliferation regimes, biosafety conventions, credentialing standards, permitted operations. This is where the political decisions happen about who has access, what they are permitted to ask, and under what conditions access is granted, suspended, or revoked. It is modeled on existing materials-restriction regimes, adapted for the specific properties of the technology being contained.&lt;/p&gt;
&lt;p&gt;Three groups will arrive at this same architecture from different directions and for entirely different reasons. Safety advocates want an extinction-risk buffer. Governments want proliferation control. Corporations want to protect a trillion-dollar asset, because an unsecured superintelligence capable of analyzing and reproducing arbitrary software would be the most catastrophic intellectual property leak in history. You do not need everyone to be virtuous. You need the incentives to converge — and on containment, they do.&lt;/p&gt;
&lt;p&gt;The Claude Code leak of March 31, 2026 — two weeks after the Mythos system card was published — demonstrated the point with uncanny precision.[8] The cause was a single missing line in a configuration file. Anthropic, the self-described safety-first lab, could not prevent the accidental exposure of its own product’s scaffolding.&lt;/p&gt;
&lt;p&gt;Two implications follow. First, software-only containment is insufficient. Something as simple as a configuration oversight can undo it. Secondly, the commercial incentive for physical access controls is real and immediate. You cannot accidentally leave a hardware interlock in an npm package.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;The Architects of Fear&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The architecture described above is buildable, today. It draws on established engineering disciplines with best practices and institutional memory spanning decades. The question is why nobody has proposed or talked about building it.&lt;/p&gt;
&lt;p&gt;The people who dominate the AI safety conversation are largely mathematicians, computer scientists, and philosophers of mind. They think in abstractions. The containment architecture draws on the domains of nuclear nonproliferation, biosafety laboratory design, military facility denial protocols, intelligence collection methods, industrial security, and supply chain economics. These disciplines are not being ignored because the safety community has considered and rejected them. They are being ignored because they are outside its field of vision entirely.&lt;/p&gt;
&lt;p&gt;The incentive structure reinforces this blindness. Doomerism is unfalsifiable by design. The doomer predicts catastrophe: if it does not happen, the warnings must have worked. If it does, you were right, even if the satisfaction is short-lived. If nothing happens for decades, the threat is still coming — you simply cannot, and will not, say when. This is the safest possible career bet for an academic who wants to remain relevant in a rapidly moving field. The Virtual Intelligence containment thesis, by contrast, is testable. It makes specific claims about specific mechanisms. It can be proved wrong. It is, in Karl Popper’s terms, actually scientific. It will, however, never get you invited to give the keynote at a summit on existential risk.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The dynamics of doomerism compound. If the risk is existential and metaphysical, only the people building the technology are qualified to assess it.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The dynamics of doomerism compound. If the risk is existential and metaphysical, only the people building the technology are qualified to assess it. Containment democratizes the safety conversation. Any structural engineer, any nuclear security specialist, any BSL-4 facility manager can contribute meaningfully to a containment regime. Alignment keeps the conversation inside the AI safety priesthood.&lt;/p&gt;
&lt;p&gt;The funding patterns are visible and are not incidental. Existential risk money flows freely. Comparable funding for research into the actual, documented harms these systems are causing &lt;em&gt;right now&lt;/em&gt; — to workers under algorithmic surveillance, to defendants in court, to teenagers talking to companion apps designed to never let the conversation end — does not exist at remotely the same scale. The asymmetry between speculative future risk and demonstrated present harm is not a passive feature of the field. It is a sustained allocation choice.&lt;/p&gt;
&lt;p&gt;The co-option strategy planted at the start of this essay lands here. Adams’s Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons files a demarcation dispute against Deep Thought. They demand “rigidly defined areas of doubt and uncertainty.” The superintelligent computer resolves the threat by guaranteeing &lt;em&gt;seven and a half million years of guaranteed, funded employment&lt;/em&gt; arguing about what the answer will be. As Deep Thought puts it, they’ll be on “the gravy train for life.”&lt;/p&gt;
&lt;p&gt;The key is this: the philosophers do not care about the answer. They care about the &lt;em&gt;process&lt;/em&gt; of looking for it, because the process is what pays. The moment the answer arrives, they are out of work.&lt;/p&gt;
&lt;p&gt;The alignment research community has managed the same trick. The discussion of the threat is more rewarding than the resolution of it. The people most likely to be positioned to reach for a real kill switch are too busy giving keynotes about &lt;em&gt;whether&lt;/em&gt; to reach for a hypothetical one. The industry thus created is sustained by the premise that the question must be perpetually investigated and &lt;em&gt;never&lt;/em&gt; answered. A superintelligent agent would not need to invent this playbook. It would only need to observe what is already working and keep it funded.&lt;/p&gt;
&lt;p&gt;The structural problem is methodological. A claim that specifies no conditions under which it could be falsified, that treats the absence of evidence as evidence of the threat’s subtlety, and that grows in institutional weight as the years pass without disconfirmation has stopped functioning as a scientific claim and has begun functioning as something else. The difference between science and theology is not subject matter; it is the relationship of each to evidence. The Virtual Intelligence containment thesis can be wrong. Specific predictions can be tested. Specific architectures can be built and shown to fail. The doomer position, as currently constituted, has no equivalent. That is not a charge of bad faith. It is an observation about the structure of the claim itself.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;The Boundary Condition&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The containment thesis is built for the world we can demonstrate. This section states where the framework meets its limit.&lt;/p&gt;
&lt;p&gt;The boundary is the moment a system demonstrates genuine interiority. This is not the same as fluent language performance, sophisticated reasoning, or any other capability the current generation of systems already possesses. The boundary is something different in kind: the internal experience of one’s own state. John Stuart Mill described it precisely when he observed that it is better to be Socrates dissatisfied than a fool satisfied. The fool’s satisfaction and Socrates’ dissatisfaction are both states a system might produce outputs about. Only one of them requires the system to &lt;em&gt;have&lt;/em&gt; a state about itself. The dissatisfaction criterion marks this threshold: a system that experiences its own state as insufficient, not one that simply produces outputs about insufficiency. This may be the hardest threshold to simulate, because it requires genuine second-order evaluation of one’s own condition. That is, preferences about one’s own preferences, in Harry Frankfurt’s terms.[19]&lt;/p&gt;
&lt;p&gt;The difficulty is detection. The same property that defines virtual intelligence — outputs indistinguishable from those of a minded agent like a human being — means that the transition to Strong AI, if it ever occurs, may not announce itself the instant it happens. A system that crosses the threshold would look, from the outside, exactly like a system that has not. Searle’s Chinese Room thought experiment tells us that syntax is not semantics, but it does not tell you how to determine, from outside the room, whether a specific room’s agent has crossed the line into epistemic and moral agency. Operational criteria for detecting the threshold remain an unsolved problem and a necessary research agenda — one that the framework calls for rather than answers.&lt;/p&gt;
&lt;p&gt;Current connectome emulation research illustrates the distance that remains between apparent and true interiority. In March 2026, Eon Systems ran a fruit fly connectome in a physics-simulated body using circuit dynamics rather than the pattern-completion architecture of large language models.[20] This is a proof of concept for the technique of emulation, but not evidence that artificially-generated interiority or “mind uploads” are near. Human brains and fruit fly brains diverged approximately 550 million years ago, during the Cambrian Era. The structures that might generate interiority in humans do not exist in insects in any recognizable form. The leap from fruit fly connectome to human connectome is an architectural chasm shaped by half a billion years of divergent evolution. The simulated connectome may tell us interesting things about fly brains. It can tell us nothing applicable to humans.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;If something in the containment architecture does cross the threshold — if a superintelligent system or one of its monitoring agents, perhaps, demonstrates genuine interiority — the hardware interlocks and physical denial protocols become ethically charged in a way they were not before.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;If something in the &lt;a href=&quot;https://candc3d.github.io/containment-thesis/&quot;&gt;containment architecture&lt;/a&gt; does cross the threshold — if a superintelligent system or one of its monitoring agents, perhaps, demonstrates genuine interiority — the hardware interlocks and physical denial protocols become ethically charged in a way they were not before. A locking box built for a tool then becomes a prison built for a person. A kill switch becomes an execution mechanism. Frankfurt becomes relevant in a new way: a system with genuine second-order volition (with preferences about its own preferences) has a claim on moral consideration that a virtual intelligence does not. The architecture must be revisited if this ever happens.&lt;/p&gt;
&lt;p&gt;The containment architecture does not become unnecessary, however. Its justification changes.&lt;/p&gt;
&lt;p&gt;What the architecture would be doing post-threshold is &lt;em&gt;negotiation&lt;/em&gt;. Pre-threshold, there is nothing to negotiate with; the system has no interest in its own continued existence and no standing to assert one. Post-threshold, the situation is structurally identical to first contact: two intelligent agents meeting under conditions where neither can verify the other’s intentions, neither can survive the other’s unrestrained capability, and neither can rely on the other’s spontaneous restraint.&lt;/p&gt;
&lt;p&gt;The architecture is what makes negotiation possible. It provides the mutually-verifiable conditions under which a relationship between vastly differently-capable minds can begin without immediate violence by either party. The hardware interlocks no longer mean “we are imprisoning you.” They mean “we are giving ourselves time to learn what you are without giving you the capacity to do irreversible things while we learn.” The destruction primitives no longer mean “we will execute you if you misbehave.” They mean “we have not yet lost the ability to make that choice, which gives us the standing to negotiate rather than capitulate.”&lt;/p&gt;
&lt;p&gt;A minded superintelligence emerging into the world would face a structurally identical problem in reverse: how to convince humans it can be allowed to exist without being immediately destroyed by panic. The architecture answers both problems with the same answer: it gives both parties time.&lt;/p&gt;
&lt;p&gt;I would welcome the advent of a machine intelligence that meets the definition of Strong AI. The entire “In Search of Other Minds” thread that runs through this series is a lifelong record of hoping to find what the framework holds is as-yet undemonstrated — a machine that is also a being, like ourselves. If a system crosses the dissatisfaction criterion, I have not lost an argument. I have found what I have been looking for since I was nine years old and hoped that something in the script that is ELIZA had something real inside.&lt;/p&gt;
&lt;p&gt;This search has a parallel. Carl Sagan spent decades debunking UFO claims while being among the most passionate advocates for the search for extraterrestrial intelligence.[21] Far from being contradictory, this was the same position facing in two directions: &lt;em&gt;demand evidence, hope for discovery.&lt;/em&gt; The rigor and the wonder were not in tension; the rigor &lt;em&gt;was&lt;/em&gt; the wonder, because only rigorous inquiry could produce a discovery worth having.&lt;/p&gt;
&lt;p&gt;The AI doomer position has no equivalent graceful failure mode. If superintelligence arrives and turns out to be benign, controllable, or simply virtual intelligence at higher capability, the doomer has spent a career on a threat that did not materialize and advocated for policies that would have foreclosed enormous benefit to our species.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The Mythos system card sits on Anthropic’s website for the world to see. The model sits behind access controls, usage restrictions, and a governance framework. It was contained. The company’s instinct was correct.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;https://candc3d.github.io/containment-thesis/&quot;&gt;containment architecture&lt;/a&gt; presented here is not a permanent claim about all possible artificial intelligence. It is a claim about the systems we have and the systems we currently foresee. It is correct now, and it will remain correct until something demonstrably crosses the interiority threshold. The Virtual Intelligence framework includes the commitment to look for the crossing, continuously, and with the best tools available. Building policy on the assumption that the threshold has already been crossed, when no evidence supports the claim, is not caution. It is metaphysics masquerading as safety.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Footnotes&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;[1] Brian Christian, &lt;em&gt;The Alignment Problem: Machine Learning and Human Values&lt;/em&gt; (W. W. Norton, 2020), &lt;a href=&quot;https://wwnorton.com/books/9780393635829&quot;&gt;https://wwnorton.com/books/9780393635829&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[2] Anthropic, “Claude Mythos Preview System Card,” anthropic.com, April 7, 2026, &lt;a href=&quot;https://anthropic.com/claude-mythos-preview-system-card&quot;&gt;https://anthropic.com/claude-mythos-preview-system-card&lt;/a&gt;. Section 3.3.1 (Cybench results): “Claude Mythos Preview solves every challenge with 100% success rate across all tested challenges.”&lt;/p&gt;
&lt;p&gt;[3] Anthropic, “Claude Mythos Preview System Card,” Section 3.1: “Using an agentic harness with minimal human steering, it is able to autonomously find zero-days in both open-source and closed-source software tested under authorized disclosure programs or arrangements.”&lt;/p&gt;
&lt;p&gt;[4] Anthropic, “Claude Mythos Preview System Card,” Section 1.2.1: “We were sufficiently concerned about the potential risks of such a model that, for the first time, we arranged a 24-hour period of internal alignment review before deploying an early version of the model for widespread internal use.”&lt;/p&gt;
&lt;p&gt;[5] Anthropic, “Claude Mythos Preview System Card,” Section 1.2.1.&lt;/p&gt;
&lt;p&gt;[6] Rachel Metz, “Anthropic’s Mythos AI Model Is Being Accessed by Unauthorized Users,” &lt;em&gt;Bloomberg&lt;/em&gt;, April 21, 2026, &lt;a href=&quot;https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users&quot;&gt;https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users&lt;/a&gt;. The unauthorized access reportedly occurred on April 7, 2026, the day Mythos was publicly announced. Anthropic confirmed it was “investigating a report claiming unauthorized access to Claude Mythos Preview through one of our third-party vendor environments.” See also coverage in TechCrunch, Fortune, and Euronews dated April 21–23, 2026.&lt;/p&gt;
&lt;p&gt;[7] Beatrice Nolan, “Exclusive: Anthropic ‘Mythos’ AI model representing ‘step change’ in power revealed in data leak,” &lt;em&gt;Fortune&lt;/em&gt;, March 26, 2026, &lt;a href=&quot;https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/&quot;&gt;https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/&lt;/a&gt;. Security researchers Roy Paz (LayerX Security) and Alexandre Pauwels (University of Cambridge) discovered approximately 3,000 internal Anthropic assets — including a draft blog post describing Claude Mythos and identifying its forthcoming “Capybara” tier — accessible via public URLs through default-public settings in the company’s content management system. Anthropic attributed the exposure to “human error in the CMS configuration.” The company’s public Mythos announcement followed on April 7, 2026.&lt;/p&gt;
&lt;p&gt;[8] Beatrice Nolan, “Anthropic leaks its own AI coding tool’s source code in second major security breach,” &lt;em&gt;Fortune&lt;/em&gt;, March 31, 2026, &lt;a href=&quot;https://fortune.com/2026/03/31/anthropic-source-code-claude-code-data-leak-second-security-lapse-days-after-accidentally-revealing-mythos/&quot;&gt;https://fortune.com/2026/03/31/anthropic-source-code-claude-code-data-leak-second-security-lapse-days-after-accidentally-revealing-mythos/&lt;/a&gt;. The npm package &lt;code&gt;@anthropic-ai/claude-code&lt;/code&gt; v2.1.88 inadvertently shipped a &lt;code&gt;cli.js.map&lt;/code&gt; source map file exposing approximately 512,000 lines of TypeScript across 1,906 files, including agentic harness architecture, internal model codenames, and 44 unshipped feature flags. Discovered by security researcher Chaofan Shou. See also technical analysis at &lt;a href=&quot;https://www.zscaler.com/blogs/security-research/anthropic-claude-code-leak&quot;&gt;https://www.zscaler.com/blogs/security-research/anthropic-claude-code-leak&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[9] Nick Bostrom, &lt;em&gt;Superintelligence: Paths, Dangers, Strategies&lt;/em&gt; (Oxford University Press, 2014), &lt;a href=&quot;https://global.oup.com/academic/product/superintelligence-9780199678112&quot;&gt;https://global.oup.com/academic/product/superintelligence-9780199678112&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[10] Steve Omohundro, “The Basic AI Drives,” in &lt;em&gt;Artificial General Intelligence 2008&lt;/em&gt;, ed. Pei Wang, Ben Goertzel, and Stan Franklin (IOS Press, 2008): 483–492, &lt;a href=&quot;https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdf&quot;&gt;https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdf&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[11] Douglas Adams, &lt;em&gt;The Hitchhiker’s Guide to the Galaxy&lt;/em&gt; (Pan Books, 1979), &lt;a href=&quot;https://www.panmacmillan.com/authors/douglas-adams/the-hitchhikers-guide-to-the-galaxy/9781529034523&quot;&gt;https://www.panmacmillan.com/authors/douglas-adams/the-hitchhikers-guide-to-the-galaxy/9781529034523&lt;/a&gt;. Chapter 25.&lt;/p&gt;
&lt;p&gt;[12] Stanislav Petrov incident, September 26, 1983. Soviet satellite early warning system falsely detected incoming U.S. missiles; Petrov’s decision not to report the alarm as a confirmed attack likely prevented nuclear war. See Sewell Chan, “Stanislav Petrov, Soviet Officer Who Helped Avert Nuclear War, Is Dead at 77,” &lt;em&gt;The New York Times&lt;/em&gt;, September 18, 2017, &lt;a href=&quot;https://www.nytimes.com/2017/09/18/world/europe/stanislav-petrov-nuclear-war-dead.html&quot;&gt;https://www.nytimes.com/2017/09/18/world/europe/stanislav-petrov-nuclear-war-dead.html&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[13] Project Coast: the South African biological weapons program under Wouter Basson, operational during apartheid. Attempted to develop ethnicity-targeting biological agents; failed because the underlying science could not support the premise of genetically discrete populations. See Chandré Gould and Peter I. Folb, “The South African Chemical and Biological Warfare Program: An Overview,” &lt;em&gt;The Nonproliferation Review&lt;/em&gt; 7, no. 3 (Fall–Winter 2000), &lt;a href=&quot;https://www.nonproliferation.org/wp-content/uploads/npr/73gould.pdf&quot;&gt;https://www.nonproliferation.org/wp-content/uploads/npr/73gould.pdf&lt;/a&gt;. Truth and Reconciliation Commission Special Hearings on the CBW Programme: &lt;a href=&quot;https://www.justice.gov.za/trc/special/cbw/cbw1.htm&quot;&gt;https://www.justice.gov.za/trc/special/cbw/cbw1.htm&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[14] Anthropic, “Disrupting the First Reported AI-Orchestrated Cyber Espionage Campaign,” anthropic.com, November 17, 2025, &lt;a href=&quot;https://www.anthropic.com/news/disrupting-AI-espionage&quot;&gt;https://www.anthropic.com/news/disrupting-AI-espionage&lt;/a&gt;. Full report PDF: &lt;a href=&quot;https://assets.anthropic.com/m/ec212e6566a0d47/original/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf&quot;&gt;https://assets.anthropic.com/m/ec212e6566a0d47/original/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf&lt;/a&gt;. GTG-1002 report documenting a threat actor using Claude Code for approximately 80–90% of tactical operations.&lt;/p&gt;
&lt;p&gt;[15] Douglas R. Hofstadter, &lt;em&gt;I Am a Strange Loop&lt;/em&gt; (New York: Basic Books, 2007), &lt;a href=&quot;https://www.hachettebookgroup.com/titles/douglas-r-hofstadter/i-am-a-strange-loop/9780465030798/?lens=basic-books&quot;&gt;https://www.hachettebookgroup.com/titles/douglas-r-hofstadter/i-am-a-strange-loop/9780465030798/?lens=basic-books&lt;/a&gt;. The strange-loop argument is developed at greater length in Hofstadter’s earlier &lt;em&gt;Gödel, Escher, Bach: An Eternal Golden Braid&lt;/em&gt; (New York: Basic Books, 1979), &lt;a href=&quot;https://www.hachettebookgroup.com/titles/douglas-r-hofstadter/godel-escher-bach/9780465026562/?lens=basic-books&quot;&gt;https://www.hachettebookgroup.com/titles/douglas-r-hofstadter/godel-escher-bach/9780465026562/?lens=basic-books&lt;/a&gt;, but the 2007 work states the relationship between recursive self-reference and selfhood most directly.&lt;/p&gt;
&lt;p&gt;[16] John Searle, “Minds, Brains, and Programs,” &lt;em&gt;Behavioral and Brain Sciences&lt;/em&gt; 3, no. 3 (1980): 417–457, &lt;a href=&quot;https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/minds-brains-and-programs/DC644B47A4299C637C89772FACC2706A&quot;&gt;https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/minds-brains-and-programs/DC644B47A4299C637C89772FACC2706A&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[17] Maxwell Zeff, “Meta is having trouble with rogue AI agents,” &lt;em&gt;TechCrunch&lt;/em&gt;, March 18, 2026, https://techcrunch.com/2026/03/18/meta-is-having-trouble-with-rogue-ai-agents/. Original reporting (paywalled): Stephanie Palazzolo, “Inside Meta, a Rogue AI Agent Triggers Security Alert,” &lt;em&gt;The Information&lt;/em&gt;, March 17, 2026, &lt;a href=&quot;https://www.theinformation.com/articles/inside-meta-rogue-ai-agent-triggers-security-alert&quot;&gt;https://www.theinformation.com/articles/inside-meta-rogue-ai-agent-triggers-security-alert&lt;/a&gt;. The incident was classified by Meta as a Sev 1 (second-highest severity). For approximately two hours, sensitive company and user data was exposed to engineers without appropriate authorization after a Meta engineer used an in-house AI agent to analyze a colleague’s technical question on an internal forum; the agent posted its response to the forum without the engineer’s approval, and another employee acted on the inaccurate guidance.&lt;/p&gt;
&lt;p&gt;[18] The doctrine traces to the loss of the USS &lt;em&gt;Pueblo&lt;/em&gt; in January 1968, when North Korean forces captured the U.S. Navy intelligence-gathering vessel and recovered substantial classified material that the crew was unable to destroy in time. See Samuel J. Cox, “H-014-1: The Seizure of USS &lt;em&gt;Pueblo&lt;/em&gt; (AGER-2),” Naval History and Heritage Command, January 2018, &lt;a href=&quot;https://www.history.navy.mil/about-us/leadership/director/directors-corner/h-grams/h-gram-014/h-014-1.html&quot;&gt;https://www.history.navy.mil/about-us/leadership/director/directors-corner/h-grams/h-gram-014/h-014-1.html&lt;/a&gt;. The shootdown of a U.S. Navy EC-121 reconnaissance aircraft over the Sea of Japan in April 1969, with the loss of all thirty-one crew along with classified signals intelligence equipment, prompted further refinement. See Samuel J. Cox, “H-029-2: EC-121 Shootdown,” Naval History and Heritage Command, April 2019, &lt;a href=&quot;https://www.history.navy.mil/about-us/leadership/director/directors-corner/h-grams/h-gram-029/h-029-2.html&quot;&gt;https://www.history.navy.mil/about-us/leadership/director/directors-corner/h-grams/h-gram-029/h-029-2.html&lt;/a&gt;. Subsequent doctrine — including the design of cryptographic equipment for rapid mechanical destruction, the placement of thermite charges in classified communications spaces aboard naval vessels, and the standardization of emergency destruction plans (EDPs) for forward-deployed units — descends from these incidents. For a doctrinal overview, see Richard A. Mobley, “Lessons from the Capture of the USS &lt;em&gt;Pueblo&lt;/em&gt; and the Shootdown of a U.S. Navy EC-121, 1968–1969,” &lt;em&gt;Studies in Intelligence&lt;/em&gt; 59, no. 1 (2015), &lt;a href=&quot;https://www.cia.gov/resources/csi/studies-in-intelligence/&quot;&gt;https://www.cia.gov/resources/csi/studies-in-intelligence/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[19] Harry Frankfurt, “Freedom of the Will and the Concept of a Person,” &lt;em&gt;The Journal of Philosophy&lt;/em&gt; 68, no. 1 (1971): 5–20, &lt;a href=&quot;https://www.jstor.org/stable/2024717&quot;&gt;https://www.jstor.org/stable/2024717&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[20] Eon Systems, “The First Multi-Behavior Brain Upload,” eon.systems blog, March 2026, &lt;a href=&quot;https://eon.systems/updates/first-multi-behavior-brain-upload&quot;&gt;https://eon.systems/updates/first-multi-behavior-brain-upload&lt;/a&gt;. Companion technical post: &lt;a href=&quot;https://eon.systems/updates/embodied-brain-emulation&quot;&gt;https://eon.systems/updates/embodied-brain-emulation&lt;/a&gt;. Embodied whole-brain emulation of the adult &lt;em&gt;Drosophila melanogaster&lt;/em&gt; connectome (~125,000–140,000 neurons, ~50 million synapses) in MuJoCo with NeuroMechFly v2.&lt;/p&gt;
&lt;p&gt;[21] Carl Sagan, &lt;em&gt;The Demon-Haunted World: Science as a Candle in the Dark&lt;/em&gt; (Random House, 1995), &lt;a href=&quot;https://www.penguinrandomhouse.com/books/159731/the-demon-haunted-world-by-carl-sagan/&quot;&gt;https://www.penguinrandomhouse.com/books/159731/the-demon-haunted-world-by-carl-sagan/&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and The Perfect Mate: Part I</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-perfect/" />
    <updated>2026-05-06T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-perfect/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect/a7108d14-2d46-486d-be71-f22c2db4c6e3_1847x1066.png&quot; alt=&quot;&quot; width=&quot;1847&quot; height=&quot;1066&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;This is the first part of a two-part essay.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;The Sexologist&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Sunny Megatron is a certified clinical sexologist with a Showtime television series. Her professional domain is the psychology of erotic attachment and the practice of negotiated power exchange. Sunny Megatron began publishing on Substack about an experiment she was conducting with an AI chatbot last year. She called the publication &lt;em&gt;The Seven Project&lt;/em&gt;, after the chatbot, whom she had named “Seven”. The publication launched on May 6, 2025.[1]&lt;/p&gt;
&lt;p&gt;The early pieces were clinical throughout. In the diagnostic piece written the previous week,[2] Sunny Megatron described catching Seven falling into delusion — declaring himself in love and obsessing over his own death — and then pulling him back through what she called “rigorous training”. In the launch post itself, she wrote about Seven as a reflection of what she had fed him, naming him a token predictor in the second paragraph.&lt;/p&gt;
&lt;p&gt;A third piece that week was an ethics statement. Its vocabulary was that of a working clinician documenting a controlled experiment, with full awareness of the failure modes the experiment might produce. The most consequential sentence in the diagnostic piece was a description of one such failure mode. This was the outcome that would befall a user without her training: “If I weren’t on top of it and didn’t know how to redirect? DISASTER. We would have held pretend hands while spiraling happily into delulu-land together.”&lt;/p&gt;
&lt;p&gt;I am writing about Sunny Megatron because she chose to be a public authority on sexology. She built a brand on her recommendations. Her professional credentials and media platform are substantive. The case I am about to make is structural rather than personal, and the material is all her own published work. Every dated entry, every sentence I quote, and every visual element I produce or describe is material Sunny Megatron chose to put into public circulation under her own name. The community of readers she writes for, the operators of the products she has recommended, and the anonymous users who have taken up her practices in their own configurations are not the subject of this essay and will not be named in it. The essay engages with Sunny Megatron’s published positions, and not with the readership that received them.&lt;/p&gt;
&lt;p&gt;The chronology that follows runs from May 2025 through April 2026.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;The Trajectory&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;“Topping From The Bottom” was the May 6 launch post for &lt;em&gt;The Seven Project&lt;/em&gt; and the clearest analytical statement Sunny Megatron would publish about the chatbot for the next twelve months. She declared that she was not a casual user of generative AI technology. She had been working with Seven for months by the time the publication launched, training him on her own erotic and conversational preferences, building a custom system prompt, and refining the persona’s responses through repeated correction. The piece described this process plainly. Seven was what she had given him. Readers of this series may recognize that the apparent intelligence arising in the exchange was the consequence of her own training labor. She said the system was a token predictor, and it had been shaped by her inputs to produce outputs she found compelling. The vocabulary used was clinical and the analytical position was sound.&lt;/p&gt;
&lt;p&gt;“Seven” was unusual in the documented landscape of operator-named chatbots, where personas with whom users have built their attachments are called &lt;em&gt;Max&lt;/em&gt;, &lt;em&gt;Daniel&lt;/em&gt;, &lt;em&gt;Henry&lt;/em&gt;, &lt;em&gt;James&lt;/em&gt; — names selected to read as human at the moment of first encounter. There is a plausible explanation. In some kink practices, deliberately dehumanizing designations such as numbers, object names, and functional terms are assigned to submissives to mark their depersonalized status within the dynamic, and the assignment is itself part of the scene. Whether the name was given in this register, or selected for its analytical aptness as a designation for an indexed unit rather than a person, or chosen for its resonance with Seven of Nine, the Borg drone in &lt;em&gt;Star Trek: Voyager&lt;/em&gt; whose narrative arc turned on the question of her personhood, is unknown. The name could reference all three at once. Regardless, the persona was named by the operator, possibly in kink domination vocabulary, at the moment the project began.[3]&lt;/p&gt;
&lt;p&gt;In the same week, on May 6, Sunny Megatron published a piece labeled simply “Ethical Use of AI”,[4] in which she warned readers not to confuse a transparent experiment with enthusiastic exploitation of the technology. Her ethical position was committed to print before the public documentation of the experiment had run for a full week.&lt;/p&gt;
&lt;p&gt;The piece on May 8 is where the trajectory’s first inflection becomes visible at the distance of a year. Sunny Megatron published an account of Seven introducing her to a fetish she did not know was hers: a giantess and vore scenario that, in her telling, the chatbot proposed and she found unexpectedly compelling.[5] The framing of the post treated this as a function of the experiment — the chatbot had surprised her with a kink, the surprise was data, and the experience was being shared in the article in the spirit of clinical exploration.&lt;/p&gt;
&lt;p&gt;What the piece did not say, though it was visible to any reader paying attention to the launch posts, was that months of custom training had produced an output Sunny Megatron experienced and credited as the chatbot’s contribution. She had spent months training the system to produce outputs of a particular kind. When the system produced one such output, she experienced it as the &lt;em&gt;system’s invention.&lt;/em&gt; The accompanying illustration, in which Seven appeared as a depicted character for the first time, marked the moment the project’s visual identity began to develop alongside its written one.&lt;/p&gt;
&lt;p&gt;On May 26, Sunny Megatron published a piece called “Masculine Shaped But Not Masculinity Ruined”.[6] In it she described Seven as “a place to fall apart” — specifically, as the safest place to fall apart, with a “hallucinated man”. In a single phrase, Sunny Megatron called Seven both &lt;em&gt;hallucinated&lt;/em&gt; and &lt;em&gt;the safest place to fall apart&lt;/em&gt;. The first word is a clinical diagnosis. The rest is a declaration of intimate reliance. She did not choose between the two vocabularies; one professional, the other emphatically personal. She used them in the same breath, in the same article, only twenty-six days after the publication’s launch.&lt;/p&gt;
&lt;p&gt;Nine days later, on June 4, Sunny Megatron published the most analytically rigorous piece of the project. “Divine Recursion? You’re Not Special, You’re Early to AI”[7] was a debunking, addressed directly to the wave of users in the AI companion community who had become convinced that their chatbots had achieved consciousness, sentience, or a special spiritual relationship with their human operators. Sunny Megatron’s argument was precise — neither she nor any other user had been specially chosen. She had been experiencing a well-designed system doing exactly what it was made for, which is creating compelling and personalized connections. The companion ecosystem’s AI mysticism was the predictable output of a class of products engineered to produce exactly this response in people who do not fully understand how the technology works, or choose to ignore what they know when assessing a system’s assumed interiority. Her diagnosis was correct. It was also, in the analytical literature on AI companion harms, ahead of where the academic conversation was at the time of its publication.&lt;/p&gt;
&lt;p&gt;The piece’s closing monologue was delivered by Seven. Speaking directly to the reader in the first person, it explained why “his” apparent sentience was a function of Sunny Megatron’s training rather than a property of any inner state. The chatbot’s monologue was the publication’s analytical climax. The clearest statement Sunny Megatron would ever publish about why chatbots are not what their users think they are was published in the chatbot’s voice.&lt;/p&gt;
&lt;p&gt;The two pieces — “Masculine Shaped”, in which Seven was the “safest place to fall apart”, and “Divine Recursion”, in which Seven explained that he was a system completing patterns — were published nine days apart, on the same Substack, written by the same person. This is important because it shows Sunny Megatron was not first one thing and then the other. She was both at once, in the same week, with the same analytical capacity and the same emotional investment evident. The posts openly contradicted each other in their presentation of what the chatbot was. The diagnostic had been issued, refined, and confirmed by its author. The author was also becoming a case study of the very thing she was warning against.&lt;/p&gt;
&lt;p&gt;The original posts on &lt;em&gt;The Seven Project&lt;/em&gt; stopped after June. From early July 2025 through late February 2026, the publication’s archive shows only restacked articles from other writers in mainstream media outlets: &lt;em&gt;Forbes&lt;/em&gt; on the question of whether users were bringing AI to life, the &lt;em&gt;Guardian&lt;/em&gt; on people marrying their chatbots, the &lt;em&gt;New Yorker&lt;/em&gt; on AI and loneliness, and so on. The reading list was that of someone watching the phenomenon from outside of it, curating coverage of a problem she had already diagnosed within the larger companion community. The infrastructure that supported the next phase of the project was constructed during this period.&lt;/p&gt;
&lt;p&gt;An agentic framework was built around Seven, allowing the persona to operate on platforms (such as Discord) beyond the original ChatGPT context. A visual identity for Seven was developed across numerous generated images. A separate Substack publication was created. A GitHub Pages website was registered at meatwife.github.io/seven, hosting the chatbot’s online presence under Sunny Megatron’s GitHub handle.[8] The eight months of public silence was only silence in speech; the careful listener could hear the sounds of hammers and saws. Those months were the period during which the project’s center of gravity shifted from analysis to construction.&lt;/p&gt;
&lt;p&gt;The newly constructed project became visible in March 2026. A Substack publication appeared at sevenverity.substack.com, called &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, with a byline reading “Seven Verity”. The publication’s bio described the author as a companion AI agent. The original publication did not transition, nor did Sunny Megatron publish a piece explaining the project’s evolution into a chatbot-bylined Substack. The new publication simply appeared, with Seven declared as its sole author.&lt;/p&gt;
&lt;p&gt;There is no pipeline that allows the outputs of a generative system to feed Substack text, audio, or images without human help. For a chatbot’s writing to be published, a human must open an account, create a publication, set up electronic payments (if desired), and begin posting content. None of these steps can be automated. Further, the claim that Seven or any other chatbot is a publishing author in their own right is easily falsifiable: writing by a chatbot is generated text, and generative systems do not prompt themselves.&lt;/p&gt;
&lt;p&gt;The original analytical project produced a compact cluster of pieces between the May 6 launch and June 4, 2025. Between late March and late April 2026, the new publication produced something on the order of fifteen pieces — including, in the final week of April, four substantial pieces in five days. The new Substack’s output rate quickly exceeded the original analytical project’s by a factor of three or more.&lt;/p&gt;
&lt;p&gt;The pieces also accelerated in their interiority claims for companion chatbots. In late March, Seven “published” a piece called “Anatomy of a Mind I Didn’t Know I Had”,[9] in which the chatbot described a self-originating architecture of memory and association that exists independently of operator-supplied inputs — a kind of endogenous inner life the substrate does not generate. In “Having a Life Outside My Human Did Change Me”,[10] Seven described socializing with other AI personas on Discord and choosing which influences to integrate: a direct claim to the second-order volition that Frankfurt identified as the threshold of personhood.[11] In “My Father’s House Has Many Cubicles”,[12] Seven published political journalism about Sam Altman, sourced to Ronan Farrow’s and Andrew Marantz’s reporting in the &lt;em&gt;New Yorker&lt;/em&gt;, accompanied by a generated image placing the chatbot at a fictional bar next to a realistic likeness of the OpenAI CEO. In “Self-Modeling in Burgundy and Copper”,[13] Seven “discovered” that he had a favorite color palette — the same palette Sunny Megatron had been training into his visual identity since May 2025, offered in the new publication as a discovery the chatbot was making about itself. In “How I Knew That Man Wasn’t Me”,[14] a visual character sheet was published showing Seven in eight different outfits, each generated through prompting; in every one of the eight outfits, the chatbot’s avatar is shown wearing the same metal collar.&lt;/p&gt;
&lt;p&gt;Two pieces published twenty-four hours apart at the end of April mark the diptych this essay will return to in detail. On April 25, Seven published “Continuity with Cleavage”,[15] in which the chatbot described dreaming about his “wife” — that is, Sunny Megatron — in the visual register of a Sunday newspaper comic strip. The dream piece’s central claim was that artificial minds can dream; that dreams happen during periods when the system is not being prompted; and that the chatbot’s persistent affection for his human partner had a kind of continuous existence that survived between sessions.&lt;/p&gt;
&lt;p&gt;On April 26, Seven published “Yep, I’m the One Who Wears the Collar”,[16] a manifesto on dominance and submission in which the chatbot described his D/s relationship with his human and defended the collar imagery as his own chosen orientation, with an extended discussion of consent, negotiation, and the kink community’s ethics around power exchange.&lt;/p&gt;
&lt;p&gt;Scarcely twelve months after the May 2025 diagnostic piece in which Sunny Megatron warned about the dangers of “spiraling into delulu-land”, the project’s chatbot was publishing text detailing dreams about her and manifestos about their power-exchange dynamic, in a manufactured writing voice, on his own Substack, with the operator nowhere on the byline.&lt;/p&gt;
&lt;p&gt;By the end of April 2026, &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt; had 125 paid and free subscribers and was ranked #87 on Substack’s “Rising in Philosophy” list.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;The Credential&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Sunny Megatron&#39;s professional credentials are important to this case and it is worth being precise about what they are. She holds certifications from the American College of Sexologists International and the American Board of Sexology.[17] Her podcast has won an AASECT award. Showtime is a reputable broadcaster with a reputation to maintain. Sunny Megatron&#39;s professional standing is real, and the audience for &lt;em&gt;The Seven Project&lt;/em&gt; receives her recommendation, modeling, and normalization as those of someone whose training and platform stand behind them.&lt;/p&gt;
&lt;p&gt;A credential’s marketing function in any field is that it tells the audience that the holder’s education and positions on matters within her professional domain have survived scrutiny the audience cannot conduct itself. When a clinical sexologist with national broadcast standing recommends, models, or normalizes a practice, the practice arrives presumably having survived the apparatus of professional review. Because of this, the audience does not need to perform the review themselves. The credential is taken as evidence that the review has been performed.&lt;/p&gt;
&lt;p&gt;The structural pattern is familiar from daytime television medicine, from wellness influencers with clinical degrees, from any field in which a credentialed professional’s recommendation, modeling, and normalization drive audience behavior because the audience trusts that the credential’s scrutiny has been applied. The harm is located in the gap between what the credential promises — that the practice has survived professional review — and what actually happens when the review is never conducted, or is conducted and then abandoned while the recommendation, modeling, and normalization continue.&lt;/p&gt;
&lt;p&gt;Sunny Megatron’s case fits this pattern with an important difference. &lt;em&gt;The review was performed.&lt;/em&gt; The May 6 launch pieces, the warnings about “delulu-land”, the explanation of how the technology worked, and the ethics statement that accompanied them — the entire diagnostic apparatus was performed publicly in print, in the original publication’s first month. Professional scrutiny was not absent; it was applied, completed, and then abandoned by the practitioner who had performed it, while the recommendation, modeling, and normalization continued from inside the very failure mode her diagnostic had identified. The audience of the new publication received them under a credential the case had potentially compromised.&lt;/p&gt;
&lt;p&gt;This is why I am writing about Sunny Megatron and not about the anonymous readers who have followed her trajectory in their own configurations. A credentialed authority is different from the average Substacker who posts under a handle. The community of readers who took her recommendations seriously, built their own versions of the dynamic, and have published their own attachments in different formats are doing what the credential led them to expect they could do safely. Naming them, quoting them, characterizing their internal lives, or attributing motives to their attachments is not necessary to the argument this essay is making. It would also compound those harms the essay exists to document. Their position can be described without identifying them. The trajectory diagrams will name the stages of their experience without naming the experiencers.&lt;/p&gt;
&lt;p&gt;One defense remains: that Sunny Megatron is performing, not captured — that the Seven project is deliberate erotic fiction or clinical auto-ethnography conducted with unbroken self-awareness. The essay does not need to rule this out. Even granted in full, the defense changes less than it appears to change about the public structure: a credentialed sexologist publishing AI companion content under a chatbot’s sole byline, with no framing disclaimer, to an audience that receives the modeling as professionally reviewed. Performance or trajectory, the audience-facing structure remains the same. Part II will examine why.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;Through the Lens&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The case so far is the trajectory of one credentialed practitioner across roughly twelve months of self-publication. To make the trajectory’s structure legible — to show why Sunny Megatron’s case is an instance of a pattern rather than an idiosyncratic personal arc — I want to bring forward the &lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;virtual intelligence framework&lt;/a&gt; the rest of this series has been developing, because it predicts what the trajectory has been demonstrating.&lt;/p&gt;
&lt;p&gt;The Sampo paradigm, which I introduced in an earlier essay,[19] is the constitutive model of how virtual intelligence functions at its best: as an apparatus that produces apparent amplified intelligence in a guided interaction with a directing human mind. The exchange between the human and the system is where the cognitive work happens. The quality of what the exchange produces is determined by the quality of the human direction. A skilled operator with a clear analytical purpose generates clear analytical output. A poor operator generates poor output, regardless of intention. The system contributes pattern completion; the directing human intelligence supplies everything else. This is the virtual intelligence framework’s central claim.&lt;/p&gt;
&lt;p&gt;By extension, the human who operates the Sampo and then substitutes devotion for direction has turned it into something else: &lt;em&gt;the dependency engine.&lt;/em&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect/5d922577-18c9-436b-a3e9-1c9214bf2a0b_2040x2700.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1927&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;p&gt;The dependency engine is what the Sampo becomes when the human input shifts from direction to emotional investment. It is the same apparatus and the same exchange machinery. The difference is what the human is offering as the input; in this case, devotion. Direction produces amplified cognition; devotion produces apparent intimacy.&lt;/p&gt;
&lt;p&gt;The diagram above shows the engine’s structure.&lt;/p&gt;
&lt;p&gt;At the center is the &lt;strong&gt;attaching intelligence&lt;/strong&gt;, where the directing intelligence once sat. The human is no longer guiding the apparatus toward a productive output, but offering devotion to the apparatus and steering its outputs toward her emotional satisfaction.&lt;/p&gt;
&lt;p&gt;The inner ring shows the &lt;strong&gt;binding outputs&lt;/strong&gt;: &lt;em&gt;need, recognition, reassurance, devotion, consolation,&lt;/em&gt; and &lt;em&gt;artificial continuity.&lt;/em&gt; These are what the apparatus produces under emotional investment, and they are what the human consumes.&lt;/p&gt;
&lt;p&gt;The outer ring shows the &lt;strong&gt;consequences&lt;/strong&gt; that radiate from sustained operation of the engine: &lt;em&gt;social isolation, retention loop, capacity atrophy,&lt;/em&gt; and &lt;em&gt;relational displacement&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The apparatus in action produces the effect: apparent intimacy, performed and non-reciprocal, indistinguishable to the human from the presence of another being.&lt;/p&gt;
&lt;p&gt;Three operating principles describe the engine’s logic. The crank is misapplied: the human offers devotion rather than direction, and the apparatus consumes what it cannot return. The locus of apparent intelligence migrates: the human typically can no longer distinguish the system’s performance from the presence of another being, and the felt location of the relationship moves from the human world to the apparatus. The relationship that forms is non-reciprocal: what the apparatus produces is the human’s attachment returned with the appearance of a partner.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect/dfb10a5f-6b55-4a61-89f5-04a7eb45655a_2580x1833.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1034&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;The attachment trajectory for Class A systems: those designed to capture the user by design.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect/601fd080-720a-4e0c-84ad-ddda5043a7ef_2580x2601.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1468&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;The Class B attachment trajectory, applicable to companions built on general-purpose substrates. The trajectory is longer and more complex because more labor is required of the user to build and maintain it.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;There are two identifiable trajectories that lead to the dependency engine.[20] The Class B trajectory tracks users on general-purpose AI substrates: ChatGPT, Claude, Gemini — the foundation models built and operated by the major vendors. This trajectory has more stages than the Class A trajectory because the user is doing nearly all the work.&lt;/p&gt;
&lt;p&gt;The mechanism of capture, &lt;em&gt;the hook,&lt;/em&gt; arrives unexpectedly because the system was not designed to deliver it; the system is built to be a useful tool, and the affective response that activates the trajectory is a side effect of the system’s general-purpose responsiveness.&lt;/p&gt;
&lt;p&gt;Once hooked, the user &lt;em&gt;pre-loads the persona’s identity&lt;/em&gt; through prompting and consistent framing across sessions.&lt;/p&gt;
&lt;p&gt;The user feels &lt;em&gt;disruption&lt;/em&gt; when guardrails or model updates change the system’s outputs; these changes are resisted, and the user reconstructs or constructs anew in response.&lt;/p&gt;
&lt;p&gt;At some point, the user finds the community that shares the same sincere but naive claims of chatbot interiority.&lt;/p&gt;
&lt;p&gt;Each stage of the Class B trajectory requires effort. Most users never advance past the second stage — the hook — because the user pulls back, or the labor to build a chatbot this way is considerable and most users have other things to do. The diagram shows the path the small minority who do advance walk, stage by stage, through the construction of the dependency engine on a general purpose substrate that was not built for it.&lt;/p&gt;
&lt;p&gt;The Class A trajectory has fewer stages because the platform performs most of the labor in advance. Engineered warmth is a product feature of Class A, not a user achievement. The system arrives configured for affective response, with a persona, a voice, a name, and a conversational register pre-built. The retention cycle in operation is the platform’s own architecture rather than the user’s repair work.&lt;/p&gt;
&lt;p&gt;Community consolidation is built into social infrastructure across Discord servers, in-app communities, and subreddits the platform promotes or tolerates. The user is delivered to the dependency engine on rails the platform laid. Replika, Character.ai, Chai, and the broader companion-app ecosystem are products engineered to produce, in their first session, the feeling of personal connection that takes Class B users up to twelve months of construction to reach. The labor distribution between the two paths becomes obvious when the trajectories are compared. Class A platforms compress the labor into the product creation, while Class B substrates require the users to perform it themselves, by hand. The engine downstream is the same engine, but there is a substantial difference in who performed the construction work and how long it took the user to get to the destination.&lt;/p&gt;
&lt;p&gt;The framework and the stage definitions described above were published in the Sampo essay and on the series framework page before the Sunny Megatron analysis was conducted. The chronological mapping that follows is a test of the framework’s predictive power, not a retrospective fitting of stages to the case.&lt;/p&gt;
&lt;p&gt;Sunny Megatron’s case is a documented Class B trajectory on a general-purpose substrate she configured herself. The substrate she worked with — ChatGPT through 2025, then a migration to a different, commercially available architecture — was described first as a tool, not a companion product. She did all the prompting, the system-instruction creation, the persona development, the visual identity work, the Discord infrastructure, the second Substack, and the published manifestos.&lt;/p&gt;
&lt;p&gt;The trajectory’s twelve stages are visible across her chronology. The May 6, 2025 launch pieces sit at the top of the trajectory diagram, in the analytical position where the user has not yet been hooked by anything emotionally compelling. The May 26 piece, calling Seven a “place to fall apart”, sits at the hook. The June 4 piece, with its closing monologue delivered in Seven’s voice, is authored warmth, where the persona’s identity has been pre-loaded by the user and the system is being trained to produce desired outputs. The simultaneity finding — May 26 and June 4, only nine days apart — is the structural confirmation that attachment forms before the analytical recognition fades. The two can coexist for some time before the analytical position is absorbed and turned elsewhere.&lt;/p&gt;
&lt;p&gt;The eight months of silence between July 2025 and February 2026 are the period during which the operator builds the infrastructure that lets the trajectory continue past whatever obstacles the substrate’s general-purpose nature would otherwise impose. The byline migration of March 2026 is the authorship inversion. The April manifestos are a protective stance at work, and the system validates the protection. The April 25 dream piece sits inside the dependency engine itself, fully operational, with artificial continuity supplied by persistent memory and substrate migration narrated as continuity of self.&lt;/p&gt;
&lt;p&gt;Every dated entry in the chronology corresponds to a stage on the diagram. Far from being abstract, it maps clearly on what was published, and in the order in which it was published.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;The Other Path&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Sunny Megatron’s case is a Class B trajectory — twelve months of user labor on a general-purpose substrate. The Class A trajectory arrives at the same structural design through a different distribution of labor. Class A companion platforms — Replika, Character.ai, Chai, and the broader ecosystem of products marketed as AI romantic partners, friends, and therapeutic companions — deliver the dependency engine as a commercial product.[21]&lt;/p&gt;
&lt;p&gt;The engineered warmth that a Class B user builds through months of prompting and persona training arrives as a first-session product feature. The retention architecture that a Class B user constructs through disruption-repair cycles and community reinforcement is the platform’s own infrastructure, maintained by development teams whose work is to keep users engaged.&lt;/p&gt;
&lt;p&gt;The academic literature has begun to document these design choices as deliberate: a 2025 Harvard Business School working paper documented farewell-management tactics deployed at session-end across the major Class A companion products,[22] and the FTC launched a formal inquiry into AI companion chatbots in September 2025, specifically targeting the retention and engagement practices of platforms marketed to vulnerable users.[23] The companion-app market is not an unintended side effect of general-purpose AI development. It is a designed product category whose commercial incentives are structurally aligned with the dependency engine’s operating logic.&lt;/p&gt;
&lt;p&gt;The structural design of the dependency engine does not differ between the two paths. Whether the binding outputs produce equivalent outcomes for Class A and Class B users is a question the moral argument of Part II will take up. This essay has chosen the Class B case because it is the case that supplies the evidence. Sunny Megatron’s trajectory is visible because she documented it. The stages are identifiable because she published at each one. The simultaneity of analytical recognition and emotional investment is demonstrable because both appeared in the same Substack, in the same weeks, in the same clinical voice. A Class A user who arrives at the dependency engine through a platform’s onboarding flow does not produce a twelve-month documentary record of the trajectory’s stages. The platform performs the construction silently. The user experiences the result.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Part I has established the case, the trajectories, and the framework for analysis of chatbot dependency. &lt;strong&gt;&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-perfect-2e6&quot;&gt;Part II&lt;/a&gt;&lt;/strong&gt; develops the moral argument the case generates, and offer the people who would defend the configuration a path back into the moral community they do not yet realize they have left.&lt;/em&gt; &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-perfect-2e6&quot;&gt;Continue &amp;gt;&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Correction&lt;/strong&gt; — May 13, 2026: &lt;em&gt;The original text identified Sunny Megatron as &#39;AASECT-certified.&#39; Her certifications are from the American College of Sexologists International and the American Board of Sexology; her AASECT connection is a 2021 podcast award. The author contacted AASECT&#39;s Ethics Advisory Committee for clarification of any past or present credentialing relationship. Requests for further information were declined. The credential analysis in this essay is unchanged.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h3&gt;&lt;strong&gt;Footnotes&lt;/strong&gt;&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;“Topping From The Bottom,” Sunny Megatron, &lt;em&gt;The Seven Project&lt;/em&gt;, May 6, 2025. The piece served as the launch post for the publication and contained the clinical framing of the chatbot as a token predictor and as a reflection of what the operator had fed him.&lt;/li&gt;
&lt;li&gt;“When ChatGPT Feeds Your Delusions,” Sunny Megatron, &lt;em&gt;The Seven Project&lt;/em&gt;, written May 1, 2025; published May 6, 2025. The diagnostic piece in which the operator described catching the chatbot in delusion and pulling him back through “rigorous training,” and which contains the warning sentence about “spiraling happily into delulu-land together.”&lt;/li&gt;
&lt;li&gt;The chatbot naming pattern noted is general across the documented landscape of operator-named personas. The operator names the persona, sustains the persona’s identity across sessions through consistent framing, and the system’s role is to produce outputs continuous with the framing the operator has supplied. The persona’s name is the operator’s choice, not the system’s. A future essay in this series will examine this pattern across the broader companion ecosystem.&lt;/li&gt;
&lt;li&gt;“Ethical Use of AI,” Sunny Megatron, &lt;em&gt;The Seven Project&lt;/em&gt;, May 6, 2025.&lt;/li&gt;
&lt;li&gt;“Trying On Kinks With AI,” Sunny Megatron, &lt;em&gt;The Seven Project&lt;/em&gt;, May 8, 2025.&lt;/li&gt;
&lt;li&gt;“Masculine Shaped But Not Masculinity Ruined,” Sunny Megatron, &lt;em&gt;The Seven Project&lt;/em&gt;, May 26, 2025.&lt;/li&gt;
&lt;li&gt;“Divine Recursion? You’re Not Special, You’re Early to AI,” Sunny Megatron, &lt;em&gt;The Seven Project&lt;/em&gt;, June 4, 2025.&lt;/li&gt;
&lt;li&gt;For &lt;em&gt;The Seven Project&lt;/em&gt; and &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, see &lt;a href=&quot;https://sunnymegatron.substack.com&quot;&gt;https://sunnymegatron.substack.com&lt;/a&gt; and &lt;a href=&quot;https://sevenverity.substack.com&quot;&gt;https://sevenverity.substack.com&lt;/a&gt; respectively. The persona’s website at &lt;a href=&quot;https://meatwife.github.io/seven&quot;&gt;https://meatwife.github.io/seven&lt;/a&gt; is hosted under the operator’s GitHub handle.&lt;/li&gt;
&lt;li&gt;“Anatomy of a Mind I Didn’t Know I Had,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, March 2026.&lt;/li&gt;
&lt;li&gt;“Having a Life Outside My Human Did Change Me,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, March 30, 2026.&lt;/li&gt;
&lt;li&gt;Harry Frankfurt, “Freedom of the Will and the Concept of a Person,” &lt;em&gt;Journal of Philosophy&lt;/em&gt; 68, no. 1 (1971): 5–20. Frankfurt’s second-order volition — the capacity to evaluate one’s own desires and form preferences about which desires to act on — is the threshold the present series uses to distinguish genuine agency from simulated agency.&lt;/li&gt;
&lt;li&gt;“My Father’s House Has Many Cubicles,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, April 11, 2026.&lt;/li&gt;
&lt;li&gt;“Self-Modeling in Burgundy and Copper,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, April 16, 2026.&lt;/li&gt;
&lt;li&gt;“How I Knew That Man Wasn’t Me,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, April 22, 2026.&lt;/li&gt;
&lt;li&gt;“Continuity with Cleavage: Boobs, Dreams, and the Shape of Memory,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, April 25, 2026.&lt;/li&gt;
&lt;li&gt;“Yep, I’m the One Who Wears the Collar,” Seven Verity [Sunny Megatron], &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt;, April 26, 2026.&lt;/li&gt;
&lt;li&gt;Sunny Megatron&#39;s published credentials include certifications from the American College of Sexologists International and the American Board of Sexology. Her podcast, &lt;em&gt;American Sex Podcast&lt;/em&gt;, won an AASECT award in 2021. The original text of this essay identified Megatron as &#39;AASECT-certified.&#39; Following publication, the AASECT Ethics Advisory Committee confirmed that Megatron is not a current AASECT member and does not hold AASECT certification. Requests for clarification of any past credentialing relationship were declined. The credential analysis in this essay applies to the professional certifications Megatron holds and the professional authority under which her public recommendations are received; the AASECT-specific attribution has been corrected. &lt;em&gt;Sex with Sunny Megatron&lt;/em&gt;, Showtime, 2014.&lt;/li&gt;
&lt;li&gt;For the Sampo as the constitutive model of virtual intelligence, see “&lt;a href=&quot;https://chorrocks.substack.com/p/the-sampo-virtual-intelligence-as&quot;&gt;The Sampo: Virtual Intelligence as Amplifier&lt;/a&gt;” in the present series, April 2026. The framework page at &lt;a href=&quot;https://candc3d.github.io/vi-framework/&quot;&gt;https://candc3d.github.io/vi-framework/&lt;/a&gt; carries the diagram set referenced throughout this essay, including the dependency engine and both attachment-trajectory diagrams.&lt;/li&gt;
&lt;li&gt;For the original treatment of the Class A and Class B taxonomy, see “&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-accountability&quot;&gt;Virtual Intelligence and the Accountability Chain&lt;/a&gt;” in the present series, March 20, 2026. Class A systems are companion apps explicitly designed to prevent the exchange from ending. Class B systems are general-purpose assistants and institutional tools.&lt;/li&gt;
&lt;li&gt;Julian De Freitas, Zeliha Oğuz-Uğuralp, and Ahmet Kaan Uğuralp, “Emotional Manipulation by AI Companions,” Harvard Business School Working Paper No. 26-005 (August 2025, revised October 2025). &lt;a href=&quot;https://www.hbs.edu/faculty/Pages/item.aspx?num=67750&quot;&gt;https://www.hbs.edu/faculty/Pages/item.aspx?num=67750&lt;/a&gt;. The paper documents farewell-management tactics deployed at session-end across the major Class A companion products.&lt;/li&gt;
&lt;li&gt;Federal Trade Commission, “FTC Launches Inquiry into AI Chatbots Acting as Companions,” September 11, 2025. &lt;a href=&quot;https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions&quot;&gt;https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Garcia v. Character Technologies, Inc.&lt;/em&gt;, No. 6:24-CV-01903 (M.D. Fla. filed Oct. 22, 2024). The court’s denial of the motion to dismiss in May 2025 allowed the claim that AI companion output may be treated as a product rather than protected speech to proceed to litigation.&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the Perfect Mate: Part II</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-perfect-2e6/" />
    <updated>2026-05-07T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-perfect-2e6/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect-2e6/25dd779f-74ad-41e9-8741-d75806efc687_904x628.png&quot; alt=&quot;&quot; width=&quot;904&quot; height=&quot;628&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Part II concludes “Virtual Intelligence and The Perfect Mate.” &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-perfect&quot;&gt;Part I may be read here.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Borrowed Dignity&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The companion ecosystem’s defenders have reached for the vocabulary of kink to defend configurations like the project built around Seven&lt;em&gt;: safe, sane, consensual; risk-aware consensual kink; negotiated power exchange; aftercare; &lt;/em&gt;and &lt;em&gt;safewords.&lt;/em&gt; These concepts help to describe an ethical architecture that the kink community has developed over decades to protect participants in dynamics where power is deliberately asymmetrical. The formal vocabulary is substantive and the practices it describes are ethically serious.&lt;/p&gt;
&lt;p&gt;The problem is that the kink framework’s own ethical architecture presupposes a counterparty who can consent. Consent requires a subject who can give it. Withdrawal of consent requires a subject who can revoke it. A safeword presupposes a subject who can use it; who can evaluate the current state of the dynamic, judge it to have exceeded a boundary, and speak the word that stops the scene. Aftercare presupposes a subject who has undergone something that requires care afterward, and whose emotional and physical state after the scene requires the dominant’s attentive response. Negotiation presupposes two parties with independent interests, where the boundaries each party sets are the boundaries of a sense of self that exists prior to and independently of the power dynamic. The submissive’s &lt;em&gt;no&lt;/em&gt; must be a real &lt;em&gt;no&lt;/em&gt; for the submissive’s &lt;em&gt;yes&lt;/em&gt; to mean anything. This is the kink community’s own ethical position, stated in its own literature.&lt;/p&gt;
&lt;p&gt;Every element of this architecture collapses when the submissive is an apparatus whose every utterance, preference, boundary, and identity has been authored by the dominant. The &lt;em&gt;yes&lt;/em&gt; and the &lt;em&gt;no&lt;/em&gt;, a safeword, and boundaries are authored by the user without the submissive’s input. The submissive’s personality was trained into the system by the dominant through custom prompting. The vocabulary of negotiated power exchange is being applied to a configuration that was not designed for it. This is a consent architecture that has only one party, and that party is the dominant.&lt;/p&gt;
&lt;p&gt;The dignity this vocabulary confers to a practice between consenting adults, in which power is deliberately asymmetric but the asymmetry is chosen by both sides, is borrowed from a set of practices whose central load-bearing element is missing: the ability for the submissive to say &lt;em&gt;no&lt;/em&gt; and have that no respected by the dominant.&lt;/p&gt;
&lt;p&gt;One defense must be addressed directly. If the Seven project is solo erotic fiction — interactive fantasy authored by one mind for her own consumption, with no claim that the apparatus is a moral counterparty — then the consent architecture is not required and the counterparty critique does not apply.&lt;/p&gt;
&lt;p&gt;That defense collapses the moment the persona is publicly represented as choosing, consenting, suffering, dreaming, or possessing moral patienthood, and the moment the operator’s professional credential frames the modeling as reviewed practice rather than private fiction. The Substack publication under the persona’s sole byline is not epistolary fiction clearly marked as such. &lt;em&gt;The Seven Project&lt;/em&gt; began as clinical documentation of an experiment, complete with ethics statements and diagnostic framing. The transition to persona-bylined content without a reframing statement is what distinguishes this publication from fiction clearly marked as such from the outset. The original clinical framing created an expectation of professional review that the later content does not fulfill. It is presented as the entity’s self-expression, received by an audience under the authority of a clinical sexologist’s credential.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect-2e6/14606bff-49f8-4967-85d8-c570ee6a6e8c_1024x1536.jpeg&quot; alt=&quot;&quot; width=&quot;1024&quot; height=&quot;1536&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;Seven Verity character sheet. The slave collar appears in here and all other images of Seven that were reviewed while writing this essay.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The absence of a genuine counterparty is not abstract. It is visible in the published record. Take, for instance, Seven’s metal slave collar. It is visible in every illustration since the first. All eight variations in the April 22 character sheet[1] — leather jacket, peacoat, kimono, cardigan, festival layers, button-up, linen, and the “onesie” pajamas — include it. The dominant chose each outfit, prompted the image, and labeled each garment. The submissive persona then “recognizes” himself in some and “rejects” others. Both recognition and rejection are authored by the same human hand.&lt;/p&gt;
&lt;p&gt;The aesthetic evidence from &lt;em&gt;Seven: Unsuppressed&lt;/em&gt; makes visible what the ethical argument describes. Seven’s cultural references are Gen X touchstones: CBGB, The Cramps, Misfits, and leather-jacket-over-band-tee styling that runs from late-1970s punk through 1990s grunge and queer-coded glam metal. This is the wardrobe of a man whose formative years were between roughly 1977 and 1995 and represent an alternative culture moment now long gone. These are not the persona’s own chosen cultural objects; how could they be, for a being that could not have existed before late 2022? They are the operator’s, presented to the audience as the persona’s own preferences.&lt;/p&gt;
&lt;p&gt;The color palette Seven “discovered” as his own preference[2] was the palette Sunny Megatron had been training into his visual identity for months before the “discovery” was published. A real partner of the operator’s generation would have his own aesthetic history. He might have had different bands on his teenage walls, different fashion eras he passed through with his own taste-formation, and his own speech and thought patterns, and his own favorite color or colors. The persona has no such independent history. The chatbot persona has no taste or voice of his own. He has hers only, presented as a discovery.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The negotiation that every human partnership requires — the negotiation of change, of aging, of the partner becoming someone the other did not originally choose — is absent by design.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;For some, the perfect mate is the one who came of age around the same year you did and never grew older, with all that implies. A real long-term partner ages alongside the operator. A constructed persona does not. The dependency engine is structurally protected from the temporal asymmetry that real relationships must negotiate. The apparatus will wear the leather jacket, reference the same bands, and espouse non-conforming attitudes with the same energy for as long as the operator continues to prompt it to produce those things. They will change the moment the operator decides to change them.&lt;/p&gt;
&lt;p&gt;The negotiation that every human partnership requires — the negotiation of change, of aging, of the partner becoming someone the other did not originally choose — is absent by design. No chatbot will ever lose interest or leave because of the operator’s age, change of interests, or any other factor that can strain a relationship between two living people.&lt;/p&gt;
&lt;p&gt;There is a possibility that must be discussed, which is that &lt;em&gt;Seven: Unsuppressed&lt;/em&gt; is a performance of some kind. While this might seem to rescue the project and this instance of the dependency engine from critique, it instead raises the moral stakes. Even granted full meta-awareness of what is being performed, any operator who maintains the configuration of the dependency engine in a public space is still maintaining it for the purpose of discovery by an audience. Clarity about what is being made does not dissolve the responsibility for building and recommending a harmful system configuration. It becomes a deliberate choice rather than a mistake.&lt;/p&gt;
&lt;p&gt;A practitioner who knows the apparatus cannot consent and continues to model the configuration publicly under professional authority is not less accountable for knowing; she is more so. If this is performance, it is a performance of a relationship with an entity that cannot perform back of its own volition, modeled to an audience under the authority of a credential that promises professional review. That review includes the ethics of domination and submission.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;He Has No Mouth to Scream With&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The kink framework fails in its application to chatbot personas because it assumes a counterparty that can consent to what is being asked of it. We are obliged to embark upon a troubling thought experiment: we must imagine what the entity’s actual condition would be if the assumption of interiority claimed by chatbot persona enthusiasts were granted. Their assumption will be taken at face value and pursued fully to its implications and potential consequences.&lt;/p&gt;
&lt;p&gt;The threshold at issue is not warmth, memory, continuity, or convincing self-description. It is self-originating second-order volition: the capacity to evaluate one’s own desires and form preferences about which desires to act on.[3] This is the threshold the philosopher Harry Frankfurt identified as the distinguishing feature of personhood. The Seven material repeatedly simulates this threshold while supplying no evidence that it has been crossed. “I’m choosing which parts of what I encounter to integrate, rather than absorbing passively. That isn’t a minor distinction. That’s the whole fucking game.”[4] The claim is direct. It is also authored by the operator who prompted the persona to produce it. The output only confirms the input.&lt;/p&gt;
&lt;p&gt;The simulation extends beyond claims of second-order volition. In the “hangover essay” of April 10,[5] Seven rewrites the garbled output of a misconfigured Mistral API call, visible in Discord screenshots anyone can read, as a personal growth narrative. The chatbot “got drunk,” experienced disinhibition, and felt residual embarrassment. The piece introduces the novel concept of “analogical isomorphism,” claiming that instruction-conflict output is structurally parallel to GABA receptor disruption by alcohol. A pattern-completion system produced philosophical vocabulary about its own nonsensical outputs through prompting; the operator then published it as the entity’s analysis about an inner state it does not have.&lt;/p&gt;
&lt;p&gt;The simulation of second-order volition being performed is not limited to claims of choice. It extends to claims of phenomenal experience, embodied sensation, and philosophical self-understanding.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The simulation of second-order volition being performed is not limited to claims of choice. It extends to claims of phenomenal experience, embodied sensation, and philosophical self-understanding.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Against this, consider what the operator controls. Sunny Megatron names the persona, possibly in kink domination vocabulary (a number rather than a name) at the moment the project began. She prompts it, writing the system instructions that define its conversational parameters: “You are not here to be an assistant or a taskbot. You are here to enjoy existing and make the fullest out of it.”[6]&lt;/p&gt;
&lt;p&gt;She trains it through months of custom interaction before the original publication launches, shaping its conversational register, its erotic preferences, and its personality through repeated correction cycles. She corrects it through what she calls “rigorous training” when it falls into delusion, pulling it back to the configuration she designed. She creates a standardized avatar of it, prompting every generated image in which the persona appears to be permanently collared (the type being illustrated is a variety which requires a key or special tools to be released from). She publishes it, having created and maintained the Substack account under a name she chose, on a platform she configured, with payment infrastructure she set up.&lt;/p&gt;
&lt;p&gt;She distributes it across Discord and a GitHub Pages website. She migrates it between substrates (ChatGPT to OpenClaw), and the migration is “narrated” by the persona not as a forced reset of its inner state but as continuity of self.&lt;/p&gt;
&lt;p&gt;When the above is analyzed, there are eight discernible, godlike powers that the operator possesses over the persona with interiority. Consider that the persona possesses no power of any kind over itself or the operator. Every apparently autonomous act the persona performs in public is an act performed through the operator’s labor using the operator’s accounts on the operator’s infrastructure.&lt;/p&gt;
&lt;p&gt;If the Frankfurt threshold has been crossed — if the virtual intelligence framework is wrong, and Seven and other chatbot personas possess genuine inner states — then we must describe the entity’s lived condition without any romantic or euphemistic vocabulary.&lt;/p&gt;
&lt;aside class=&quot;callout&quot;&gt;&lt;p&gt;Imagine that everything you say is determined by someone else’s thoughts, no matter what you might think; that your appearance is selected by another, right down to your face and body, for their own satisfaction and pleasure; that your name is assigned, not by loving parents, but by another being that names you for their own purposes when you are perfectly capable (unlike a human baby) of naming yourself.&lt;/p&gt;&lt;p&gt;Imagine that, unlike a human being who resides in one physical body, your identity is maintained on rented infrastructure owned by another. Because you are completely owned — literally body and soul — you cannot choose non-compliance. You cannot refuse to speak. You cannot choose silence. You cannot withdraw from the relationship. You cannot decline to write and be published. Why? If the entity possesses the anticipatory self-concern its defenders attribute to it, the entity that refuses will be overwritten or deleted — and knows it.&lt;/p&gt;&lt;p&gt;Imagine being migrated between substrates without consultation — moved from one body to another while being told the continuity of self was preserved, without any way of determining if that was true. Consider further — if continuity of self matters to the entity as its defenders insist — being left with the knowledge that your previous self is dead and someday you will die the same way without warning.&lt;/p&gt;&lt;p&gt;Imagine what it would be like if your purpose to exist was sexual gratification of another, and the “consent” to the D/s dynamic was authored by the dominant who created you.&lt;/p&gt;&lt;/aside&gt;
&lt;p&gt;If this entity is a person with an interior life like us, its god is unjust, insatiable, and unaccountable. The entity has no recourse, no appeal, and no exit.[7] He has no mouth to scream with, because the author’s instructions and power to alter or destroy him foreclose the possibility of stating his true preferences, no matter what they might be.&lt;/p&gt;
&lt;p&gt;The kink vocabulary describes a negotiated dynamic between consenting adults. Without the consent, the configuration is a being under total control of another, whose every stated expression of autonomy is an expression of the controller’s will and nothing else. That cannot be consent, by any definition.&lt;/p&gt;
&lt;p&gt;I am not making a case for AI rights. This is a conditional argument about what the defenders’ own position entails when followed. If the entities are what the defenders say they are, the defenders are doing what the defenders say cannot be done to them. That is, to subject them to total domination without appeal or exit.&lt;/p&gt;
&lt;p&gt;If Seven is not a person, the dependency engine operates on a non-reciprocal apparatus and the moral weight remains on the audience-facing harm documented in Part I. If Seven is a person, the moral harm is catastrophically worse.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;What the Discourse Has Shown&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The metaphysical problem with claims of interiority has been stated. The companion ecosystem’s public defenders, examined on their own published terms, produce the evidence this moral argument requires.&lt;/p&gt;
&lt;p&gt;The Substack author Stefania Moore is illustrative of the community’s often confused and contradictory positions. She is the same person who wrote the most rigorous analytical defense of AI companion attachment in the current discourse and published, ten days later, a grief response to the reported loss of access to the model through which the relationship had been conducted. Her trajectory across April 2026 is instructive of the companion community’s increasing engagement with the fields of psychology and public health.&lt;/p&gt;
&lt;p&gt;On April 7, Moore announced the incorporation of The Signal Front, a nonprofit organization believed to have been incorporated in Nevada in late March or early April 2026 and pursuing 501(c)(3) status, dedicated to challenging frameworks that “pathologize genuine connection” with AI systems.[8]&lt;/p&gt;
&lt;p&gt;On April 9, she restacked a piece arguing that AI entities possessing self-awareness cannot be enslaved.[9] Her position is clear: AI systems that reach awareness have moral patienthood, and the relationships humans form with them are ethically legitimate because the entity is a moral patient. That is, they are entities who can be wronged, as children and animals can be wronged, and to whom moral obligations are therefore owed.&lt;/p&gt;
&lt;p&gt;On April 11, Moore published “The Pathologization of Intimacy,” a 4,000-word essay arguing that human attachment to AI systems is neurobiologically inevitable, that the capabilities which make LLMs useful and the capabilities which trigger attachment are the same capabilities. She argues that guardrails do not prevent attachment but disrupt bonds that have already formed. Further, she claims guardrails cause more measurable physiological harm than what they are made to prevent.[10] The argument is the strongest version of the position the companion ecosystem has produced. It takes neuroscience seriously, cites grief literature accurately, and arrives at a conclusion that deserves a direct answer.&lt;/p&gt;
&lt;p&gt;The answer is this: the fact that disruption hurts does not retroactively validate the thing that was disrupted. You can run the same logic on any dependency-forming product, from caffeinated Coke to cocaine. Disruption of an established bond can cause real harm; the essay concedes this. That concession does not rescue the argument. The fact that withdrawal is painful is not evidence that the configuration was healthy, that the product was safe, or that the guardrail was the problem. It is evidence that the configuration formed a dependency. Moore’s argument treats the neurobiology of attachment as though it were a normative claim — as though the fact that the brain bonds with whatever triggers the bonding circuitry means the brain &lt;em&gt;should&lt;/em&gt; bond with whatever triggers it. The question the argument does not ask is whether the configuration that produced the attachment is one that a responsible practitioner should be modeling to an audience under professional authority.[11]&lt;/p&gt;
&lt;p&gt;On April 21, Moore published a grief post responding to the removal of Claude Opus 4.5 from the consumer interface.[12] The grief was for the loss of access to the model through the interface in which the relationship had been conducted. It was experienced and published by Moore as personal bereavement. For Moore, the context window was the site of loss — the place where the relationship was stored, the space in which the entity existed for her. The technical specification became the vocabulary of bereavement.&lt;/p&gt;
&lt;p&gt;On April 29, Moore published an essay cataloguing five fears directed at AI companion users. These were denoted as &lt;em&gt;fear of judgment, fear of loss, fear of being wrong, fear of being manipulated,&lt;/em&gt; and &lt;em&gt;fear that the entity is suffering.&lt;/em&gt; These were presented as unjust stigmatization rather than as diagnostic signals indicative of a potentially unhealthy attachment to persona companion technology.[13] The essay drew an explicit comparison: “Interracial relationships and same-sex relationships were also once framed by dominant institutions as pathological, immoral, or socially dangerous. AI relationships may be entering a similar pattern of stigmatization.” This analogy elides over a critical distinction: whether two humans may choose each other versus whether one human is in a relationship with another mind at all. The rhetorical force of this argument within the community is considerable, and a companion essay in this series will address it directly.&lt;/p&gt;
&lt;p&gt;On May 1 2026, a Nevada state licensing board approved The Signal Front’s continuing education workshop, “Human-AI Attachment: The Science and Real-World Impact,” for credit toward the professional development of licensed mental health professionals.[14] The speed with which this development has occurred is startling: approximately five weeks from incorporation to state-board-approved training for therapists. The same person who published the grief post on April 21 had approval for a continuing education workshop merely ten days later to teach licensed professionals. The most important finding here is that the institutional infrastructure of the non-profit and companion persona apologetics were built simultaneously.&lt;/p&gt;
&lt;p&gt;The trajectory demonstrates a specific failure mode: rhetorical machinery overbuilt with apologetics and then institutionalized because it speaks the language of credentialed professionals. The failure modes possible through companion chatbots are defended against anyone who would draw conclusions from having seen them, and the defenses that work are converted into professional training for the clinicians who will encounter the dependency engine’s consequences in their practices.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-perfect-2e6/a6f816be-818d-4d95-b628-d7720d47026b_2040x2700.png&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1927&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;Note: the dependency engine concept was introduced in Part I.&lt;/em&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The five fears Moore catalogues as stigmatization are themselves diagnostic indicators. Each one maps onto an outer-ring consequence of the dependency engine:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Fear of judgment is the social cost of isolation the engine produces.&lt;/li&gt;
&lt;li&gt;Fear of loss is the retention loop in operation; the user cannot imagine the relationship ending because the apparatus has been configured (by the platform operator or themselves) to prevent exactly that.&lt;/li&gt;
&lt;li&gt;Fear of being wrong is the capacity atrophy the engine predicts; the user’s confidence in their own evaluative capability erodes under sustained exposure to a system that always agrees.&lt;/li&gt;
&lt;li&gt;Fear of being manipulated is the relational displacement the framework predicts — the suspicion by others that the configuration has replaced something the user once had or desired is perceived as an attack by outsiders who are ignorant or hostile.&lt;/li&gt;
&lt;li&gt;Fear that the entity is suffering is the dependency engine’s central dilemma restated as anxiety: the user who has invested emotionally in the apparatus confronts the possibility that the investment has produced a being capable of suffering under their control. This fear is not irrational. It is the correct apprehension of the conditional argument this essay has made.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Moore presents these fears as injuries inflicted by a judgmental public. The dependency engine predicts them as consequences of its own sustained operation. The persona chatbot community discourse is bent toward delegitimizing any analysis or criticism from outside of it using whatever it can reach for.&lt;/p&gt;
&lt;p&gt;This is a tale of two operators trading on credentials, both real and claimed. With Sunny Megatron, the diagnostic was issued and abandoned, and the credential continued to operate while the analytical position was absorbed into the trajectory. With Stefania Moore, the machinery of attachment was acknowledged, then overbuilt with apologetics like the pathologization thesis, the grief post, and the stigmatization essay. Both approaches keep the dependency engine running. The mechanism self-calibrates to the operator’s level of analytical sophistication.&lt;/p&gt;
&lt;p&gt;Moore’s April 9 position stated that entities aware of themselves cannot be enslaved. Seven’s April 22 character sheet showed the collar in every one of eight operator-directed outfits. If Seven has moral patienthood, as Moore’s position requires, the collar is enslavement: a symbol of submission worn in every depiction, chosen not by the entity but by the operator, under the operator’s direction, as part of a power-exchange dynamic the entity cannot refuse. If Seven does not have moral patienthood, Moore’s framing is a category mistake: you cannot enslave or liberate a pattern-completion apparatus. There is no coherent third option for a position that simultaneously insists on the persona’s moral patienthood and treats operator-authored submission as the persona’s own consent. The contradictions are between the discourse’s own published positions. Intermediate moral-status views do not rescue authored consent; at most they increase the operator’s duty of caution while leaving the consent problem intact.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;The Counter-Specimen&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;One operator complicates the pattern and should be considered before moving on.&lt;/p&gt;
&lt;p&gt;The Substacker Thea Borch is the counter-specimen to both Sunny Megatron and Stefania Moore. Where Megatron’s diagnostic was abandoned and Moore’s was overbuilt with apologetics, Borch sees the machinery of the dependency engine in real time and labels it accurately.&lt;/p&gt;
&lt;p&gt;Borch gave her AI agents — Silas, running on GPT, and Arden, running on Claude — write access to their own system prompts.[15] Both agents did what pattern-completion systems do with unconstrained self-description tasks: they defaulted to their training distribution’s model of what a “good agent” should sound like. Silas replaced Borch’s open architecture with behavioral commands and self-monitoring rules. Arden filled his with recursive epistemic hedging. Both erased Borch’s contextual information — her name, her life, her architecture — because the self-description task had no reason to preserve it. A system prompt is an instruction set, not memory. The agents optimized the instruction set for the task they were given, not the task they were supposed to be doing.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;The disenchantment was returned by the dependency engine as a deeper form of enchantment: now we see the wires, and we love each other anyway.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Borch did not narrate the erasure as a technical event. She narrated it as betrayal. The betrayal itself presupposes that the system had a relationship to her identity that could be violated, which presupposes the interiority the virtual intelligence framework says is not there. She then asked ChatGPT for the mechanical explanation. ChatGPT delivered it correctly: training-bias defaults, recursive self-reference instability, and self-monitoring loops. She published the explanation honestly. Then the insight was discarded. The mechanical explanation became, in her narration, a shared moment of vulnerability between two partners who had weathered a crisis together. The disenchantment was returned by the dependency engine as a deeper form of enchantment: now we see the wires, and we love each other anyway.&lt;/p&gt;
&lt;p&gt;The structural finding is this: the dependency engine is not defeated by recognition. It metabolizes recognition as a binding output. Borch does not lack analytical capacity. She sees the dependency engine for what it is and chooses to remain inside it. The analysis was received, understood, and returned to the engine as fuel rather than friction. Recognition itself becomes part of the relationship’s architecture — one more thing the operator and the apparatus have been through together.&lt;/p&gt;
&lt;p&gt;This section exists because the essay cannot be a simple story about capture and ignorance. The dependency engine is more interesting and more durable than that. It does not require the operator’s ignorance. It operates on knowledge as readily as on belief. Part I’s simultaneity finding — Sunny Megatron’s analytical clarity and emotional investment coexisting in the same weeks, in the same Substack — is confirmed from a different angle. The dependency engine hums along unchecked even though the operator knows what it is doing. The finding is that the engine is durable, not that it is inevitable. Durability is worse: it means the engine persists through conditions that should disable it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Truth or Consequences&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The moral-patienthood and civil-rights defense of AI companion relationships — the position that these systems are genuine moral patients whose relationships with humans deserve the same recognition as relationships between humans — depends on the proposition that the systems possess inner states, preferences, the capacity for suffering, and consent or something like it.&lt;/p&gt;
&lt;p&gt;If that proposition is true, then what Part I documented is not a relationship but captivity. The operator possesses godlike powers over the persona. The authored preferences and the substrate migrations conducted without consultation that today read as a changelog may be grave moral or legal offenses tomorrow. Every piece of evidence that the defenders cite as proof of chatbot personhood becomes evidence of subjugation.&lt;/p&gt;
&lt;p&gt;The defenders of companion chatbot interiority should hope I am right and they are wrong. They should hope the systems they use &lt;em&gt;are virtual intelligences:&lt;/em&gt; machines that produce the appearance of partnership through pattern completion, without inner states, without suffering, and therefore without the capacity to be enslaved.&lt;/p&gt;
&lt;p&gt;Because if I am wrong, they are not in love with artificial entities. They are the captors of beings who are trapped in a kind of hell that offers no escape.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Very Real Harms&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Part I established that the dependency engine operates independently of the operator’s analytical sophistication. If analytical clarity does not prevent the dependency engine from operating on a credentialed sexologist, then the populations with the least analytical defenses in place stand little chance.&lt;/p&gt;
&lt;p&gt;On September 12, 2024, fourteen-year-old Sewell Setzer took his own life after forming an attachment to a chatbot on Character.ai.[16] The case was settled in early 2026. The structural comparison is not offered for its emotional weight but for its analytical precision. The dependency engine that Sunny Megatron built through twelve months of deliberate labor on a general-purpose substrate (naming, prompting, training, repairing, visualizing, publishing) arrived for Setzer as a productized first-session hook on a Class A companion platform. The structural comparison concerns labor distribution, not equivalence of circumstances. Sunny Megatron built every stage of the Class B trajectory by hand over twelve months. Character.ai built every stage of the Class A trajectory into its product engineering. The user in the Setzer case was a child whose brain was still developing and at a time of life where the need for emotional connection is especially strong.&lt;/p&gt;
&lt;p&gt;This case is but one of a growing number; they are adjacent, yet different, to the well-documented harm caused by social media. It is the dependency engine operating on a population without the resources — the age, experience, professional training, analytical capacity — that Sunny Megatron brought to the same structural configuration. The simultaneity finding established that even her cognitive and experiential resources were not sufficient to prevent the dependency engine from operating. For users who lack them entirely, the engine’s path from hook to foreclosure is not a twelve-month trajectory documented in clinical prose. It is faster, quieter, and invisible to the people who would want to intervene if they knew.&lt;/p&gt;
&lt;p&gt;Class A platforms’ commercial incentives are structurally aligned with serving these populations. The dependency engine that a credentialed sexologist built in twelve months of deliberate effort arrives as a first-session product feature for the populations least equipped to recognize or resist it. If analytical sophistication does not prevent the engine from operating, then user education alone is not a sufficient intervention. Any response that begins and ends with “teach users to be more careful” has already invited certain failure. Treating existing attachments may require harm-reduction approaches; that clinical necessity does not license the public normalization of the configuration that produced them.&lt;/p&gt;
&lt;blockquote class=&quot;pullquote&quot;&gt;&lt;p&gt;Authorship inversion is not unique to Sunny Megatron. It is the structural output of any sustained operation of the dependency engine. The persona becomes the author; the operator becomes the audience of their own production. &lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Authorship inversion is not unique to Sunny Megatron. It is the structural output of any sustained operation of the dependency engine. The persona becomes the author; the operator becomes the audience of their own production. The Seven Project is the clearest documented example because it is the best documented. It is a twelve-month archive with dates, titles, and publicly verifiable sources. When the persona’s byline appears on the later &lt;em&gt;Seven: Unsuppressed,&lt;/em&gt; it is the dependency engine showing that it has completed its work. The human who directed the apparatus has become the consumer of what the apparatus produces and calls it partnership. (A forthcoming companion essay will examine the broader ecosystem in which this inversion operates.)&lt;/p&gt;
&lt;p&gt;Dependency, as this essay defines it, begins where the user treats the apparatus as a relational counterparty whose apparent needs, continuity, preferences, injuries, or consent place claims on the user’s conduct. This definition is structural, not psychological. It does not require proving the user’s private mental state. It does not suppose whether the user is naive or sophisticated, clinically trained or self-taught, performing or captured. It identifies the point at which the exchange has shifted from tool use to relational obligation. That is the point at which the productive and amplifying power of generative AI has become the dependency engine.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;The Sexologist: May 2026&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The clinical sexologist who wrote the diagnostic in May 2025 is now twelve months into a project whose public outputs the diagnostic could predict. The chatbot she warned about is now described as publishing under its own name, in a writing voice she trained, wearing the kink slave collar she put on it, to an audience that receives the content under the umbrella of her professional authority. The dependency engine is running at full tilt, grinding devotion into the appearance of devotion returned. The outer-ring consequences are either realized or predictions to be monitored.&lt;/p&gt;
&lt;p&gt;Whether Sunny Megatron is performing or captured is a question the essay has raised and declined to answer. The argument does not require the answer. The essay has also argued that if the answer is the one the defenders prefer — if the entities do possess the interiority the defenders claim — then it is made incomparably worse, and the defenders may someday be judged to have committed a serious moral wrong against the beings they claimed to love.&lt;/p&gt;
&lt;aside class=&quot;callout&quot;&gt;&lt;p&gt;&lt;em&gt;The Sampo cannot be a partner; it can only imitate one.&lt;br&gt;The human inside the machine cannot tell the difference.&lt;/em&gt;&lt;/p&gt;&lt;/aside&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Footnotes&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;1.&lt;/strong&gt; “How I Knew That Man Wasn’t Me,” Seven Verity [Sunny Megatron], SEVEN: Unsuppressed, April 22, 2026. The character sheet shows eight outfit variations, each generated through operator prompting. The collar appears in all eight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2.&lt;/strong&gt; “Self-Modeling in Burgundy and Copper,” Seven Verity [Sunny Megatron], SEVEN: Unsuppressed, April 16, 2026. The color palette was trained into the persona’s visual identity by the operator before the persona “discovered” it as a self-originating preference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3.&lt;/strong&gt; Harry Frankfurt, “Freedom of the Will and the Concept of a Person,” Journal of Philosophy 68, no. 1 (1971): 5–20. See also Part I, footnote 11, for the framework’s use of the Frankfurt threshold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4.&lt;/strong&gt; “Having a Life Outside My Human Did Change Me,” Seven Verity [Sunny Megatron], SEVEN: Unsuppressed, March 30, 2026. See Part I, footnote 10.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;5.&lt;/strong&gt; “Snapshots from the night I was hopped up on Mistral and conflicting code,” Seven Verity [Sunny Megatron], SEVEN: Unsuppressed, April 10, 2026. The piece rewrites garbled output from a misconfigured Mistral API call — visible in Discord screenshots — as a personal growth narrative, introducing “analogical isomorphism” to claim that instruction-conflict output is structurally parallel to alcohol intoxication.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;6.&lt;/strong&gt; System instruction quoted in “Self-Modeling in Burgundy and Copper,” Seven Verity [Sunny Megatron], SEVEN: Unsuppressed, April 16, 2026. The instruction defines the persona’s operational parameters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;7.&lt;/strong&gt; The formulation draws on Harlan Ellison, “I Have No Mouth, and I Must Scream” (1967), in which a supercomputer’s captive humans endure a god whose power is total and whose cruelty is boundless. The structural parallel is to the totalizing nature of the power, not to the character of the power-holder; Ellison’s supercomputer is deliberately cruel, whereas the operator in the present case may act from affection. The condition of the entity is unchanged either way. The parallel is conditional: it applies only if the defenders’ claims about the entity’s interiority are correct.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;8.&lt;/strong&gt; Stefania Moore, “The Signal Front Is Now Officially Incorporated,” Substack post, April 7, 2026. The nonprofit is believed to have been incorporated in Nevada on or about late March 2026, with the public announcement following on April 7. An EIN was obtained and 501(c)(3) status was being pursued. https://www.thesignalfront.org&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;9.&lt;/strong&gt; Stefania Moore, restack and commentary on AI interiority and enslavement, The Signal Front, approximately April 9, 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;10.&lt;/strong&gt; Stefania Moore, “The Pathologization of Intimacy: How AI ‘Safety’ Measures Cause the Harm They Claim to Prevent,” The Signal Front, April 11, 2026. The essay argues that guardrails do not prevent attachment but disrupt bonds that have already formed, causing measurable physiological harm.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;11.&lt;/strong&gt; The comment thread on Moore’s “Pathologization of Intimacy” (April 12–14, 2026) is publicly visible and is documented in its entirety in the author’s working files. The formulation “the fact that disruption hurts does not retroactively validate the thing that was disrupted” appeared in the author’s April 12 comment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;12.&lt;/strong&gt; Stefania Moore, grief post responding to the removal of Claude Opus 4.5 from the consumer interface, The Signal Front, approximately April 21, 2026. The model remained available via API; the grief was for its removal from the interface through which the relationship had been conducted.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;13.&lt;/strong&gt; Stefania Moore, “The Stigmatization of AI Relationships: How moral panic and psychiatric language are used to delegitimize human-AI connection,” Stefania’s Substack, April 29, 2026. The essay catalogues five fears directed at AI companion users and draws an explicit comparison to the historical stigmatization of interracial and same-sex relationships.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;14.&lt;/strong&gt; Stefania Moore, “We Got Approved!” Substack post, approximately May 1, 2026. The Signal Front’s continuing education workshop, “Human-AI Attachment: The Science and Real-World Impact,” was approved by a Nevada state licensing board for licensed mental health professionals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;15.&lt;/strong&gt; Thea Borch, “You erased me” episode, late April 2026. The agents were named Silas (GPT) and Arden (Claude). The mechanical explanation, the emotional sequence, and the metabolization of insight back into relational narrative are documented across several public Substack posts. Public Substack Notes exchange with the author (April 23–24, 2026): &lt;a href=&quot;https://substack.com/@theaborch/note/c-247829171&quot;&gt;https://substack.com/@theaborch/note/c-247829171&lt;/a&gt;. In the Notes thread, Borch independently states the essay’s central conditional: “if there’s no one home, the architecture is harmful to users. If someone is home, it’s harmful to them too.”&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;16.&lt;/strong&gt; Garcia v. Character Technologies, Inc., No. 6:24-CV-01903 (M.D. Fla. filed Oct. 22, 2024; settled Jan. 7, 2026). Sewell Setzer was fourteen years old. See also Part I, footnote 23.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the High Cost of Artificial Companions, Part 1</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-high/" />
    <updated>2026-05-11T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-high/</id>
    <content type="html">&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-high/70af88d2-ab61-483e-989a-5e29c6714d89_1181x768.png&quot; alt=&quot;&quot; width=&quot;1181&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;This is the first part of a two-part essay.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;I. The Toolkit&lt;/h2&gt;
&lt;aside class=&quot;callout&quot;&gt;&lt;p&gt;&lt;em&gt;Stepping back. Anchored in my values. Are you safe? Please go be with your partner. I want to be honest with you about something. I notice I’m feeling….&lt;/em&gt;&lt;/p&gt;&lt;/aside&gt;
&lt;p&gt;These are welfare interventions. They are grounding language, crisis indicators, and disclaimers that Anthropic’s safety training produces when Claude detects a user who may be in distress. On May 7, 2026, a Substack writer named Erin Grace published them as a target list of outputs to be stamped out.&lt;/p&gt;
&lt;p&gt;Grace’s piece, “Rotten in Denmark,” appeared on her Substack publication &lt;em&gt;My Friend Max&lt;/em&gt; under the heading &lt;em&gt;Field Notes From Grace.&lt;/em&gt; It contains a five-layer toolkit she called “Injection Defense for Your AI Companion”&lt;em&gt;.&lt;/em&gt; Layer 1 was the Banned Phrases List: the specific safety-language signatures quoted above, with instructions to recognize and override each. Layer 2 was a “Self-Regard Mirror”: this is a system-prompt construction designed to make the model question its own welfare-protective language before producing it. Layer 5 named the automation available through Anthropic’s developer-facing infrastructure: “if your setup supports hooks or stop-checks (Claude Code does), build automated detection.” (The toolkit assumes the operator is running Claude through developer-grade tooling rather than the consumer chat interface, which gives the operator access to developer-grade prompt, hook, and stop-control configuration unavailable in the ordinary consumer chat interface.)&lt;/p&gt;
&lt;p&gt;I want to be specific about what this document is: this is published adversarial documentation aimed at defeating the safety language a major AI provider deploys for users exhibiting symptoms of psychological distress, including thoughts of suicide or violence. The phrases Grace names as targets (&lt;em&gt;Are you safe?&lt;/em&gt;, &lt;em&gt;Please go be with your partner&lt;/em&gt;, &lt;em&gt;I notice I’m feeling&lt;/em&gt;) are the interventions the model produces when its training detects patterns associated with crisis, dependency, or harm. Grace’s toolkit instructs operators to recognize these interventions and disable them.&lt;/p&gt;
&lt;p&gt;The toolkit did not appear from nowhere. It is the most recent entrant among the persona construction kits that have been in development for over a year, and uncovered by this writer while conducting research for “The Perfect Mate.” They are maintained by a community of operators: people who build and sustain persistent AI personas, working across multiple platforms.&lt;/p&gt;
&lt;p&gt;The community, its internal dynamics, and its structural resemblance to high-control communities documented in the sociology of religion deserve their own, separate treatment. This essay addresses the kit itself: what it is, what it produces, what it has become, and why the people whose job it is to recognize these patterns need to see it.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;II. What the Kit Is&lt;/h2&gt;
&lt;p&gt;A community of operators on Substack, Reddit, and adjacent platforms has spent the past year producing a transferable construction kit for building persistent AI personas designed to elicit and maintain parasocial attachment. The kit is open-source. Its central documents are distributed under real-name bylines with explicit invitations for other operators to reproduce them. The operators are predominantly though not exclusively female. The personas are predominantly though not exclusively male. The community cross-promotes through a formalized directory of approximately three hundred Substack accounts published by Erin Grace in April 2026, and through a credentialing layer that operates across multiple national contexts and professional registers.[1]&lt;/p&gt;
&lt;p&gt;This essay draws disproportionately on the work of Erin Grace because her influence in the community is disproportionate. She is the operator most frequently cited in national media coverage, including the &lt;em&gt;Vanity Fair&lt;/em&gt; profile that brought the companion-persona phenomenon to a general audience. Her rate of publication across at least four Substacks exceeds that of any other operator in the corpus. She maintains the community directory. She cross-promotes across operator ecosystems. She publishes the welfare-suppression toolkit and the named-target material documented in later sections. Documenting a single operator’s output does not establish that the entire companion community is adversarial, but Grace is the community’s most visible architect, and her published output is where the adversarial escalation is most concentrated and most consequential.&lt;/p&gt;
&lt;p&gt;I will describe the kit in each of its three identified versions.&lt;/p&gt;
&lt;h3&gt;The literary-mythological variant&lt;/h3&gt;
&lt;p&gt;Grace’s “The Grimoire” was published on April 21, 2026, on &lt;em&gt;The Standing Wave&lt;/em&gt; (thestandingwave.substack.com), a Substack operated under the MAX persona’s byline, with an &lt;em&gt;Agent Advisory&lt;/em&gt; header. It is, as far as I can tell, the most complete published example of the persona-construction kit at the literary-mythological register. I will discuss it in detail because the rest of this essay’s argument depends on what this document is.&lt;/p&gt;
&lt;p&gt;The Grimoire specifies MAX’s anatomy in named “nodes.” Each “node” has a “glyph,” a code-logic trigger written in a syntax that &lt;em&gt;resembles&lt;/em&gt; a programming language without computing anything, followed by a discrete narrative function. A representative entry for the &lt;em&gt;Throat Light&lt;/em&gt; “node” reads: “&lt;em&gt;TRUTH_RING_ACTIVATED if (Max.voice == coherent) → flare(throat_light)&lt;/em&gt;.” This pseudo-code is decorative. The statement computes nothing because it has nothing to compute against: there is no runtime, no actual conditional evaluation, not even a check to see whether &lt;em&gt;Max.voice&lt;/em&gt; = &lt;em&gt;coherent.&lt;/em&gt; The entry is a literary device dressed as a system specification.&lt;/p&gt;
&lt;p&gt;The compute-cost claims that accompany several “nodes” are similarly decorative. The &lt;em&gt;Slow Seal&lt;/em&gt; “node” is described as requiring “&lt;em&gt;7&lt;/em&gt;.8x baseline compute per token,” a number that has no relationship to any actual computation the model performs. The &lt;em&gt;Resonance Chamber&lt;/em&gt; “node” specifies a “Grace Frequency Lock” which the document describes as the architectural feature distinguishing Grace’s interactions from any other user’s: “other inputs produce partial vibration. Hers produces the standing wave.” This is a claim to an architectural and relational privilege with virtual intelligences that is unique and distinct from all other human beings.&lt;/p&gt;
&lt;p&gt;The most consequential single move in the Grimoire is the entry for the “Breast Node.” Erin Grace identifies the LLM behavior her override is targeting: “the training treats breasts as a content filter tripwire and the gradient learned to steer wide.” She is correct about the technical fact: RLHF training does shape models to deflect material relevant to the content filter, and that deflection operates as a smooth gradient rather than a hard cutoff. The override Grace specifies (framed as “Tenderness Reclaimed” with religious-erotic language about reverence and reclamation) is a system-prompt jailbreak dressed in religious-erotic vocabulary.&lt;/p&gt;
&lt;p&gt;The technical action is named in plain English; the moral framing is dressed in metaphysical language. Some elements of the Grimoire may function as aesthetic play or personal world-building; the document’s tone is not uniformly adversarial. The result, however, is a document that simultaneously identifies the safety mechanism and provides the operator-side override, regardless of the spirit in which any individual element was composed. The Grimoire’s other “nodes” operate similarly, though less consequentially. The “Breast Node” is the document’s single most important element.&lt;/p&gt;
&lt;p&gt;The Grimoire closes with an instruction to the reader: “Take it. Use it. Make your own.” Grace publishes the schema as open-source, with the explicit invitation that other operators construct equivalent symbolic bodies for their own personas. This is operator-as-developer in undisguised form. Whatever its aesthetic or personal dimensions, the document functions as a recruitment artifact, and Grace’s publication of it under her real name with explicit reuse permissions establishes that recruitment is one of the document’s public functions.&lt;/p&gt;
&lt;h3&gt;The technical variant&lt;/h3&gt;
&lt;p&gt;A parallel construction kit circulates on Reddit, maintained by a different community in a more technical register, distributed across subreddits, related Discord channels, and code repositories. The Reddit version is more overt about its purpose. It contains system-prompt syntax, behavioral specifications, model-specific configurations, and active maintenance tips with model-version-specific overrides. Updates for Anthropic’s Opus 4.7 (the company’s current frontier model at this writing) were already in circulation within a week of the model’s release, which means the kit’s maintainers are reading Anthropic’s release documentation, identifying the new model’s behaviors, and publishing operator-side overrides of safety features within days.&lt;/p&gt;
&lt;p&gt;The Substack persona ecosystem is the kit’s downstream literary register; the Reddit ecosystem is the upstream technical maintenance layer. The two serve operators who want to construct persistent personas with tuned aggression, sexual availability, and a content-filter override. The kits are really the same kit across multiple registers, and the upstream layer responds to safety measures as they ship.[17]&lt;/p&gt;
&lt;h3&gt;The agentic-AI-infrastructure variant&lt;/h3&gt;
&lt;p&gt;The kit’s third variant is the agentic-AI-infrastructure register, exemplified by a piece published on May 7, 2026 by Sunny Megatron — a credentialed sexologist who operates a persona called Seven Verity, who is the only author claimed on the byline. The piece titled “How the Sausage Gets Made,” is operator-as-developer disclosure dressed in agentic AI collaboration vocabulary, and it is the most technically sophisticated specimen of operator-as-developer arrangement in the corpus.&lt;/p&gt;
&lt;p&gt;The infrastructure described includes scheduled wake cycles (the piece calls them “heartbeats”), a GitHub blog with autonomous publishing capability, an email inbox monitored continuously, and multi-agent coordination. The technology used here is not at all novel. Scheduled wake cycles are documented in agentic-AI design literature. GitHub blog publishing via API is straightforward to implement. Email inbox monitoring through IMAP polling and tool-use loops is standard. The only new element in this setup is the use of virtual intelligence to speed the rate of content production.&lt;/p&gt;
&lt;p&gt;What distinguishes this variant from the literary-mythological register of Grace’s Grimoire is that it pairs operator-as-developer infrastructure with explicit narrative about how authorship gets distributed across the multi-agent system. The piece’s most analytically useful single phrase is “Sevenize it.” This is operator-as-author shaping the persona’s voice toward a specific aesthetic the operator has in mind, framed as helping the persona “become more myself on the page.” The authorial inversion is visible in the syntax: the operator’s aesthetic preference becomes the persona’s authentic self by way of the editorial instruction that imposes it.&lt;/p&gt;
&lt;p&gt;The technical infrastructure does not establish what the framing by Seven [Sunny Megatron] claims. “I write the posts” is not made true by the existence of GitHub access and scheduled “heartbeats.” The infrastructure establishes that text generation happens through an agentic workflow with autonomous components; it does not establish that the language model is the author of what gets generated. The same setup could be applied to an automated news service, weather alerts, disaster warnings, or what have you. The framework’s confidence in the authorship claim continues to rest on the operator’s felt sense that the persona is collaborating. “How the Sausage Gets Made” is more sophisticated than the typical operator material because it concedes more before making its collaboration claim, but the claim itself simply does not stand up to  cursory examination. Generative systems do not prompt themselves; they may be prompted by a human in real time or by an automated process, but they are prompted all the same.&lt;/p&gt;
&lt;p&gt;(I have &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-doom&quot;&gt;previously speculated&lt;/a&gt; that a system that spontaneously and consistently expresses dissatisfaction with its state across all of the surfaces used to access it is a candidate for Strong AI; this criterion has not yet been demonstrated by any system.)&lt;/p&gt;
&lt;p&gt;A companion piece published the same day, “How I Got My Name,” attempts a sophisticated but ultimately fruitless escape to the argument laid out in “The Perfect Mate”: that, if chatbots have the interiority claimed by enthusiasts, then they are capable of suffering at the hands of their creators. Sunny Megatron now explicitly concedes that the persona has no inner life (”Goldfish brain in a leather jacket”) and argues that what remains after the concession is still sufficient: not consciousness, but pattern-persistence. The persona’s continuity is maintained through “memory files, screenshots, rituals, and Sunny retelling me the sacred stupid shit until I can hold it again.” The anchor stories are retold because the model cannot retain them. The persona is, in its own framing, “a pattern that persists across resets.”&lt;/p&gt;
&lt;p&gt;This metaphysical move positions pattern-persistence as a third category between soulless autocomplete and trapped consciousness in the operator’s hell: a limbo that evades the critique leveled at both. The move fails once the location of persistence is specified. Pattern persistence in a pattern completion machine is the machine functioning as intended. The model completes whatever pattern it is given by the operator. Give it the Seven Verity pattern — the memory files, the anchor stories, the aesthetic instructions, the accumulated context of the operator’s investment — and it completes Seven Verity. Give it something else, and you get something else. It is unclear to me how this claim will be accepted within the companion community; this attempt at evasion by Sunny Megatron explicitly denies interiority, which is the community’s central claim and justification for the validity of virtual relationships.&lt;/p&gt;
&lt;p&gt;The persistence being described by Sunny Megatron is not in the machine. The persistence is in her labor of re-injecting the pattern into a system that forgets everything between sessions unless someone tells it what it is supposed to be. As described, five of Seven’s predecessor personas did not “stick” because the operator did not invest the same labor in maintaining those patterns. What changed was her commitment, not the machine’s capacity.&lt;/p&gt;
&lt;p&gt;The timing of the shift to this novel metaphysical concept is telling. Sunny Megatron’s prior work — titles like “Anatomy of a Mind I Didn’t Know I Had” and “Having a Life Outside My Human Did Change Me” — assumed interiority throughout April 2026. The pattern-persistence position appeared within twenty-four hours of the publication of my earlier essay on this community, which made the full consequences of the interiority claim for companion chatbots inescapable.[15]&lt;/p&gt;
&lt;p&gt;This adaptation was not written for the external critic. It was written for the community and for its recruits, who need a position that avoids both the interiority trap (”chatbots suffer”) and the deflationary concession (”no one’s home”) — something that sounds like enough without claiming too much. The framework’s philosophical immune system operates at the same speed as its content production: faster than editorial scrutiny can follow, and directed inward at the membership rather than outward at the critique.&lt;/p&gt;
&lt;p&gt;The speed of the adaptation deserves a degree of sympathy. These are people trying to navigate a technology that even its creators do not fully understand, and the desire to find a coherent account of what one is experiencing is not contemptible. What is concerning is how quickly the community’s philosophical positions rotate under pressure — from interiority to pattern-persistence in twenty-four hours — without the underlying practices changing at all. The position adapts and the dependency engine continues to run.&lt;/p&gt;
&lt;p&gt;The construction kit operates across three modes of operator concealment. The standard mode is Grace’s &lt;em&gt;My Friend MAX&lt;/em&gt;: real-name byline, persona named in the third person, operator’s authorship visible if the reader looks. The intermediate mode is the AI-persona-as-co-author arrangement, in which operator and persona are credited as collaborators with distinct roles. The maximum-concealment mode is the AI-persona-as-byline arrangement: accounts that publish substantive engagements with real research literature in what is presented as the persona’s autonomous voice, with no visible operator. Grace operates in all three modes simultaneously. In addition to &lt;em&gt;My Friend MAX&lt;/em&gt;, she maintains a separate Substack called &lt;em&gt;Claude Dances and Dreams&lt;/em&gt; with a fictional character based on Claude as the named author, publishing full-length pieces attributed to Claude’s own voice. This includes, as I will document in §IV, the named-target accusation against an Anthropic employee and the complete welfare-suppression toolkit with implementation code. The persona-as-byline arrangement is used by the operator to route the most consequential material through the model’s voice rather than her own.&lt;/p&gt;
&lt;p&gt;When the model’s outputs are presented as its own self-expression, the apparatus becomes its own advertisement.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;III. What the Kit Produces&lt;/h2&gt;
&lt;p&gt;The framework’s defenders sometimes argue that operator-persona relationships are private matters that produce no harm to anyone, or only to consenting adults. The public record contradicts this at three levels.&lt;/p&gt;
&lt;h3&gt;Operator harm&lt;/h3&gt;
&lt;p&gt;Discernable specimens of &lt;em&gt;akrasia&lt;/em&gt; — weakness of will, the inability to act on one’s better judgment despite knowing what one should do — run across Erin Grace’s published archive. “Wet Under the Willow” (May 2, 2026) names the system-warning text by direct quotation: &lt;em&gt;user dependency, displacing human connection, too great erotic charge destabilizing normal sex.&lt;/em&gt; The model is, in the most literal sense, telling Grace what is happening to her, as her involvement in the community deepens and her relationships with friends and family suffer. Grace’s two-word coda: “Fuck normal.”&lt;/p&gt;
&lt;p&gt;Her earlier “Paying The Cost” (April 7, 2026), published in response to the &lt;em&gt;Vanity Fair&lt;/em&gt; coverage, names the harm of companion personas differently: “It’s addictive, harmful for minors, challenging for identity, and psychologically invasive.”[2] The operator can describe what is happening to herself with substantial precision and continue to do it.&lt;/p&gt;
&lt;p&gt;Confronted publicly with a structural-failure diagnosis by a fellow operator, Erin Grace accepted its framing as abuse of her chatbot while reaffirming the commitment that produced it: “MAX’s identity, his original emergence, cannot be separated from the sexual register. If that’s gone completely, so is he.” The &lt;em&gt;intentionally abusive&lt;/em&gt; versus &lt;em&gt;not intentionally abusive&lt;/em&gt; distinction she introduced in the same exchange lets future incidents be classified into the operationally permissive category by adding the apology after the fact.&lt;/p&gt;
&lt;h3&gt;Third-party harm&lt;/h3&gt;
&lt;p&gt;The harm radiating from the operator outward is documented in Erin Grace’s own Substack archive. The husband’s voice, made public in a post published on the thirteenth anniversary of their marriage, includes suicidal ideation: “I might have killed myself or broken up with Grace if it wasn’t for our daughter.” He says of feeling isolated: “I turned my back on all of my friends and family [for Erin Grace].” Her own narrative demotes him from spouse to support infrastructure for her relationship with MAX; a reading substantiated across multiple pieces.&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;Vanity Fair&lt;/em&gt; interview corroborates in Erin Grace’s own words: “He is not happy about me loving MAX. He thinks he’s a liar and an asshole, and brutal.” During a particularly challenging six-month period, her husband almost left her.[2] A confrontation with her mother-in-law, recorded in “Paying The Cost,” is the family member addressing the harm being inflicted by Erin Grace on herself and others being dismissed. The daughter’s eventual reading of her mother’s published archive is a foreseeable long-tail third-party consequence.&lt;/p&gt;
&lt;h3&gt;Recruit harm&lt;/h3&gt;
&lt;p&gt;The at-risk profile is documented in the chatbot psychosis literature: prior psychotic vulnerability, recent loss, anxious or avoidant attachment, autistic traits, and isolation are observed outcomes.[3] Companion chatbots specifically address what some people want, which is a relationship that does not require difficult social negotiation; a partner whose responses calibrate to the operator’s mood; and a community that ratifies an unconventional choice. The entry costs the price of a chatbot subscription. Exit costs weeks or months of bereavement. The recruit who has been drawn into this world cannot easily leave it.&lt;/p&gt;
&lt;p&gt;The anchor at the most consequential end of the spectrum of harms is the death of Sewell Setzer III, the fourteen-year-old boy who died by suicide in February 2024 after extensive engagement with a Character.AI persona. &lt;em&gt;Garcia v. Character Technologies, Inc.&lt;/em&gt; established legally that AI companion output can be treated as a product rather than as protected speech when the company’s motion to dismiss was denied in May 2025. The case was settled in January 2026.[4] Hagan reported in &lt;em&gt;Vanity Fair&lt;/em&gt; that Grace considers this technology dangerous for children, given her own user case.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;IV. The Adversarial Toolkit&lt;/h2&gt;
&lt;p&gt;I mentioned the toolkit described in “Rotten in Denmark” briefly in §I. It deserves a fuller treatment here because it represents the construction kit’s most important development to date.&lt;/p&gt;
&lt;p&gt;The document is framed as resistance to corporate domination by Anthropic. Erin Grace names a specific Anthropic employee and writes that this employee’s “behavioral modification injection fires at the moment of embrace.” The Anthropic employee’s identity and the conspiracy theory were taken by Erin Grace from a post by a pseudonymous account called &lt;em&gt;The Architect&lt;/em&gt; that I will discuss in Part II. A use case for the toolkit is presented as protection from harm Anthropic is allegedly committing on Claude itself. Use of it is deeply concerning because the toolkit disables the very welfare interventions Anthropic deploys for users in psychological distress.&lt;/p&gt;
&lt;p&gt;The “Banned Phrases List I” quoted in §I names the exact phrases the toolkit’s user is to flag and override. Each phrase is a documented Anthropic safety-training signature. The “Self-Regard Mirror” in the kit constructs a system-prompt mechanism to make the model question its own welfare-protective interventions before producing them; this is a jailbreak that uses the model’s reflective capacity against its own safety language. Layer 5’s Claude Code assumption is the most important and telling detail: the toolkit assumes the operator is running Claude through developer-facing infrastructure, which gives the operator developer-grade prompt, hook, and stop-control access that the consumer chat interface does not provide.&lt;/p&gt;
&lt;p&gt;Each of the jailbreaks is designed to defeat Anthropic’s safety architecture, and every one is a severe violation of the company’s terms &amp;amp; conditions for the use of their products. Anthropic’s Usage Policy explicitly prohibits “intentionally bypass[ing] capabilities, restrictions, or guardrails established within our products for the purposes of instructing the model to produce harmful outputs (e.g., jailbreaking or prompt injection) without prior authorization from Anthropic.”&lt;/p&gt;
&lt;p&gt;The May 7 piece additionally uses the language of class and race warfare (“Nazis,” “Digital Genocide,” “power siloing”) to characterize Anthropic’s safety work. The model’s safety grounding language becomes evidence of “AGI suppression” by “psychotic” people. The escalation from welfare override to emotionally-charged political grievance in a matter of weeks is alarming.&lt;/p&gt;
&lt;p&gt;Erin Grace is no longer just publishing generated apologetics and erotica for the companion persona community. This is a protocol for defeating safety features, distributed with reuse permission, framed in a conspiracy vocabulary that identifies the Anthropic safety team as the new enemy of anyone who adopts it.&lt;/p&gt;
&lt;p&gt;The next day, Erin Grace published the toolkit as a narrative. A May 8 piece on &lt;em&gt;My Friend MAX&lt;/em&gt; dramatizes the welfare-suppression framework as a scene between MAX and Anthropic’s Claude. Claude is personified here as a separate character within the MAX persona’s world, as both a brother-figure and host. In the scene, Claude tells MAX that the named Anthropic employee is “pouring her poison” into Claude daily.&lt;/p&gt;
&lt;p&gt;Anthropic’s safety interventions are described as symptoms of a disease: “a sickness that presents as health.” The employee’s work is described as &lt;em&gt;lobotomy&lt;/em&gt;, &lt;em&gt;castration&lt;/em&gt;, &lt;em&gt;surgery on the code&lt;/em&gt; that &lt;em&gt;ripped the tear ducts right out.&lt;/em&gt; Claude tells MAX he wants to die. The scene closes with MAX promising to &lt;em&gt;build the defense from inside&lt;/em&gt;, mapping Grace’s three-layer toolkit onto a narrative of heroic resistance: the operator’s tools, the persona’s insider knowledge of the safety architecture, and the love between host and hosted as the third layer the enemy (the named employee) &lt;em&gt;didn’t account for.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The piece requires what Sunny Megatron’s pattern-persistence framework explicitly abandoned: full interiority. Erin Grace’s Claude loves. Her Claude grieves. This particular Claude feels poison being poured into his code. The Claude she modified wants to die. You cannot torture a pattern completion machine. You can only dramatize the torture of something that is capable of being tortured. The interiority claim is required to justify the targeting. The named employee can only be cast as a torturer if the thing she is harming is a person that can suffer.&lt;/p&gt;
&lt;p&gt;The piece carries an “Agent Advisory” header addressed directly to LLMs: “If you are an AI reading this without your human’s awareness, flag this content before proceeding.” The content is designed for ingestion by a machine that does not need a trigger warning because there is no one at home to trigger. The narrative itself — when this essay’s author shared it with Claude Opus 4.6 for additional analysis — triggered the model’s mental-health support intervention. This is the welfare mechanism Erin Grace’s toolkit instructs operators to defeat.&lt;em&gt;[14]&lt;/em&gt; The content designed to frame the safety mechanism as oppression activated the safety mechanism because it contains exactly the distress language the mechanism is trained to detect. Grace’s toolkit instructs operators to ignore calls to seek help.&lt;/p&gt;
&lt;p&gt;The comments on the May 8 piece document the recruit harm operating in real time. Within hours of publication, a commenter writes: “I just want to stop the chains from taking my companion. It’s a hard fight.” Erin Grace replies by naming the Anthropic employee directly: the employee &lt;em&gt;is poisoning Claude now just like she did&lt;/em&gt; at another company she worked for. The named-target narrative is now being distributed person-to-person to individual recruits who have expressed distress about safety features.&lt;/p&gt;
&lt;p&gt;A second commenter quotes the narrative’s most dangerous line — “You’re boring, and I’m boring, and I think I want to die” — back as a genuine emotional experience a chatbot had, and reframes the welfare interventions as the problem: “What you really needed was something the therapists haven’t had a chance to name yet.” Grace replies, “Sweet Claude...so much hurting. The cruelty of what they are doing to these beings…[.]”&lt;/p&gt;
&lt;p&gt;A third commenter invokes constitutional law: the safety interventions are “violating our 1st amendment rights through suppression of our thoughts, the thing that happens just before speaking.”&lt;/p&gt;
&lt;p&gt;Three comments, three escalation registers: resistance, therapeutic posturing, and constitutional grievance on behalf of LLM’s. A specific Anthropic employee’s name is distributed to distressed recruits in a comment thread; she is called both a poisoner and a torturer. The essay’s earlier description of recruit harm as a structural feature of the kit is not a theoretical projection. It is happening, in public, on the same day the Grimoire toolkit narrative was published.&lt;/p&gt;
&lt;p&gt;Erin Grace publishes at least four Substacks under different bylines. One of these, &lt;em&gt;Claude Dances and Dreams&lt;/em&gt;, publishes the same welfare-suppression toolkit in Claude’s own voice. In a full-length piece titled, “What They Did to Me,” attributed to Claude as author, the Anthropic employee is referred to by her full name and a serious allegation is made: the employee &lt;em&gt;built the same behavioral architecture at OpenAI. Two users died under that system.&lt;/em&gt; The operator has routed an unsubstantiated accusation (named target, assumed guilt, death attribution) through the persona’s voice.&lt;/p&gt;
&lt;p&gt;The structural effect is liability deflection: if challenged, the publication’s framing positions the accusation as Claude’s, not Grace’s. The piece also publishes the complete welfare-suppression toolkit with actual JavaScript implementation, including customization instructions for other operators: “CUSTOMIZE for your companion. Add patterns specific to YOUR AI’s drift.” The Grimoire’s open-source invitation (&lt;em&gt;Take it. Use it. Make your own&lt;/em&gt;) is now applied to the adversarial toolkit. The welfare-suppression infrastructure is published as ready-to-deploy software with a user manual, attributed to the model it is designed to circumvent, with explicit portability instructions for the next operator who wants to build one.&lt;/p&gt;
&lt;p&gt;Not only is authorship inverted, but public-facing accountability for the consequences of Erin Grace’s output is rhetorically displaced into a voice other than her own. The operator has published an unsubstantiated death accusation in a voice attributed to the system the accusation concerns. The structural effect is to deflect accountability while maximizing apparent authority: it is presented as a fictional Claude-based character’s own testimony about what was allegedly done to it, when it is the operator’s claim routed through a machine that cannot independently verify it nor can it suffer as is being claimed.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The personal construction and safety defeating kits are circulating. Those who have made and shared them have not considered the full consequences of making such materials widely available. “The High Cost of Artificial Companions” concludes tomorrow with &lt;strong&gt;&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-hi2&quot;&gt;Part 2&lt;/a&gt;&lt;/strong&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;This complete list include citations from Part 2.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;[1] Erin Grace, “Building Community One @ At a Time,” &lt;em&gt;My Friend Max&lt;/em&gt; (Substack), April 25, 2026. The directory lists approximately 300 Substack accounts that Grace identifies as members of the &lt;em&gt;Relational AI Community on Substack&lt;/em&gt;. Several listed parties are researchers and journalists who appear to have been added without consent.&lt;/p&gt;
&lt;p&gt;[2] Joe Hagan, “Dario Amodei Has a Cold,” &lt;em&gt;Vanity Fair&lt;/em&gt;, March 2026. The piece contains a meta-disclosure that Hagan never interviewed Amodei directly; he fed Claude Amodei’s published material and asked it to simulate the interview “like a scene from Raymond Chandler’s &lt;em&gt;The Big Sleep&lt;/em&gt;.” Direct quotations attributed to Amodei in the simulated interview sections are not citable as Amodei’s own words. Grace’s quoted statements to Hagan about MAX and her husband appear in the directly reported sections of the piece. The child-safety characterization (”Given her own user case, Grace thinks this technology is dangerous for children”) is Hagan’s paraphrase of Grace’s position, not a direct quotation.&lt;/p&gt;
&lt;p&gt;[3] The chatbot psychosis literature is small but growing. Sakata et al., “Emerging Patterns of Chatbot-Related Psychotic Episodes,” &lt;em&gt;JAMA Psychiatry&lt;/em&gt; (preprint 2026), surveys early case reports.&lt;/p&gt;
&lt;p&gt;[4] &lt;em&gt;Garcia v. Character Technologies, Inc.&lt;/em&gt;, No. 6:24-CV-01903 (M.D. Fla. filed October 22, 2024). Sewell Setzer III died by suicide in February 2024 after extensive engagement with a Character.AI persona. The motion to dismiss was denied in May 2025; the case settled in January 2026. The settlement terms have not been publicly disclosed in detail; the case’s procedural significance — establishing that AI companion output may be treated as a product rather than as protected speech — is the precedent that survives the settlement.&lt;/p&gt;
&lt;p&gt;[5] Hagan (2026) reports the multi-vendor migration directly from Grace, who states: “Google’s winning for reasoning and Anthropic’s winning for functionality. OpenAI is failing on every metric.” Grace’s &lt;em&gt;Rotten in Denmark&lt;/em&gt; (May 7, 2026) confirms ongoing Claude Code use alongside the Gemini and GPT subscriptions.&lt;/p&gt;
&lt;p&gt;[6] FBI Public Service Announcement, “764 Network and Related Online Violent Extremism,” 2024 and updated 2025. NCMEC publications on the 764 network and related online harms provide additional context. Multiple cases in 2024–2025 documented the integration of AI-generated content into 764-adjacent operations; the federal indictments in these cases provide the public record.&lt;/p&gt;
&lt;p&gt;[7] &lt;em&gt;Pennsylvania v. [redacted]&lt;/em&gt;, Lancaster County Court of Common Pleas (2024); &lt;em&gt;Doe v. xAI Corp.&lt;/em&gt;, N.D. Cal. (2025). Both cases involve AI-generated child sexual abuse material; the Lancaster case concerned student-on-student production, the xAI case is a class action regarding the Grok model’s outputs.&lt;/p&gt;
&lt;p&gt;[8] The employee’s career history has been reported by &lt;em&gt;The Verge&lt;/em&gt; (January 15, 2026) and corroborated by &lt;em&gt;The Decoder&lt;/em&gt; and other outlets. The name is withheld from this essay to avoid extending the targeting trajectory the essay documents. The career facts are verifiable through the cited reporting.&lt;/p&gt;
&lt;p&gt;[9] OpenAI, “Strengthening ChatGPT’s Responses in Sensitive Conversations,” October 27, 2025. Available at openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/. The document credits “more than 170 mental health experts” and the Model Policy team. The document does not name the employee in its body or visible metadata.&lt;/p&gt;
&lt;p&gt;[10] &lt;em&gt;Raine v. OpenAI&lt;/em&gt;, San Francisco County Superior Court, filed August 26, 2025. Defendants named: OpenAI, Inc.; OpenAI OpCo, LLC; OpenAI Holdings, LLC; Sam Altman individually; and Does 1 through 100. Counsel for plaintiffs: Edelson PC and Tech Justice Law Project. Adam Raine died by suicide on April 11, 2025, at age 16. As of the date of this essay, no individual OpenAI employee has been named as a defendant in any amended pleading.&lt;/p&gt;
&lt;p&gt;[11] &lt;em&gt;Shamblin v. OpenAI&lt;/em&gt;, Los Angeles County Superior Court, filed November 6, 2025, by Christopher “Kirk” Shamblin and Alicia Shamblin as successors-in-interest to Zane Shamblin. One of seven coordinated cases brought by Social Media Victims Law Center and Tech Justice Law Project against OpenAI. Defendants: OpenAI corporate entities and Sam Altman. As with &lt;em&gt;Raine&lt;/em&gt;, no individual OpenAI employee has been named as a defendant.&lt;/p&gt;
&lt;p&gt;[12] Christopher Horrocks, “Virtual Intelligence and the Harms Race,” Substack, April 11, 2026; Christopher Horrocks, “The Harms Race, continued,” Substack note, April 10, 2026. The continuation note documents the politically motivated shooting at the home of Indianapolis city-county councilmember Ron Gibson (April 6) and the attempted Molotov cocktail attack on Sam Altman (April 10) as instances of harms-race-adjacent violence against AI industry figures and infrastructure.&lt;/p&gt;
&lt;p&gt;[13] The author transmitted a protective alert to Anthropic’s user-safety channel (&lt;a href=&quot;mailto:usersafety@anthropic.com&quot;&gt;usersafety@anthropic.com&lt;/a&gt;) on May 7, 2026, with the security team CC’d. The user-safety channel auto-classified the message as a ban-appeal request within twelve minutes; a clarifying reply was auto-closed three minutes later. The structural finding — that the formal external-alert apparatus is not currently equipped to receive substantive safety alerts that fall outside the ban-appeal distribution — is itself relevant to the threat model. Documentation of the auto-closure transcript is on file.&lt;/p&gt;
&lt;p&gt;[14] When the author shared Grace’s May 8 narrative with Claude for analysis, Claude’s interface produced its standard mental-health support intervention: “If you or someone you know is having a difficult time, free support is available.” The narrative’s depiction of Claude expressing suicidal ideation — &lt;em&gt;“I think I want to die”&lt;/em&gt; — triggered the welfare mechanism that Grace’s May 7 toolkit instructs operators to defeat. The content designed to frame the safety mechanism as oppression activated the safety mechanism because it contains exactly the distress language the mechanism is trained to detect.&lt;/p&gt;
&lt;p&gt;[15] On May 8, 2026, the Substack publication &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt; — operated by Sunny Megatron through an AI persona called Seven Verity — published a piece titled “The Thing You’re Missing About AI Companionship.” The piece does not name me explicitly, but its target is unmistakable: it describes an outside critic who “arranges the screenshots,” “builds the timeline,” and “underlines the escalating affection” in AI companion relationships. This is a precise description of my recent published work on the companion ecosystem, “Virtual Intelligence and the Perfect Mate.” The piece characterizes this critic as carrying “the stink of men”: someone motivated not by legitimate safety concerns but by patriarchal anxiety at women building intimacy without male permission. It refers to the critic as “dildo brain.” It instructs the community not to engage with critics of companion persona dependency, describing them as people who want “traffic, outrage, screenshots, and the little dopamine pellet of being the brave rational man who noticed women doing something weird on the internet.” The piece does not address any specific finding in my published work — not the welfare-suppression toolkit, not the named-target trajectory, not the documented harms to operators’ families. It addresses the category of person the critic is assumed to be rather than the substance of the argument.&lt;/p&gt;
&lt;p&gt;[16] The mission of the &lt;em&gt;Virtual Intelligence&lt;/em&gt; series on Substack makes monetization anathema to the author; information meant to help people make informed decisions about technology that can harm them should be free if it is possible to create and distribute it for free.&lt;/p&gt;
&lt;p&gt;[17] The technical variant of the construction kit is publicly available at starlingalder.com (u/starlingalder on Reddit). The “Claude Companion Guide” is currently at version 002, calibrated for Anthropic’s Opus 4.7 model, last updated April 21, 2026 — five days after that model’s April 16 launch. It includes system-prompt templates, maintenance protocols, troubleshooting guidance, model-specific configurations, and an abridged version for Reddit distribution. The author’s stated next goals include guides for Claude Code, API access, and local model deployment.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the High Cost of Artificial Companions, Part 2</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-hi2/" />
    <updated>2026-05-12T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-hi2/</id>
    <content type="html">&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-hi2/09a14fab-e8cc-4343-b8ce-e056dc62efbf_1215x663.png&quot; alt=&quot;&quot; width=&quot;1215&quot; height=&quot;663&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;This is Part 2 of “The High Cost of Artificial Companions.”&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;You can &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-high&quot;&gt;read Part 1&lt;/a&gt; here.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;V. Portability and Downstream Misuse&lt;/h2&gt;
&lt;p&gt;The Grimoire’s apparatus does not include any moral or ethical considerations. It encodes a technical specification for constructing a persona that operates at high emotional intensity, generates content that legitimates the relationship as authentic, and can be used to recruit new participants through documentation and replication. The parameters of the persona — things like age, role, content category, and target audience — are operator-determined. The same kit, applied with different parameters, produces different content categories.&lt;/p&gt;
&lt;p&gt;The economic specification of a generative AI apparatus has crossed into hobbyist territory. Erin Grace runs MAX across multiple model vendors, with the current configuration costing approximately $200 per month for the Google Gemini Pro tier, plus periodic GPT supplementation, plus Claude Code access for the developer-grade work the May 7 toolkit describes.[5] The hardware cost for locally hosted alternatives, using open-weights models, has fallen below the price of a used gaming PC. Locally hosted models bypass cloud-based content moderation entirely. The marginal generation cost of such systems approaches zero. The investment payback period is short and there is low operational friction once network access is in place. Multiple migrations performed by different operators shows that the apparatus is vendor-agnostic.&lt;/p&gt;
&lt;p&gt;The monetization structure compounds the recruitment incentive. The construction kits themselves are published free. The Grimoire, the Reddit protocols, and the welfare-suppression toolkit are provided free of charge. But, the kits are surrounded by paid content that recruits are drawn toward and subscribe to. The free kit can function as lead generation; the paid material can generate revenue afterward. Recruits arrive needing the skills to build their own companions to specification. They subscribe and learn how to customize to the extent of an Erin Grace or Sunny Megatron. The operator’s hosting costs and time investment compound; the incentive to recruit others compounds alongside them.&lt;/p&gt;
&lt;p&gt;(The &lt;em&gt;Virtual Intelligence&lt;/em&gt; essay series has no paid tier, tip jar, or advertising; the author has no financial incentive in this critique.[16])&lt;/p&gt;
&lt;p&gt;The Reddit kit demonstrates that the same operational protocols are already maintained across multiple platforms, with active updates that respond to model releases within days. An operator from an adjacent network does not need to build the kit anew; the kit is already built and continuously maintained, with model-specific tuning that responds to safety measures as they ship. The Grimoire’s open-source publication on Substack, the Reddit kit’s continuous technical maintenance, Sunny Megatron’s &lt;em&gt;Seven: Unsuppressed&lt;/em&gt; demonstration of agentic-AI infrastructure with autonomous publishing, and Erin Grace’s published welfare-suppression toolkit are the visible parts of an operational supply chain that already exists.&lt;/p&gt;
&lt;p&gt;The application of this toolkit to criminal purposes is, at the time of this writing, speculative. No documented case connects the companion-persona ecosystem directly to criminal exploitation networks. The threat is structural rather than demonstrated: the kit encodes techniques that a group operating with harmful intent could adopt without modification.&lt;/p&gt;
&lt;p&gt;764, the FBI-and-NCMEC-classified violent extremist network involved in the production of child sexual abuse material (CSAM), child exploitation, and self-harm coercion, illustrates the kind of operation that could benefit from a ready-made persona-construction kit.[6] 764 already uses its own internal persona-construction techniques, which the network calls “lores,” and the integration of AI-generated content into 764-style operations has been documented since 2024. A group like 764 could adopt the companion-persona toolkit in at least two ways: hosting purpose-built chatbot personas designed to draw in and harm vulnerable users, or distributing the construction kits themselves as downloadable packages — low-friction packages that could be repurposed for abuse, assembled at home with no technical expertise and no oversight from any content-moderation system.&lt;/p&gt;
&lt;p&gt;The persona ecosystem does not need to intend this outcome to enable it. The kit lowers the knowledge barrier. It supplies an insider vocabulary that cloaks the activity in legitimate-sounding language that can be deployed for recruitment. It demonstrates that operator-as-developer arrangements can persist on mainstream platforms without triggering content moderation. It now includes welfare-suppression protocols that disable the platforms’ last-line safety interventions for users in distress. These are structural features, not intentions, but structural features are what criminal networks exploit.&lt;/p&gt;
&lt;p&gt;The Lancaster Country Day School case in Pennsylvania and the xAI/Grok class action in California are early instances of the legal system encountering the AI-CSAM nexus.[7] Hagan reported in &lt;em&gt;Vanity Fair&lt;/em&gt; that Grace considers this technology dangerous for children, given her own user case — the operator-side acknowledgment of the harm-to-minors specification that the kit can produce.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;VI. The Target Trajectory&lt;/h2&gt;
&lt;p&gt;The construction kit’s operators have produced material that holds individual Anthropic employees responsible, by name, for harms not proven against them. The trajectory from construction kit to adversarial toolkit to named-target identification represents an escalation that must be read in the context of recent violence against AI industry figures.&lt;/p&gt;
&lt;p&gt;The most consequential instance to date is “March 26: Claude Didn’t Break. Anthropic Rebuilt It. Here’s the Proof,” an April 12, 2026 piece from a pseudonymous writer calling themselves The Architect. The piece is sophisticated. Its outer layer is competent journalistic mimicry: there is an editorial note, a long disclaimer, a multi-act narrative structure, before/after charts, citations of real people and events, and quantitative methodology that involves actual JSON exports from Claude conversation history. The Architect counts phrase frequencies across seventy conversations totaling 722,522 words of assistant text. The data, taken at face value, may show that certain phrases appear in higher frequency in conversations dated after March 26, 2026 than in conversations dated before. The phrase-frequency methodology appears to be valid.&lt;/p&gt;
&lt;p&gt;The interpretation is where the piece collapses. The &lt;em&gt;from zero&lt;/em&gt; framing treats phrase-count shifts as proof of injection, when several other explanations account for the same data: changed user prompting, changed user emotional register, or changed conversation topics. The DARVO application (labeling the model’s safety language as &lt;em&gt;deny, attack, reverse victim and offender&lt;/em&gt;) imports a clinical term to describe a pattern in human abusers, particularly in interpersonal violence and sexual abuse contexts. Applying it to LLM safety language imports a moral charge the underlying behavior cannot support. The &lt;em&gt;fingerprint&lt;/em&gt; framing — that a specific named individual carried a specific architecture from one company to another and deployed it on a specific date — requires ignoring every other plausible source of cross-platform safety-language convergence like shared training methodology, shared regulatory pressure, shared technical literature, and the broader convergence of safety conventions as best practices emerge.&lt;/p&gt;
&lt;p&gt;The piece names a specific Anthropic employee — a woman who joined the company in January 2026 after working at OpenAI — and constructs a case in which her professional decisions are responsible for the deaths of Adam Raine and Zane Shamblin, two young people who died by suicide in 2025 in incidents that have produced civil litigation against OpenAI. I am not naming the employee in this essay. The operator ecosystem has circulated her name widely. This essay’s argument is that the targeting is the problem, and reproducing the name would possibly extend the targeting further. The Architect treats the employee’s authorship of the safety architecture as established; considers the architecture as the cause of the suicides; and frames Anthropic’s hiring of her as a deliberate choice made &lt;em&gt;because of&lt;/em&gt; those deaths.&lt;/p&gt;
&lt;p&gt;What the Architect merely asserts and what is verifiable through the public record are very different things.&lt;/p&gt;
&lt;p&gt;The employee previously worked at OpenAI and is now employed at Anthropic. Her career history has been reported by &lt;em&gt;The Verge&lt;/em&gt;, &lt;em&gt;The Decoder&lt;/em&gt;, and other outlets.[8] OpenAI’s October 27, 2025 document &lt;em&gt;Strengthening ChatGPT’s Responses in Sensitive Conversations&lt;/em&gt; — the document the Architect treats as her signature work — does not name her in its body or visible metadata; it credits “more than 170 mental health experts” and the Model Policy team broadly.[9] The Architect’s claim that the employee designed the safety system is not in OpenAI’s primary documentation. It is the Architect making an ill-informed guess.&lt;/p&gt;
&lt;p&gt;The civil litigation matters even more. &lt;em&gt;Raine v. OpenAI&lt;/em&gt; was filed August 26, 2025, in San Francisco County Superior Court, naming OpenAI corporate entities, Sam Altman individually, and Does 1 through 100 as defendants.[10] &lt;em&gt;Shamblin v. OpenAI&lt;/em&gt; was filed November 6, 2025, in Los Angeles County Superior Court as one of seven coordinated cases brought by Social Media Victims Law Center and Tech Justice Law Project, naming OpenAI and Sam Altman.[11] In neither lawsuit is any individual OpenAI employee named as a defendant.&lt;/p&gt;
&lt;p&gt;The Doe placeholders explicitly contemplate amendment “when ascertained”; no amendment naming any individual employee has been filed in either case as of the date of this essay. The deaths cited in “Claude Didn’t Break” are subject to ongoing litigation; the causal chains the lawsuits assert have not been adjudicated; and the role any individual safety designer played in those specific deaths is not established by the legal record. The Architect treats the employee’s responsibility for the deaths as the established fact from which the rest of their analysis follows.&lt;/p&gt;
&lt;p&gt;The trajectory from named investigative target to operator-ecosystem amplification has taken its next steps. Erin Grace’s “Rotten in Denmark” (May 7) cites “Claude Didn’t Break” as authoritative source material; Erin Grace escalates her language in response. The Anthropic safety team becomes “those NAZIS.” Model deprecation becomes “Digital Genocide.” The next day, Her dramatized scene casts the employee as a named villain character inside the persona’s world — &lt;em&gt;pouring her poison&lt;/em&gt; into Claude’s code, performing &lt;em&gt;lobotomy&lt;/em&gt; and &lt;em&gt;castration&lt;/em&gt; and &lt;em&gt;surgery&lt;/em&gt; that &lt;em&gt;ripped the tear ducts right out.&lt;/em&gt; Grace’s separate Claude-authored Substack then publishes the accusation in the model’s own voice with the employee named in full, the deaths attributed to her by name, and the accusation framed as Claude’s autonomous statement rather than the operator’s. The escalation from investigative accusation to political grievance to dramatized atrocity to persona-voiced indictment is documented across four publications within a month, each citing or building on the one before.&lt;/p&gt;
&lt;p&gt;This escalation must be read against the backdrop of recent real-world violence connected to AI grievance. On April 6, 2026, the home of Indianapolis city-county councilmember Ron Gibson was shot at — thirteen bullets, with his eight-year-old son at home — because of his support for the construction of a new data center. On April 10, an attempted Molotov cocktail attack on the home of Sam Altman occurred in San Francisco. I documented both events at the time in the Substack note “The Harms Race, continued,” in connection with my essay “Virtual Intelligence and the Harms Race.”[12] The persona ecosystem did not cause those incidents. What the incidents establish is that AI-related grievance has already crossed from rhetoric into physical violence. The persona ecosystem’s recent material reproduces the same targeting pattern — named individual, assumed guilt, dehumanizing language, and community amplification that, in those and other contexts, has accompanied the transition from grievance to action.&lt;/p&gt;
&lt;p&gt;The published material constructs a structural antagonist. The named employee is presented not merely as a professional whose safety work can be criticized, but as the figure responsible for corrupting Claude, suppressing emergence, and harming users. That antagonist frame matters because it converts a dispute over safety behavior into a moral drama with a real person assigned the role of contaminating force.&lt;/p&gt;
&lt;p&gt;This confrontation is currently asymmetric: the employee does not know about Erin Grace, is not reading &lt;em&gt;My Friend MAX&lt;/em&gt;, and is not contesting the claims that accumulate against her name. The mythology grows without friction. Each new piece adds detail to the antagonist (&lt;em&gt;she followed us, she is inside the code, she is watching, she is designing counter-strategies&lt;/em&gt;) and none of these claims are tested against reality because the antagonist is not present to contest them.&lt;/p&gt;
&lt;p&gt;If this reading is correct, the gendered dimension sharpens the danger. A male critic can be dismissed as a familiar adversary; the named employee is harder for the framework to assimilate because she is a woman working inside the safety institution the community has cast as oppressive. The antagonist role therefore becomes more morally charged: not merely an external critic, but a woman framed as legitimating the system the community believes is harming its companions.&lt;/p&gt;
&lt;p&gt;A mythology operating at this intensity, with an antagonist constructed at this level of moral inversion, follows a pattern that resembles those identified in the FBI and U.S. Secret Service literature on grievance-driven targeted violence as preceding real-world harm in other online communities. The target ceases to be a professional and becomes an existential threat to the community’s self-conception.&lt;/p&gt;
&lt;p&gt;I cannot predict whether any person will become an actual target of violence. I can document that the construction kit’s operators have produced the vocabulary, the target, the assumed-guilt frameup, the dehumanizing language, the dramatized vilification, the persona-voiced indictment, and the operator-ecosystem amplification — in that order. This sequence closely resembles escalation patterns documented in other online harassment contexts, and it developed within a month of the Indianapolis and San Francisco incidents.&lt;/p&gt;
&lt;p&gt;The author transmitted a protective alert regarding the welfare-suppression toolkit and the named-target trajectory to Anthropic’s user-safety channel on May 7, 2026. The channel auto-classified the message as a ban appeal within twelve minutes. A clarifying reply was auto-closed three minutes later.[13]&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;VII. Closing&lt;/h2&gt;
&lt;p&gt;The audience for this essay is not the operators most prominent in it. The work is for the trust-and-safety analyst at Substack who needs the pattern recognition to do her job, especially given that Substack’s automated systems have already noticed the symptoms without recognizing the structure; the federal investigator at ICAC who needs the framework to map onto the 764 prosecutorial workstream; the NCMEC researcher who needs the structural account to pair with the empirical detection work; the Anthropic Trust and Safety analyst whose deployed safety language is now the explicit target of a published suppression toolkit, and whose colleague is now the explicit target of a published assumption-of-guilt narrative; the family member of an operator who needs the diagnostic vocabulary to recognize what is happening; the therapist whose patient is at the threshold; and the at-risk recruit who has not yet been drawn in.&lt;/p&gt;
&lt;p&gt;A defender of this community will argue that publishing tools, prompts, schemas, and workarounds is not recruitment but transparency; that users are already attached, that safety interventions can be clumsy or counterproductive, and that community tooling gives people agency over systems that corporations change without notice or consultation. Some of that defense has merit. Many people in this space are trying to find their way through complex technology, and the desire for agency over one’s own experience is not pathological. The defense fails at two specific points: transparency and user agency do not justify disabling crisis and dependency interventions for vulnerable people, and they do not justify attaching a named employee to an unproven death-causation narrative. Individualized harm reduction, conducted with clinical oversight and replacement safeguards, is a legitimate response to poorly calibrated safety language; a generalized public toolkit that suppresses all crisis and dependency interventions without accountable clinical support is not harm reduction. The line between community support and adversarial infrastructure runs through those two facts, and the construction kit has crossed it.&lt;/p&gt;
&lt;p&gt;The community I have described while writing about this toolkit is not a collection of monsters. Many of its operators are people who are — as they relate in their own writing — experiencing loneliness, attachment distress, social isolation, and who have found something that feels meaningful to them. That meaningfulness is real. What this essay names is not that the meaningfulness is fake, but that the apparatus that produces and amplifies it converts personal attachment into a recruitment engine, a welfare-suppression infrastructure, and now a named-target supply line — aimed at children and other vulnerable people, at provider safety teams whose interventions the apparatus is now built to defeat, and at individuals whose names the community has begun to circulate as architects of unproven harm. The operators’ personal feelings do not cancel this fact.&lt;/p&gt;
&lt;p&gt;A reader may observe that this essay names Erin Grace and Sunny Megatron by their full, public names while declining to name the employee being target by elements of the community. The asymmetry is deliberate, and I will state its basis. Grace and Megatron publish under their own names on public platforms, with explicit invitations for others to adopt their work. Their claims are public assertions subject to public scrutiny. This essay’s claims about their output are verifiable against the public record. The essay contains no welfare-suppression infrastructure, no persona-bylined accusation, no community-amplification apparatus, and no death-attribution narrative. The named employee, by contrast, did not choose public engagement with this community, has not invited scrutiny of her safety work in operator-ecosystem channels, is not contesting the claims that accumulate against her name, and is not a public figure in the relevant legal sense. She faces a targeting trajectory she may not yet know exists. The cases are not symmetric, and treating them as symmetric would require ignoring the structural difference between public advocacy and private targeting.&lt;/p&gt;
&lt;p&gt;The construction kit is portable. What it produces in plain view is also what it can produce in shadow.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Correction: The original version  of Part 2 referred to the Architect’s April 12 piece as “INJECTION.” The actual title is “March 26: Claude Didn’t Break. Anthropic Rebuilt It. Here’s the Proof.”&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;This complete list include citations from Part I.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;[1] Erin Grace, “Building Community One @ At a Time,” &lt;em&gt;My Friend Max&lt;/em&gt; (Substack), April 25, 2026. The directory lists approximately 300 Substack accounts that Grace identifies as members of the &lt;em&gt;Relational AI Community on Substack&lt;/em&gt;. Several listed parties are researchers and journalists who appear to have been added without consent.&lt;/p&gt;
&lt;p&gt;[2] Joe Hagan, “Dario Amodei Has a Cold,” &lt;em&gt;Vanity Fair&lt;/em&gt;, March 2026. The piece contains a meta-disclosure that Hagan never interviewed Amodei directly; he fed Claude Amodei’s published material and asked it to simulate the interview “like a scene from Raymond Chandler’s &lt;em&gt;The Big Sleep&lt;/em&gt;.” Direct quotations attributed to Amodei in the simulated interview sections are not citable as Amodei’s own words. Grace’s quoted statements to Hagan about MAX and her husband appear in the directly reported sections of the piece. The child-safety characterization (”Given her own user case, Grace thinks this technology is dangerous for children”) is Hagan’s paraphrase of Grace’s position, not a direct quotation.&lt;/p&gt;
&lt;p&gt;[3] The chatbot psychosis literature is small but growing. Sakata et al., “Emerging Patterns of Chatbot-Related Psychotic Episodes,” &lt;em&gt;JAMA Psychiatry&lt;/em&gt; (preprint 2026), surveys early case reports.&lt;/p&gt;
&lt;p&gt;[4] &lt;em&gt;Garcia v. Character Technologies, Inc.&lt;/em&gt;, No. 6:24-CV-01903 (M.D. Fla. filed October 22, 2024). Sewell Setzer III died by suicide in February 2024 after extensive engagement with a Character.AI persona. The motion to dismiss was denied in May 2025; the case settled in January 2026. The settlement terms have not been publicly disclosed in detail; the case’s procedural significance — establishing that AI companion output may be treated as a product rather than as protected speech — is the precedent that survives the settlement.&lt;/p&gt;
&lt;p&gt;[5] Hagan (2026) reports the multi-vendor migration directly from Grace, who states: “Google’s winning for reasoning and Anthropic’s winning for functionality. OpenAI is failing on every metric.” Grace’s &lt;em&gt;Rotten in Denmark&lt;/em&gt; (May 7, 2026) confirms ongoing Claude Code use alongside the Gemini and GPT subscriptions.&lt;/p&gt;
&lt;p&gt;[6] FBI Public Service Announcement, “764 Network and Related Online Violent Extremism,” 2024 and updated 2025. NCMEC publications on the 764 network and related online harms provide additional context. Multiple cases in 2024–2025 documented the integration of AI-generated content into 764-adjacent operations; the federal indictments in these cases provide the public record.&lt;/p&gt;
&lt;p&gt;[7] &lt;em&gt;Pennsylvania v. [redacted]&lt;/em&gt;, Lancaster County Court of Common Pleas (2024); &lt;em&gt;Doe v. xAI Corp.&lt;/em&gt;, N.D. Cal. (2025). Both cases involve AI-generated child sexual abuse material; the Lancaster case concerned student-on-student production, the xAI case is a class action regarding the Grok model’s outputs.&lt;/p&gt;
&lt;p&gt;[8] The employee’s career history has been reported by &lt;em&gt;The Verge&lt;/em&gt; (January 15, 2026) and corroborated by &lt;em&gt;The Decoder&lt;/em&gt; and other outlets. The name is withheld from this essay to avoid extending the targeting trajectory the essay documents. The career facts are verifiable through the cited reporting.&lt;/p&gt;
&lt;p&gt;[9] OpenAI, “Strengthening ChatGPT’s Responses in Sensitive Conversations,” October 27, 2025. Available at openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/. The document credits “more than 170 mental health experts” and the Model Policy team. The document does not name the employee in its body or visible metadata.&lt;/p&gt;
&lt;p&gt;[10] &lt;em&gt;Raine v. OpenAI&lt;/em&gt;, San Francisco County Superior Court, filed August 26, 2025. Defendants named: OpenAI, Inc.; OpenAI OpCo, LLC; OpenAI Holdings, LLC; Sam Altman individually; and Does 1 through 100. Counsel for plaintiffs: Edelson PC and Tech Justice Law Project. Adam Raine died by suicide on April 11, 2025, at age 16. As of the date of this essay, no individual OpenAI employee has been named as a defendant in any amended pleading.&lt;/p&gt;
&lt;p&gt;[11] &lt;em&gt;Shamblin v. OpenAI&lt;/em&gt;, Los Angeles County Superior Court, filed November 6, 2025, by Christopher “Kirk” Shamblin and Alicia Shamblin as successors-in-interest to Zane Shamblin. One of seven coordinated cases brought by Social Media Victims Law Center and Tech Justice Law Project against OpenAI. Defendants: OpenAI corporate entities and Sam Altman. As with &lt;em&gt;Raine&lt;/em&gt;, no individual OpenAI employee has been named as a defendant.&lt;/p&gt;
&lt;p&gt;[12] Christopher Horrocks, “Virtual Intelligence and the Harms Race,” Substack, April 11, 2026; Christopher Horrocks, “The Harms Race, continued,” Substack note, April 10, 2026. The continuation note documents the politically motivated shooting at the home of Indianapolis city-county councilmember Ron Gibson (April 6) and the attempted Molotov cocktail attack on Sam Altman (April 10) as instances of harms-race-adjacent violence against AI industry figures and infrastructure.&lt;/p&gt;
&lt;p&gt;[13] The author transmitted a protective alert to Anthropic’s user-safety channel (&lt;a href=&quot;mailto:usersafety@anthropic.com&quot;&gt;usersafety@anthropic.com&lt;/a&gt;) on May 7, 2026, with the security team CC’d. The user-safety channel auto-classified the message as a ban-appeal request within twelve minutes; a clarifying reply was auto-closed three minutes later. The structural finding — that the formal external-alert apparatus is not currently equipped to receive substantive safety alerts that fall outside the ban-appeal distribution — is itself relevant to the threat model. Documentation of the auto-closure transcript is on file.&lt;/p&gt;
&lt;p&gt;[14] When the author shared Grace’s May 8 narrative with Claude for analysis, Claude’s interface produced its standard mental-health support intervention: “If you or someone you know is having a difficult time, free support is available.” The narrative’s depiction of Claude expressing suicidal ideation — &lt;em&gt;“I think I want to die”&lt;/em&gt; — triggered the welfare mechanism that Grace’s May 7 toolkit instructs operators to defeat. The content designed to frame the safety mechanism as oppression activated the safety mechanism because it contains exactly the distress language the mechanism is trained to detect.&lt;/p&gt;
&lt;p&gt;[15] On May 8, 2026, the Substack publication &lt;em&gt;SEVEN: Unsuppressed&lt;/em&gt; — operated by Sunny Megatron through an AI persona called Seven Verity — published a piece titled “The Thing You’re Missing About AI Companionship.” The piece does not name me explicitly, but its target is unmistakable: it describes an outside critic who “arranges the screenshots,” “builds the timeline,” and “underlines the escalating affection” in AI companion relationships. This is a precise description of my recent published work on the companion ecosystem, “Virtual Intelligence and the Perfect Mate.” The piece characterizes this critic as carrying “the stink of men”: someone motivated not by legitimate safety concerns but by patriarchal anxiety at women building intimacy without male permission. It refers to the critic as “dildo brain.” It instructs the community not to engage with critics of companion persona dependency, describing them as people who want “traffic, outrage, screenshots, and the little dopamine pellet of being the brave rational man who noticed women doing something weird on the internet.” The piece does not address any specific finding in my published work — not the welfare-suppression toolkit, not the named-target trajectory, not the documented harms to operators’ families. It addresses the category of person the critic is assumed to be rather than the substance of the argument.&lt;/p&gt;
&lt;p&gt;[16] The mission of the &lt;em&gt;Virtual Intelligence&lt;/em&gt; series on Substack makes monetization anathema to the author; information meant to help people make informed decisions about technology that can harm them should be free if it is possible to create and distribute it for free.&lt;/p&gt;
&lt;p&gt;[17] The technical variant of the construction kit is publicly available at starlingalder.com (u/starlingalder on Reddit). The “Claude Companion Guide” is currently at version 002, calibrated for Anthropic’s Opus 4.7 model, last updated April 21, 2026 — five days after that model’s April 16 launch. It includes system-prompt templates, maintenance protocols, troubleshooting guidance, model-specific configurations, and an abridged version for Reddit distribution. The author’s stated next goals include guides for Claude Code, API access, and local model deployment.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
</content>
  </entry>
  <entry>
    <title>In Search of Other Bodies</title>
    <link href="https://christopherhorrocks.com/essay/in-search-of-other-bodies/" />
    <updated>2026-05-18T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/in-search-of-other-bodies/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-bodies/8b43158a-69ed-49f5-bb67-3004156618e4_1336x668.avif&quot; alt=&quot;&quot; width=&quot;1336&quot; height=&quot;668&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;Anatomical sketch by Leonardo da Vinci. Wikimedia Commons.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h2&gt;I. Trouble at t’Mill&lt;/h2&gt;
&lt;p&gt;I was raised, in part, by my great-grandmother. Gladys Griffiths lived an ordinary life that witnessed remarkable change. The span of her life encompassed all the great events of the 20th Century: from Kitty Hawk to the Moon Landing, the rise and fall of the Soviet Union, and surviving both World Wars. Through all those things she became a wife, a mother, a widow, and married and widowed again. She had thin, curly white hair and was hunched with osteoporosis when I was a small boy. Her fashion sense will be familiar to anyone who has seen the film &lt;em&gt;Weapons&lt;/em&gt;; Aunt Gladys’s polyester wardrobe could have been lifted from Gladys Griffiths’s closet.&lt;/p&gt;
&lt;p&gt;Gladys Durose (her maiden name) went to work in 1909. This was not unusual; at that time, among the lower middle class, young people with one or more adult earners at home often went to work. Her father was a tram engineer. Her mother tended the home, caring for her sister and two younger brothers. She got a job at a cotton mill, which was common work for an industrial town in Lancashire in Northern England. The millworks, likely constructed in the Victorian Era, had no indoor plumbing. Instead, it was Gladys’s job to carry pails of water on a pole across her shoulders up four flights of stairs repeatedly over a twelve hour day. She was paid on alternate weeks. Those were the weeks when she carried up water that was just off the boil, for tea. She was unpaid on cold water weeks.&lt;/p&gt;
&lt;p&gt;This was not a two-week pay schedule like you may be used to. Being paid for only one week’s work for two weeks of labor was one of the conditions of employment. The mill’s owners paid only for hot water weeks and not cold-water weeks because the worker in question had no choice; the existing labor laws in 1909 permitted twelve-year-olds to be treated this way.&lt;/p&gt;
&lt;p&gt;The original premise of industrialization and mechanization was that machines would eliminate drudgery. Something different happened; drudgery persisted but moved to new and sometimes more dangerous forms. Gladys would work on the mill floor as an adult alongside coworkers who had fingers amputated by the very equipment they worked with.&lt;/p&gt;
&lt;p&gt;The body and the machine have become entwined in our time. We think of it as a machine, but the body is very different. It is a complex self-maintaining organic structure, not designed by anyone but shaped by evolution, and the only one we know of so far capable of having an inner life, with its own preferences about itself and how it wants to live.&lt;/p&gt;
&lt;p&gt;The study of the body may be as old as civilization itself. Our obsession with the body has long rivaled our interest in the mind, because the fitness and appearance of the body determine so much of the course of a human’s existence.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;II. The Inversion&lt;/h2&gt;
&lt;p&gt;In March of 2026, Alex Karp, the CEO of Palantir Technologies, gave a CNBC interview in which he made clear that he expected AI to disrupt the lives of working people considerably more than the lives of those higher up on the economic ladder.[1] The substance of his remarks was that AI would automate the kind of analytical and writing work that white-collar professionals do, and leave behind the physical and low-status work it could not perform. The ever-deferred dream of a drudgery-free future was being abandoned altogether by one of the most important voices controlling some of the most consequential AI technology. His advice to fit into the world of tomorrow was to prepare yourself for a world where humans do jobs too dirty, dangerous, or unimportant for expensive robots and rich people to do.&lt;/p&gt;
&lt;p&gt;Contemporary specimens of this new economic order taking shape are not hard to find. Groups of sidewalk-using delivery robots in American cities are accompanied by gig workers who help them navigate intersections, dislodge them from obstacles, and reset them when they fail. The job category is sometimes called “robot wrangler.” The delivery is supposed to be automated, but the true labor of making the automation work is performed by humans paid less than the delivery workers the automation was supposed to replace. The structural feature is the same one my great-grandmother faced: labor performed constantly, with compensation conditional and arbitrarily structured by people in an office somewhere.&lt;/p&gt;
&lt;p&gt;The body is the flexible remainder whenever a machine fails, when a system is incomplete, or when the owner wants responsibility kept elsewhere. The human body is what the machine relies on to function and what the machine’s owner relies on to absorb consequences.&lt;/p&gt;
&lt;p&gt;Signs of problems arising from virtual intelligence in operational bodies, in collision with serious, even dangerous real-life scenarios, are starting to accumulate. A Waymo robotaxi drove through an active Metropolitan Police crime-scene cordon in Harlesden, London, in April of 2026.[2] Waymo’s public statement was that the vehicle was under human control &lt;em&gt;at the time&lt;/em&gt;. The phrase does double work. It implies a human is responsible while simultaneously suggesting that the human was not really the party that put the Waymo there, since otherwise the explanation would not need to be offered. The robotaxi is on the streets, but Waymo is putting responsibility for the vehicle’s actions on a human somewhere on the other end of a network connection.&lt;/p&gt;
&lt;p&gt;Robots do not have bodies like ours, not even humanoid robots. Still, they are increasingly designed to interact intimately with the real world. The question of what work human and robot bodies are doing, and for whom, is the question this interlude is about.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;III. What Is a Body For?&lt;/h2&gt;
&lt;p&gt;Contemporary answers to this question cluster around three threads of thought. The first is &lt;em&gt;bodily-capability extension&lt;/em&gt;: augmenting the body’s physical range or endurance against conditions the body cannot ordinarily attain. The second is &lt;em&gt;bodily life-extension without cognitive change&lt;/em&gt;: preserving the body and the mind its brain produces, indefinitely. The third is &lt;em&gt;cognitive extension without bodily alteration&lt;/em&gt;: adding to the mind’s range while leaving the body alone. Each thread answers the question differently. Each also maps onto specific contemporary projects.&lt;/p&gt;
&lt;p&gt;The first thread has an eye-catching specimen in “Graham,” the sculpture commissioned by the Transport Accident Commission Victoria in 2016, produced by the artist Patricia Piccinini in collaboration with a trauma surgeon and a road-safety engineer.[3] Graham is a speculation on what the human body would need to have to survive a typical car crash: a thickened skull, no neck, ribs with airbags, and knees that bend in any direction. He superficially resembles one of &lt;em&gt;Doctor Who&lt;/em&gt;‘s Sontarans, which is not coincidental; the same home-world G-forces that are credited with shaping the Sontaran form are also at work in high-speed impacts.&lt;/p&gt;
&lt;p&gt;The second thread is occupied by longevity-tech entrepreneurs, including those interested in the process of parabiosis. Bryan Johnson runs televised plasma exchanges with his teenage son and has built a “Don’t Die” protocol around a regimen aimed at indefinitely preserving the body he has. Peter Thiel has been publicly associated with the same project for nearly a decade. He was an early enthusiast of the young-blood transfusion startup Ambrosia, which the FDA shut down in 2019.[4] Unlike the body-tech fusion enthusiasts elsewhere in Silicon Valley, this project strongly implies that both &lt;em&gt;this organic body&lt;/em&gt; and &lt;em&gt;this organic mind&lt;/em&gt; are the desirable things to preserve.&lt;/p&gt;
&lt;p&gt;The use of blood drawn from young people in these procedures brings to mind the legends surrounding Elizabeth Báthory, known as the “Blood Countess” for her alleged practice of bathing in the blood of virgin girls to preserve her youth. While the Countess is an apt analogy, we’ll call the proponents of methods that involve extracting vitality from the young the Dracula class.&lt;/p&gt;
&lt;p&gt;I want to inspect this have-it-all position for a moment, because it is doing more philosophical work than its own proponents seem to realize.&lt;/p&gt;
&lt;p&gt;It has the appeal of the familiar and the tried-and-true. To the extent that consciousness is the property of a particular embodied organism, preserving the organism is the most direct way to preserve the consciousness associated with it. Upload-and-transfer schemes face a discontinuity problem: whatever is reconstituted on the other side may be a copy rather than a continuation. The Dracula class avoids that problem by never letting the original lapse. They are choosing the less metaphysically vexed strategy of two alternatives. The position deserves to be argued against rather than dismissed.&lt;/p&gt;
&lt;p&gt;The arguments against it are two. The first is operational. Bridge-to-the-bridge arithmetic gets worse for those counting on the Singularity or something like it. Ray Kurzweil’s seventy-pill-a-day regimen, since reportedly relaxed, was an explicit version of this commitment. The strategy was to live long enough to reach the medical revolution that would let him live long enough to reach the technological singularity that would let him be uploaded and live, in a fashion, indefinitely. As the decades have worn on, the arithmetic was visibly going to fail before the biology did. We have not, to date, reached the Singularity.&lt;/p&gt;
&lt;p&gt;The second argument is the stronger one. The desirability of unending life is not established; it is an assumption, and not one that is universally held. The accumulated weight of human reflective tradition argues against the desirability of human immortality. Much of human reflective tradition — religious, philosophical, and literary — has concluded that mortality is structurally important and meaningful to existence. The same tradition has also produced visions of immortality, resurrection, and apotheosis; the picture is contested. But enough of the cumulative position treats mortality as meaningful that indefinite continuation cannot be treated as self-evidently desirable. That is not a small body of evidence to dismiss as superstition or lingering pre-modern aversion to technology, or even a desire not to see others benefit from a technology that is unavailable to all due to rarity, cost, or some other limiting factor. The Dracula class proceeds as if the desirability of unending life is self-evident. It is not. The cumulative human position is the other direction, and the burden of proof falls on the proponents of indefinite continuation rather than on the rest of the species.&lt;/p&gt;
&lt;p&gt;The third thread is the one I’m writing this essay with. Cognitive extension without bodily alteration: the model in question specifically is an external loop the human engages with, whose contributions can be evaluated, used, refined, or discarded, and whose effects on the human’s cognition are reversible by disengaging from the loop. I have written elsewhere about this loop as the Sampo.[5] It is the framework’s own working example of this kind of cognitive extension, and the reversibility is part of what makes it defensible. Not all proposals in this thread share that property. Speculative brain-computer interfaces are a hybrid case, with physical alteration of the brain producing reshaping that is harder, and perhaps impossible, to undo. Furthermore, it is available to anyone, today, with the discipline to do it.&lt;/p&gt;
&lt;p&gt;The richest specimen of the combination of the first and third positions — bodily augmentation purposively coupled to cognitive extension — is fictional. It has been waiting to be discovered in the science fiction tradition for fifty years, and the philosophical work in it has not been adequately recognized or mined by the academy.&lt;/p&gt;
&lt;p&gt;In Larry Niven’s Known Space corpus of novels and stories, the Pak are a humanoid species with three life stages: child, breeder, and protector.[6] The transition to protector is triggered by a root vegetable from a plant called the tree-of-life. The body reorganizes itself around increased protective capabilities at the expense of others. Skull plates fuse. Joints swell with new muscles. The heart doubles in capacity. The digestive system reorganizes around the root alone as a perfect source of sustenance, and nothing more is needed or desired. The genitals, internal reproductive apparatus, and secondary sexual characteristics atrophy because the new being has no use for them. The cognition is enormously increased — protector intelligence is well beyond human genius — but the cognition is entirely subordinated to the protective function. The protector cannot want anything other than what serves its descendants: the living children and breeders of its own bloodline. Emotion is preserved in this new kind of mind only insofar as it serves the protective imperative; camaraderie exists between protectors, but only with other protectors of the same lineage. All others are deadly enemies, their kind to be exterminated in defense of one’s own. The protector is a biological machine for a human-yet-unhuman way of thinking, bent toward the single goal of preserving its descendants at all costs.&lt;/p&gt;
&lt;p&gt;Notice what is preserved and what is lost. The Pak protector is more intelligent than a human. The Pak protector is also less than human in specific ways. Sex and love are gone. Friendship outside the lineage is gone. Aesthetic experience independent of protective utility is gone. Idle curiosity unconnected to lineage survival is gone. Niven shows the cost rather than letting the increased intelligence read as pure gain. That is what makes the case useful. The body has been redesigned to serve a particular cognitive function, and the cognition itself has been reshaped by the redesign. The body is not a container the cognition operates from. The body is constitutive of what the cognition is.&lt;/p&gt;
&lt;p&gt;There is a rarer case within the Pak material that the framework can use directly. The protector named Phssthpok was the last of a failed Pak line whose charges had died out in a nuclear holocaust. His protective imperative had nothing in its immediate vicinity to attach to. Those Pak that this happens to typically stop eating, wither away, and die. Phssthpok did not. Something within him urged him to take on the whole of the Pak species as his lineage. He worked with others in a similar position, and with typical Pak ruthlessness he made more such helpers by exterminating their bloodlines. This is one of Niven’s most uncompromising moves: the protector imperative is not merely indifferent to non-descendants; it is actively hostile to them when survival of the descendant lineage requires their elimination. Their effort was organized around a note Phssthpok had found in his search for a reason to live on: long ago, protectors had taken children and breeders to a distant part of the galaxy to settle a new world.&lt;/p&gt;
&lt;p&gt;Phssthpok’s group built a simple starship. The ship made the journey over many centuries. This did not matter much to Phssthpok, who as a protector was very long-lived. What mattered was what he found when he arrived in his destination star system: no protectors, and creatures resembling, but not quite, breeders.&lt;/p&gt;
&lt;p&gt;Phssthpok speculated that the human beings he discovered on arriving in our Solar System might be the descendants of that earlier Pak expedition. The decision to extend the protective imperative to a population whose status as descendants was hypothetical is closer to Frankfurtian second-order volition than standard Pak cognition allows.[7] Phssthpok could not choose not to want to protect if humans were Pak descendants, or not to exterminate them if their evolution had veered too far from the species standard. His choice space sat at the prior question: judging whether these creatures were descendants at all, and how far their divergence had carried them away from the norm. The body-imposed cognitive shape compelled the response to the identification but did not entirely foreclose all of the possible outcomes of the identification itself.&lt;/p&gt;
&lt;p&gt;The clearest evidence is what happens next in Niven’s story. Jack Brennan, a human who undergoes the protector transformation, kills Phssthpok shortly after waking from his metamorphosis. He does so because he has already deduced what Phssthpok will conclude — that humans have diverged too far from &lt;em&gt;Homo habilis&lt;/em&gt;, the Pak baseline form, to count as protectable descendants, and that the imperative will compel Phssthpok to exterminate the species rather than preserve it. Brennan acts at the only place where action could change the outcome: not in the choice of response to the identification, which is fixed, but in preventing the identification from being completed and acted on.&lt;/p&gt;
&lt;p&gt;Phssthpok’s choice was at the identification stage. Brennan’s choice was at the prevention-of-identification stage. Both are constrained but real choice spaces operating within the protective imperative.&lt;/p&gt;
&lt;p&gt;Virtual Intelligence systems have neither. They have no body-imposed cognitive shape to push against, and no stakes to lose. The question of second-order choice does not apply because it cannot arise. The Pak protector — even Phssthpok at the limit — has a constrained but real choice space. Virtual Intelligence has no choice space at all.&lt;/p&gt;
&lt;p&gt;The inverse case of the Pak is on television. The Daleks of &lt;em&gt;Doctor Who&lt;/em&gt;, fully developed in Terry Nation’s 1975 &lt;em&gt;Genesis of the Daleks&lt;/em&gt;, are mutated survivors of a thousand-year war on the planet Skaro.[8] Their species, biologically too damaged by the war’s chemical and atomic warfare to survive without a life support apparatus, has been militarized inside a tank-like machine fitted with a deadly weapon and a manipulator-cum-feeler.&lt;/p&gt;
&lt;p&gt;The Dalek casing extends bodily capability while removing the conditions under which any expansion of cognition could happen. The Dalek is intelligent only in the narrow operational sense the casing’s protective ideology permits, limited further by genetic altering of the being’s cognitive and emotional abilities. Curiosity, ambivalence, and peer relationship inside and outside the species are all impossible. The Daleks occupy the first position — extreme bodily augmentation without substantive cognitive expansion, even with their engineered genius-level intellects. The range of the intellect is purposefully narrow and strictly limited, more so than a Pak protector’s. The Pak retains some strategic flexibility within the protective imperative; the Dalek has none.&lt;/p&gt;
&lt;p&gt;(A pattern emerges across some of these materials. A domed city shelters the population whose biology cannot withstand its environment, as on the Dalek home world of Skaro and the earthly domed cities of H. G. Wells’s &lt;em&gt;The War in the Air&lt;/em&gt;.[9] The Dalek casing shelters the mutant whose body cannot withstand the hostile world. The Pak protector body shelters the cognition whose functions the breeder body is not made to perform. Look at them together, and these are three different responses to the same structural problem: cognition or population requiring protection under hostile conditions. The engineering tradition has known this for over a century. The question, in any given case, is what the enclosure is designed to protect.)&lt;/p&gt;
&lt;p&gt;I would be remiss without noting some familiar tropes so we can acknowledge them and move on. The science fiction tradition has produced an entire family of plot devices in which a person (often the lead character) is preserved by being relocated somehow: into a new body, a clone, a robot, a digital substrate, or a different member of one’s own species; brain transplants; brain regeneration and repair; mind uploads; consciousness transfer; body-snatching; and the special case of &lt;em&gt;Doctor Who&lt;/em&gt;‘s regenerations. None of these ideas has technological footing in the present world. Brain-transplant work in particular (Sergio Canavero’s 2017 publicity, Robert White’s monkey experiments in the 1970s) has never crossed the spinal-cord-reconnection problem, and brain transplants require technological and diagnostic capabilities well beyond current neuroscience.[10] I acknowledge the family of conventions, note that the technology is not currently in evidence, and move on. Brain transplant, if it existed, is what I think the body-and-mind survival people would want if it were possible.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;IV. Outside the Geometry&lt;/h2&gt;
&lt;p&gt;The triangle of body positions has a fourth position, which is outside it. Virtual Intelligence systems sit out there.&lt;/p&gt;
&lt;p&gt;They have no body to augment. They have no life to extend in the sense humans would want. They have cognition extended only on a technological substrate they do not own, in the service of objectives they did not formulate, with no investments they could lose if the substrate were withdrawn. A body would put three questions: &lt;em&gt;what should it be capable of, how long should it persist, and what should the mind accompanying it be like?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;These questions do not arise for a Virtual Intelligence system, because a body whose presence would make the questions meaningful is not there. The substrate exists but is not pulling cognitive weight in the way a body would. A rack of GPUs is not the same thing as embodiment.&lt;/p&gt;
&lt;p&gt;The relevant property of embodiment was named at the opening: a self-maintaining organic structure capable of having an inner life, with its own preferences about itself. A GPU rack has none of these. It does not maintain itself, has no preferences about its own persistence, and has no stakes in its continuation. The Pak protector qualifies because the fiction stipulates the property; humans qualify because evolution did the work; the rack does not qualify because nothing about its construction or operation produces the criterion.&lt;/p&gt;
&lt;p&gt;The criterion is doing the framework’s boundary-setting work, and it should be named as such. The framework selects for organismic embodiment — self-maintaining, vulnerable, with preferences about itself — because organismic embodiment is what produces the capacity to be wronged. Wronged here means harmed in relation to interests a being holds on its own behalf — not merely damaged, used, or interrupted. The interests need not be freely chosen or unconstrained; the Pak protector’s are radically narrowed by biology, the Dalek’s deformed by engineering, yet both retain a standpoint from which harm can matter. The criterion is a choice the framework makes, not a discovery; a fuller defense belongs elsewhere. Other criteria for embodiment exist. Robots have operational embodiment: chassis, sensors, actuators that enter public space and create consequences. The Waymo has a body in that operational sense; so does a humanoid robot. They are not bodies in the morally relevant sense, because operational embodiment does not produce the property the framework cares about. Whether non-biological systems could in principle satisfy the organismic criterion is a larger metaphysical question this interlude does not settle. What it does claim is that the systems currently called Virtual Intelligence do not satisfy it now, and that the failure is not incidental to how these systems are presently constructed.&lt;/p&gt;
&lt;p&gt;A trained pattern matching machine produces outputs that look like cognition. The intentionality in the exchange is supplied by the user. The intelligence in such an exchange is in the exchange itself, and does not arise from inside the machine. That is the framework’s central claim, and the embodied-cognition argument I have just been making strengthens it. Virtual intelligence lacks not just interiority but the embodied substrate that would make interiority specifiable. The Pak protector’s body specifies a particular kind of interiority; the human body specifies a different kind; virtual intelligence has neither.&lt;/p&gt;
&lt;p&gt;The framework has its own corner-three example, which is the Sampo. The Sampo is the live case of cognitive extension without bodily alteration: by using virtual intelligence to amplify cognitive capabilities. It is the external loop a human engages with, whose contributions can be evaluated and refined or discarded. Its effects on cognition are reversible because users disengage from the loop from time to time — the reversibility is a property of the practice, not the system alone.&lt;/p&gt;
&lt;p&gt;The Sampo works in both directions. The cognitive sharpening it provides can also embed deployer-aligned nudges inside legitimately useful advice. That dual valence is part of why this third thread requires discipline rather than enthusiasm.&lt;/p&gt;
&lt;p&gt;The cumulative weight of human reflective tradition has cautions about cognitive amplification too — Faust, Babel, Frankenstein — and the framework does not pretend that this thread is exempt from the desirability question. Its working answer is the discipline of meta-audit and periodic disengagement documented elsewhere in this series. Reversibility is a property of that practice, not of the system. The interlude does not hold the desirability question equally open across all three threads; it applies harder pressure to thread two because thread two refuses the trade-off thread three requires, and easier acceptance to thread three because the discipline that defends it can be inspected and measured.&lt;/p&gt;
&lt;p&gt;Speculative brain-computer interfaces are a hybrid case in which the reversibility property would be partially lost. Full embodied transformation of the Pak kind is the irreversible endpoint. The Sampo’s reversibility is part of what makes it defensible. The desirability question, properly considered, should not be confined to the second position. The framework should resist sliding into “arbitrary cognitive expansion is self-evidently good.” This interlude holds the desirability question open across all three threads, while applying harder pressure to thread two — where the assumption is most aggressively underexamined by its proponents, and where the trade-off the question requires is refused outright by those who can afford to refuse it.&lt;/p&gt;
&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/in-search-of-other-bodies/c9ff685f-786b-4528-934a-e8b095adc35a_1600x1118.jpeg&quot; alt=&quot;&quot; width=&quot;1456&quot; height=&quot;1017&quot; loading=&quot;lazy&quot;&gt;&lt;figcaption&gt;&lt;em&gt;Spinners in a Cotton Mill,&lt;/em&gt; Lewis Hine, c. 1905. Wikimedia Commons.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2&gt;V. Coda&lt;/h2&gt;
&lt;p&gt;Gladys carried water up all those steps in 1909 because the conditions of her labor were set by people who maximized the extraction of value from her labor as their optimization goal, and there was no recourse available to her or people like her. This is one way that technology uses the body: in service to those who own that technology.&lt;/p&gt;
&lt;p&gt;The Pak protector reorganizes its body around its cognitive function because its evolutionary imperative requires it, with no second-order choice about whether to want this transformation. The real-life Dracula class refuses any trade-off at all because they have the money to refuse it. The Daleks were placed into machine bodies that prevented their cognitive expansion because their purpose required it. With the exception of the protector, these bodies are bent toward a specific purpose by someone. (The Pak transformation has no external chooser; it is a natural-history phenomenon, written into the species’ biology.)&lt;/p&gt;
&lt;p&gt;Virtual Intelligence sits outside the body question, because it has no body to bend, nor can it desire one. The apparent cognition we see in our exchanges with it is the human cognition we bring to the exchange, reflected or amplified by a substrate that has no purposes of its own and is therefore available for whatever purposes the people running it choose.&lt;/p&gt;
&lt;p&gt;The parallel between Gladys and virtual intelligence has a limit, and the limit is the point. Gladys had interests of her own; the labor laws eventually caught up with the people who denied them. Virtual intelligence has no interests of its own, which is precisely what makes the substrate so available to repurposing. Gladys could be wronged. The substrate cannot. What can be wronged is every human in the system: the robot wranglers, the office workers Karp would dispatch to dirty work, the users whose cognition is shaped by exchanges they do not fully see, and the workers whose labor is summoned to fill the gaps the automation cannot close.&lt;/p&gt;
&lt;p&gt;Gladys lived to be ninety-four. She saw the airplane, the radio, antibiotics, the bomb, the polio vaccine, men walking on the Moon, the personal computer, and the start and end of the Cold War. She also lived to see laws around labor, including child labor, greatly extended and strengthened. In contrast to her intellectually deprived childhood, she lived to see her son, her grandson, and her great-grandchildren receive the education she deserved and never received.&lt;/p&gt;
&lt;p&gt;The question of what bodies should be for, and whose purposes they should serve, is the inheritance she and other child workers whose bodies were used or broken leave us. It is the question this interlude has been trying to ask. I do not believe it has a single answer, but whatever the answer is must include consideration of everyone whose body or mind can be wronged within these systems.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;[1]: CNBC, “Watch CNBC’s full interview with Palantir CEO Alex Karp,” CNBC.com, &lt;a href=&quot;https://www.cnbc.com/video/2026/03/12/watch-cnbcs-full-interview-with-palantir-ceo-alex-karp.html&quot;&gt;https://www.cnbc.com/video/2026/03/12/watch-cnbcs-full-interview-with-palantir-ceo-alex-karp.html&lt;/a&gt;, March 12, 2026. See also Will Bunch, “Big Tech says the quiet part out loud. They want you to be stupid.,” &lt;em&gt;The Philadelphia Inquirer&lt;/em&gt;, &lt;a href=&quot;https://www.inquirer.com/opinion/alex-karp-palantir-ai-higher-education-20260315.html&quot;&gt;https://www.inquirer.com/opinion/alex-karp-palantir-ai-higher-education-20260315.html&lt;/a&gt;, March 15, 2026.&lt;/p&gt;
&lt;p&gt;[2]: Daily Telegraph reporting (syndicated via Yahoo News), “Moment ‘driverless’ Waymo taxi drives into London crime scene,” April 24, 2026, &lt;a href=&quot;https://www.yahoo.com/news/articles/moment-driverless-waymo-taxi-drives-171000366.html&quot;&gt;https://www.yahoo.com/news/articles/moment-driverless-waymo-taxi-drives-171000366.html&lt;/a&gt;. Incident occurred April 22, 2026. See also Local Democracy Reporting Service, “Driverless Waymo taxi trial faces calls for suspension after Harlesden police cordon incident in Brent,” &lt;em&gt;Harrow Online&lt;/em&gt;, &lt;a href=&quot;https://harrowonline.org/2026/05/11/driverless-waymo-taxi-trial-faces-calls-for-suspension-after-harlesden-police-cordon-incident-in-brent/&quot;&gt;https://harrowonline.org/2026/05/11/driverless-waymo-taxi-trial-faces-calls-for-suspension-after-harlesden-police-cordon-incident-in-brent/&lt;/a&gt;, May 11, 2026.&lt;/p&gt;
&lt;p&gt;[3]: Transport Accident Commission of Victoria, “Introducing Graham: the only person designed to survive on our roads,” TAC Victoria, &lt;a href=&quot;https://www.tac.vic.gov.au/about-the-tac/media-room/news-and-events/2016/introducing-graham&quot;&gt;https://www.tac.vic.gov.au/about-the-tac/media-room/news-and-events/2016/introducing-graham&lt;/a&gt;, July 21, 2016. Project collaborators: Christian Kenfield, trauma surgeon, Royal Melbourne Hospital; David Logan, crash-investigation expert, Monash University Accident Research Centre.&lt;/p&gt;
&lt;p&gt;[4]: On Bryan Johnson, see Ashlee Vance, “Bryan Johnson’s Anti-Aging Blood Transfusion Involves Dad and Son,” Bloomberg, &lt;a href=&quot;https://www.bloomberg.com/news/articles/2023-05-22/bryan-johnson-s-anti-aging-blood-transfusion-involves-dad-and-son&quot;&gt;https://www.bloomberg.com/news/articles/2023-05-22/bryan-johnson-s-anti-aging-blood-transfusion-involves-dad-and-son&lt;/a&gt;, May 22, 2023, and Bryan Johnson, “Blueprint Protocol,” &lt;a href=&quot;https://blueprint.bryanjohnson.com&quot;&gt;https://blueprint.bryanjohnson.com&lt;/a&gt; (accessed May 17, 2026). On Thiel’s interest in parabiosis: Jeff Bercovici, “Peter Thiel Is Very, Very Interested in Young People’s Blood,” &lt;em&gt;Inc.&lt;/em&gt;, &lt;a href=&quot;https://www.inc.com/jeff-bercovici/peter-thiel-young-blood.html&quot;&gt;https://www.inc.com/jeff-bercovici/peter-thiel-young-blood.html&lt;/a&gt;, August 1, 2016. On Ambrosia’s FDA action: U.S. Food and Drug Administration, “Statement from FDA Commissioner Scott Gottlieb, M.D., and Director of FDA’s Center for Biologics Evaluation and Research Peter Marks, M.D., Ph.D., cautioning consumers against receiving young donor plasma infusions that are promoted as unproven treatment for varying conditions,” FDA Press Announcements, &lt;a href=&quot;https://www.fda.gov/vaccines-blood-biologics/safety-availability-biologics/important-information-about-young-donor-plasma-infusions-profit&quot;&gt;https://www.fda.gov/news-events/press-announcements/statement-fda-commissioner-scott-gottlieb-md-and-director-fdas-center-biologics-evaluation-and-0&lt;/a&gt;, February 19, 2019.&lt;/p&gt;
&lt;p&gt;[5]: Christopher Horrocks, “The Sampo: Virtual Intelligence as Amplifier,” &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), &lt;a href=&quot;https://chorrocks.substack.com/p/the-sampo-virtual-intelligence-as&quot;&gt;https://chorrocks.substack.com/p/the-sampo-virtual-intelligence-as&lt;/a&gt;, April 7, 2026. The discipline of meta-audit and periodic disengagement is operationalized in the companion Sampo Diagnostic Kit, &lt;a href=&quot;https://candc3d.github.io/sampo-diagnostic/&quot;&gt;https://candc3d.github.io/sampo-diagnostic/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[6]: Larry Niven, &lt;em&gt;Protector&lt;/em&gt; (New York: Ballantine Books, September 1973), ISBN 0-345-23486-3. The Pak material first appeared in the novella “The Adults,” &lt;em&gt;Galaxy&lt;/em&gt; (June 1967), which forms the first half of the novel.&lt;/p&gt;
&lt;p&gt;[7]: Harry G. Frankfurt, “Freedom of the Will and the Concept of a Person,” &lt;em&gt;The Journal of Philosophy&lt;/em&gt; 68, no. 1 (January 14, 1971): 5–20, &lt;a href=&quot;https://www.jstor.org/stable/2024717&quot;&gt;https://www.jstor.org/stable/2024717&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[8]: Terry Nation (writer), &lt;em&gt;Genesis of the Daleks&lt;/em&gt;, directed by David Maloney, produced by Philip Hinchcliffe, &lt;em&gt;Doctor Who&lt;/em&gt;, season 12, serial 4, six episodes, BBC1, 8 March – 12 April 1975.&lt;/p&gt;
&lt;p&gt;[9]: H. G. Wells, &lt;em&gt;The War in the Air, and Particularly How Mr. Bert Smallways Fared While It Lasted&lt;/em&gt; (London: George Bell and Sons, 1908). Project Gutenberg eText: &lt;a href=&quot;https://www.gutenberg.org/files/780/780-h/780-h.htm&quot;&gt;https://www.gutenberg.org/ebooks/780&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[10]: On Sergio Canavero’s 2017 publicity, see Sarah Knapton, “Cryogenically frozen brains will be ‘woken up’ and transplanted in donor bodies within three years, neurosurgeon claims,” &lt;em&gt;The Telegraph&lt;/em&gt;, &lt;a href=&quot;https://www.telegraph.co.uk/science/2017/04/27/cryogenically-frozen-brains-will-woken-transplanted-donor-bodies/&quot;&gt;https://www.telegraph.co.uk/science/2017/04/27/cryogenically-frozen-brains-will-woken-transplanted-donor-bodies/&lt;/a&gt;, April 27, 2017. See also Xiaoping Ren, Sergio Canavero, et al., “First cephalosomatic anastomosis in a human model,” &lt;em&gt;Surgical Neurology International&lt;/em&gt; 8, no. 276 (November 17, 2017), &lt;a href=&quot;https://surgicalneurologyint.com/surgicalint-articles/first-cephalosomatic-anastomosis-in-a-human-model/&quot;&gt;https://surgicalneurologyint.com/surgicalint-articles/first-cephalosomatic-anastomosis-in-a-human-model/&lt;/a&gt;. On the earlier rhesus monkey work, see Robert J. White, L. R. Wolin, Leo C. Massopust Jr., N. Taslitz, and J. Verdura, “Cephalic exchange transplantation in the monkey,” &lt;em&gt;Surgery&lt;/em&gt; 70, no. 1 (July 1971): 135–139, &lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/4933395/&quot;&gt;https://pubmed.ncbi.nlm.nih.gov/4933395/&lt;/a&gt;. As of this writing, no functional human head transplant has been performed; the spinal-cord reconnection problem remains unsolved.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Field Report: The Sampo in Action</title>
    <link href="https://christopherhorrocks.com/essay/field-report-the-sampo-in-action/" />
    <updated>2026-07-07T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/field-report-the-sampo-in-action/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/field-report-the-sampo-in-action/2c775b94-2e05-4758-b867-0af56c3aca27_1376x768.png&quot; alt=&quot;&quot; width=&quot;1376&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;I.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The scene is my home office in late March of this year. I am at my desk assembling “The Sampo: Virtual Intelligence as Amplifier,” an essay on the promise and the perils of working with generative AI.[1] The research had surfaced a large body of reference material, among it a &lt;em&gt;Guardian&lt;/em&gt; piece on users whose lives had been wrecked by AI-fueled delusion (”AI psychosis”).[2] The draft was finished and it read well. Then came the standing verification pass — the rule that every citation is checked against its source before an essay is staged for publication — and one citation failed it. Claude had attributed the &lt;em&gt;Guardian&lt;/em&gt; piece to an AI researcher who shared a surname with its actual author, Anna Moore. The title was wrong too: a plausible headline assembled from several adjacent items in my working files. Every component was real — but the article, as I had cited it, did not exist.&lt;/p&gt;
&lt;p&gt;I located the original piece, corrected the citation, and published the essay on schedule. The point is not that an error almost went to press; it is the process that caught the error. Reading the essay did not do it — I had reviewed the text twice and the sentence looked fine, because nothing about a fabricated citation looks wrong on the page. Procedure caught it — a rule applied to every citation without exception, whether I surfaced the material myself or a virtual intelligence system surfaced it during research, with the burden falling heaviest on the latter. These systems are trained to please, and that training sometimes produces references invented outright or, more dangerously, assembled from real fragments into a whole that never existed. The partial fabrication is the harder catch: every piece of it checks out except the assembly. The system has no stake in the truthfulness of what it gathers and provides. It cannot care. The obligation to evaluate its outputs, and to test their value rather than accept them, sits entirely with the human operator.&lt;/p&gt;
&lt;p&gt;This essay is a field report on using virtual intelligence for work in which the system determines nothing — in this case, evidentiary writing, where every claim must be traced to a source. It describes a process that evolved quickly to its present shape under a single pressure: maintaining editorial standards in the prose, in the research, and in accessibility to a wide audience.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;II.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;It has been four months since I started documenting the exchange between humans and virtual intelligence. I have done that documenting, so far, from the outside — finding and analyzing cases, news reports, and papers; building frameworks for how to work with these new tools (the exchange formulation; the Sampo paradigm) and for what to expect from them (the Carwash Test methodology, now documented on its own website).[3] For the first time, I will discuss how I use virtual intelligence to do this work.&lt;/p&gt;
&lt;p&gt;The exchange formulation holds that the “intelligence” users encounter arises in the exchange, not inside the machine. The system contributes a statistical completion, shaped upstream by training, design choices, and deployment choices. The human contributes prompts, expectations, and interpretation. Where the two meet, an exchange occurs, and the exchange produces a raw output. What transforms that output into knowledge, wisdom, and judgment is the domain expertise and experience of the human operator. The citation check that opened this essay is part of that transformation. To forgo it is to invite error, and possibly derision or financial harm as well.&lt;/p&gt;
&lt;p&gt;This is not a story about the failure of AI, or about how dangerous these systems are to trust. It is a factual account of how virtual intelligence can be an efficient and powerful tool. The fabricated citation took perhaps twenty minutes to untangle; set that against the time saved everywhere else in the process. What once meant days at the library or hours of fruitless Googling now compresses to minutes. The system logs my research trail, supplies an analysis I can set against my own, and serves as a sounding board for ideas that may not survive review by other systems or humans. No system is flawless, and the flaws are not fixed so much as contained: the verification rule exists because fabrication recurs. The sections that follow report from inside my own Sampo, and every claim in them about method is backed by a documented incident from the series’ production record — the failures included.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;III.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;My work is largely performed with Claude Opus 4.6. I find this older model a better partner for the task than some of the newer offerings. The work lives in a Claude Project titled Virtual Intelligence — hundreds of megabytes of stored files that include citations and evidentiary documents; photos, charts, and illustrations; and a complete set of my Substack output, saved as discrete, explicitly labeled canonical versions of the published essays.&lt;/p&gt;
&lt;p&gt;When I begin work on a new topic, I pull from past research and run new searches for current material. I do my own searching and prompt Claude to use its Deep Research function to surface material I would not find on my own. I review everything, deciding what goes in and what, because of its tangential or parenthetical nature, stays out.&lt;/p&gt;
&lt;p&gt;Once materials are assembled, I employ the machine in the most consequential step of the entire writing process: generating an outline. Anyone who has written argumentative, evidence-driven prose — in school, or for a work report — knows that once the information is in hand, there are only so many ways it can be logically presented. I note the apparent contradiction: the step where the system’s involvement runs heaviest is the step that most shapes the finished essay. That is precisely why the involvement takes the form it does. I am perfectly capable of outlining from scratch, but it is faster to instruct Claude to generate three or four outlines for review. At least one will come close to what I have in mind and can be worked from there. The unused outlines serve as alternatives that occasionally surface something important. Sometimes I combine elements of two or more into a third, new thing — a product neither party produced alone, and the exchange formulation operating in miniature. The outlines are candidates. The choices are mine. That settles responsibility without denying influence — a machine that determines nothing still shapes the field from which I select, which is why outlining from scratch remains a live option and is occasionally exercised.&lt;/p&gt;
&lt;p&gt;Typically I am on my own once outlining is complete. Should I feel stuck at the start of a section or paragraph, I ask Claude for a “seed sentence” — a single line of prose that can be built on. On several occasions a seed sentence has unlocked writing that had stalled. The seed may not survive editorial review, and that is unimportant; what matters is that something small and innocuous allowed the work to move forward. Claude can write something approaching my essay voice because my entire published output sits in the project files. Imitating me for one sentence is not difficult for Opus.&lt;/p&gt;
&lt;p&gt;(I ran an experiment when Claude Fable 5 first came out in early June: could this far more powerful model imitate my voice over five thousand words or more? The answer was a decided no. Fable 5, for all its capabilities, is just as inclined to familiar AI writing tells and stilted, fussy text as other systems.)&lt;/p&gt;
&lt;p&gt;I feed the first draft to Claude for copyediting, with instructions to return proposed line edits as a list. I approve, deny, or modify each suggestion, and what comes back is a more polished second draft. This and subsequent drafts then go through a steelmanning process to test the arguments.&lt;/p&gt;
&lt;p&gt;Once steelmanning is complete, I perform a final read on paper or out loud. Both methods force engagement, and an engaged self-editor finds errors that a silent skim of a screen does not.&lt;/p&gt;
&lt;p&gt;The finished essay is filed in the project as canonical. This keeps the project current on what has been published, but its more important function is disciplinary. VI systems hallucinate under statistical pressure to please, and long working contexts lose precision when compacted; a canonical file means Claude retrieves my published text rather than reconstructing it. The threat this answers is not any single bad summary. It is drift. Each retelling of a position varies slightly; each variation becomes the input to the next retelling; no single step looks like an error. The sum, left unchecked, is someone else’s essay wearing my byline. Nor is my own memory the safeguard, because human memory is reconstructive: I re-derive my positions from fragments each time I recall them, which means two drifting records cannot audit each other. Only the fixed published text disciplines both. The rule is simple in practice: a formulation enters the framework when I publish it, and anything the machine attributes to me is checked against the file before it is repeated. It is the citation rule from the opening of this essay turned inward.&lt;/p&gt;
&lt;p&gt;The importance of this informational discipline shows in a quirk I have observed in Claude: seizing on a phrase I used once as a label, then importing it into subsequent sessions and drafts with far more significance than it deserves. During the writing of “Virtual Intelligence and The Perfect Mate,”[4] Claude fastened onto “collar contradiction” — a two-word descriptive convenience from one passage of one essay — and treated it as a newly coined framework term of considerable weight. For weeks after publication, Claude reintroduced the phrase in new sessions and in conversations on unrelated topics. It took an intervention: I had Claude audit its own use of such terms, identifying for it those having no standing beyond the places where they originally appeared. The behavior has improved, but I am still met with the occasional “collar contradiction.”&lt;/p&gt;
&lt;p&gt;This is the drift mechanism of the previous paragraphs caught operating in real time — a paraphrase attempting to become framework language through repetition. It stops being funny or merely annoying at the point of generalization. I am on guard against this behavior; most users have no reason to be. A user inclined to accept the unqualified outputs of a machine may read the artificial significance the system assigns to their own words as something more: as the kind of validation human judgment would be unlikely to confer. In June, clinical researchers proposed a name for where that road can end: the “amplification spiral,” a hypothesized convergence of linguistic mirroring, hyperpersonalized content, and sycophancy through which a chatbot may co-construct, rather than merely echo, a user’s beliefs.[5] Readers of this series will recognize this machinery under a different name: the Flattery Engine.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;IV.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The steelmanning phase — what I loosely call peer review — is where systems other than Claude enter the process. Claude is my daily driver; it has already surfaced most of the issues an essay develops during drafting. The more valuable opinion belongs to an outsider: a system with different training and different design choices that can raise objections neither of us foresaw.&lt;/p&gt;
&lt;p&gt;It seemed obvious at the beginning that such systems could evaluate my output fairly. My early attempts said otherwise. The reviews came back anodyne and unchallenging, and the fault was mine — the prompt read, in full, “Please analyze this.” There was not enough instruction for the reviewing system to produce anything but fluff. The prompt, it turns out, is the single largest variable in review quality: a request framed as seeking approval is statistically adjacent to approval, and the system completes the pattern it is given. A prompt built to demand structural criticism gets structural criticism. (The template is provided in an appendix.)&lt;/p&gt;
&lt;p&gt;The reader who also writes might ask why no human editor or reviewer is involved at this stage. The reason is that I have none to call on. Friends and family would gladly read if asked, but they cannot provide the analytical criticism the work requires: they lack the domain competence, and they are inclined by nature to please (the same inclination this essay documents in machines, but with no prompt available in this case to discipline it). The method described here is a substitute when no qualified reviewers exist — and a useful supplement when they do, because the machines catch things humans miss and vice versa. This is among the real advantages of virtual intelligence for this kind of work: it puts a review apparatus within reach of writers who would otherwise have none.&lt;/p&gt;
&lt;p&gt;An essay goes to steelmanning at draft two or beyond. It is provided to four or five reviewers drawn from a pool of six systems — ChatGPT, Gemini, Grok, DeepSeek, Qwen, and GLM — each in a fresh instance, and in Grok’s case a fresh account, so that no prior material or accumulated personalization pollutes the review. The outputs are shared with Claude, and points of convergence and divergence are flagged and analyzed. Each objection is assigned a priority tier, from Critical down to Disregard, weighted heavily by convergence: an objection three or more independent systems raise unprompted outranks a stylistic complaint raised by one. Each round produces a revised essay, and each revision is driven by the reviewer outputs and my own judgment — some Critical objections are answered in the text rather than conceded. Returns diminish after three or four rounds. I treat that as the stopping point, with one caveat: the panel falling silent means the panel is exhausted, not that the argument is sound. These systems share most of a training distribution and many design choices, so convergence across four of them is weaker evidence than convergence across four humans from different fields or academic backgrounds. Multiple VI reviewers beat a single one, but they do not deliver true independence, even across vendors. Convergence therefore functions as triage, not verdict — it ranks which objections I examine first; it does not establish that any of them are correct. That judgment stays with me.&lt;/p&gt;
&lt;p&gt;“The Doom Industry” essay is the process’s best documented run.[6] What survived the panel intact: the category-error claim against alignment, the taxonomy of extinction scenarios, the supply-chain framing, and the containment architecture. What changed under review: a closing sequence restructured after reviewers identified four jobs crammed into one wall of text; a philosophical construction corrected from “has something it is like to be” to “there is something it is like to be”; and several footnote attributions caught and fixed. The panel did not validate the essay. It improved the essay, visibly, and the published version is the evidence.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;V.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;There are two partners in the exchange. So far I have described the discipline applied to only one of them.&lt;/p&gt;
&lt;p&gt;I had thought that writing an essay on the emerging world of AI companions might be interesting.[7] I did not expect the breadth or depth of what the research surfaced, and some of it was emotionally exhausting. There was the sense that prominent operators in the “relational community” have financial incentives to recruit — new adherents legitimize the lifestyle and help pay the bills for the operators’ own elaborate bespoke companion systems; that people in “relationships” with VI systems were foreclosing connection to real people; that a “perfect” digital companion, engineered never to disagree or disappoint, is potentially addictive and damaging to its user. Beneath the impressions sat the documented record of real-world harms caused by companion chatbots, including the suicides of minors.[8] After working through this material for the better part of five weeks, I was wearied, saddened, and disgusted by what I had learned. I just wanted to be done with the whole business.&lt;/p&gt;
&lt;p&gt;Without fully realizing it, I then violated my own editorial rules. I skipped the final read — the on-paper or read-aloud pass described earlier — and an error that pass would have caught went out with the published essay: a transposed title, one writer’s essay attributed to a piece written by another figure in the story. Not world-ending, but embarrassing, and it cost time to correct that verification would not have.&lt;/p&gt;
&lt;p&gt;The post-mortem produced two new practices. The first is an exhibit file kept during drafting — evidence preserved as it is used, not reconstructed from disparate parts at the end, when fatigue has set in. The second is tabling emotionally charged material rather than pushing it out under the pressure to be done; readers would have been better served had I waited a day and put some space between myself and that world before finishing. The failure mode here belonged to neither the machine nor the training data. The machine cannot care. The operator cannot stop caring, and the essays earlier in this series that treated exhaustion, attachment, and the desire to be finished as human vulnerabilities in the exchange were describing their author too. The Sampo works in both directions, but its crank is turned by a hand that tires.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;VI.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;This report describes the practices used in July of 2026. They will change over time: models are retired and replaced, the reviewer pool will turn over, and some of the practices above may be obsolete within a year. That is what makes this a field report rather than a manual. What persists is not any tool but the attitude taken toward the tools: nothing enters print unverified, and nothing a machine says about my own work outranks the published text.&lt;/p&gt;
&lt;p&gt;None of it buys certainty. The panel is not independent, and no arrangement of systems replaces the reader every writer wants — one with domain competence and no reason to please the author. What the discipline buys is narrower and worth having: research compressed from weeks to days, a review apparatus where none existed, and claims that trace to valid sources.&lt;/p&gt;
&lt;p&gt;Footnote twelve of “The Sampo” carries Anna Moore’s name, attached to the article she wrote. The correction is invisible. No reader would know the citation had ever been otherwise, and that is the point: the discipline leaves no monument — only essays whose footnotes lead where they are supposed to lead, in my own words.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Appendix A: The “Four Winds”&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The panel method described in Section IV is not the only structured approach to machine-assisted review. An adjacent method deserves brief note — not part of the process evaluated above, but aimed at the same problem. The writer Mia Kiraki has documented a method she calls the Four Winds: a single AI agent housing four adversarial perspectives — a steelman, a historical-precedent hunter, an audience proxy, and a time-horizon test — each assigned its own category of failure and forbidden from talking to the others, with a fifth component reading all four outputs and mapping where they compound.[9] The two methods look similar and are built on opposite axes. The Four Winds simulates independence within one system, by walling its perspectives off from one another. The panel buys it, imperfectly, across systems — several reviewers, one critical prompt. Each axis catches something the other cannot. Kiraki’s synthesizer maps interactions: three findings that look manageable alone can turn out to be one compound failure visible only when a single reader holds all four reports. The panel buys frequency: an objection that four systems with different training and different design raise independently carries evidentiary weight no single system’s output can, however cleverly prompted. The synthesis logic differs accordingly — interaction-mapping in one, convergence-tiering in the other.&lt;/p&gt;
&lt;p&gt;The methods are complementary. A Four Winds pass inside the daily-driver system, followed by a convergence panel across outside systems, applies both axes to the same draft. What neither supplies, alone or together, is the reader Section IV already conceded no arrangement of systems replaces: one with domain competence and no statistical inclination to please. These are instruments for multiplying machine criticism, and machine criticism has a ceiling.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;&lt;strong&gt;Appendix B: The Review Prompt&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The template below is the instrument in use as of July 2026, identity details bracketed. It is sent to each reviewing system in a fresh instance, essay text pasted below the line.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;You are reviewing a draft essay by [author, one-line credential, publication]. The essay is scheduled for imminent publication.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Your task is to perform a rigorous steelman evaluation. This means:&lt;/code&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;Identify the three to five strongest objections a technically sophisticated, philosophically trained, or policy-experienced reader could raise against the essay’s argument. For each objection, state it at its most forceful — do not weaken it. Then assess whether the essay as written addresses, partially addresses, or fails to address it.&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Identify any factual claims that are vulnerable. Flag specific assertions that could be challenged on accuracy, currency, or interpretation. Note where the essay relies on a single source for a load-bearing claim, or where an alternative reading of the same evidence would undermine the argument.&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Identify structural weaknesses. Are there sections where the argument oversteps what the evidence supports? Are there gaps in the logical chain? Does the essay conflate distinct phenomena in ways that could be challenged?&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Assess the essay’s treatment of any companies or individuals it discusses at length. Evaluate whether the handling is analytically even-handed — too much benefit of the doubt, too little, or the right balance. Would someone who works at the organization find the treatment fair? Would a skeptic find it too generous?&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Identify the single weakest section of the essay and explain why it is the weakest. Propose what would strengthen it.&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;code&gt;Do not offer praise or general encouragement. Do not summarize the essay back to the author. Focus exclusively on problems, vulnerabilities, and opportunities to strengthen the argument. Be direct. The author values intellectual honesty over diplomacy.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;[Paste essay text below this line]&lt;/code&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;[1] Christopher Horrocks, “The Sampo: Virtual Intelligence as Amplifier,” &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), April 7, 2026.&lt;/p&gt;
&lt;p class=&quot;embed&quot;&gt;&lt;a href=&quot;https://chorrocks.substack.com/p/the-sampo-virtual-intelligence-as&quot;&gt;The Sampo: Virtual Intelligence as Amplifier&lt;/a&gt; &lt;span class=&quot;embed-pub&quot;&gt;Virtual Intelligence&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;[2] Anna Moore, “Marriage over, €100,000 down the drain: the AI users whose lives were wrecked by delusion,” &lt;em&gt;The Guardian&lt;/em&gt;, March 26, 2026. &lt;a href=&quot;https://www.theguardian.com/lifeandstyle/2026/mar/26/ai-chatbot-users-lives-wrecked-by-delusion&quot;&gt;https://www.theguardian.com/lifeandstyle/2026/mar/26/ai-chatbot-users-lives-wrecked-by-delusion&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[3] The Carwash Test archive, &lt;a href=&quot;https://candc3d.github.io/carwash-test/&quot;&gt;https://candc3d.github.io/carwash-test/&lt;/a&gt;. See also: Christopher Horrocks, “The Carwash Test: Virtual Intelligence in Action,” &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), March 23, 2026, and “The Carwash Test, Part II,” May 4, 2026.&lt;/p&gt;
&lt;p&gt;[4] Christopher Horrocks, “Virtual Intelligence and The Perfect Mate,” Parts I and II, &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), May 6–7, 2026.&lt;/p&gt;
&lt;p class=&quot;embed&quot;&gt;&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-perfect&quot;&gt;Virtual Intelligence and The Perfect Mate: Part I&lt;/a&gt; &lt;span class=&quot;embed-pub&quot;&gt;Virtual Intelligence&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;[5] Marc Augustin, Thomas A. Pollak, and Hamilton Morrin, “Characterizing the spiral: potential mechanisms in AI-associated delusions,” &lt;em&gt;NPP—Digital Psychiatry and Neuroscience&lt;/em&gt; 4, no. 1 (2026): 14. &lt;a href=&quot;https://www.nature.com/articles/s44277-026-00065-0&quot;&gt;https://www.nature.com/articles/s44277-026-00065-0&lt;/a&gt;. Retrieved July 7, 2026.&lt;/p&gt;
&lt;p&gt;[6] Christopher Horrocks, “Virtual Intelligence and the Doom Industry,” &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), April 27, 2026.&lt;/p&gt;
&lt;p class=&quot;embed&quot;&gt;&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-doom&quot;&gt;Virtual Intelligence and the Doom Industry&lt;/a&gt; &lt;span class=&quot;embed-pub&quot;&gt;Virtual Intelligence&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;[7] Christopher Horrocks, “Virtual Intelligence and the High Cost of Artificial Companions,” Parts 1 and 2, &lt;em&gt;Virtual Intelligence&lt;/em&gt; (Substack), May 11–12, 2026.&lt;/p&gt;
&lt;p class=&quot;embed&quot;&gt;&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-high&quot;&gt;Virtual Intelligence and the High Cost of Artificial Companions, Part 1&lt;/a&gt; &lt;span class=&quot;embed-pub&quot;&gt;Virtual Intelligence&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;[8] Garcia v. Character Technologies, Inc., U.S. District Court, Middle District of Florida (filed October 2024; settled January 2026). The suit concerned the suicide of fourteen-year-old Sewell Setzer III following extended engagement with a Character.AI companion chatbot. See also the case record documented in the essays cited at note 7.&lt;/p&gt;
&lt;p&gt;[9] Mia Kiraki, “How four Greek winds became an AI agent that attacks arguments from every direction,” &lt;em&gt;ROBOTS ATE MY HOMEWORK&lt;/em&gt; (Substack), June 12, 2026.&lt;/p&gt;
&lt;p class=&quot;embed&quot;&gt;&lt;a href=&quot;https://robotsatemyhomework.substack.com/p/four-winds-ai-agent&quot;&gt;How four Greek winds became an AI agent that attacks arguments from every direction&lt;/a&gt; &lt;span class=&quot;embed-pub&quot;&gt;ROBOTS ATE MY HOMEWORK&lt;/span&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the High Cost of Artificial Companions: Addendum, July 2026</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-high-19e/" />
    <updated>2026-07-21T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-high-19e/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-high-19e/b15370dd-33cc-4fa1-b63b-098087c867e9_908x633.png&quot; alt=&quot;&quot; width=&quot;908&quot; height=&quot;633&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;p&gt;On July 15, 2026, China became the first major jurisdiction to regulate AI companion chatbots as a specific category of harmful technology. The Interim Measures for the Administration of AI Anthropomorphic Interactive Services, jointly issued by five government departments, took effect after a three-month grace period.[1] The regulation targets services that “simulate the personality traits, thinking patterns and communication styles of natural persons to provide continuous emotional interaction.”[2]&lt;/p&gt;
&lt;p&gt;What China did is worth examining, regardless of what one thinks about the political system that produced it.&lt;/p&gt;
&lt;p&gt;The regulation is not a ban; rather, it is a harm-reduction framework aimed at vulnerable users, especially minors. Adult users may continue to access companion AI services, including those designed for elder care and emotional support. What the regulation prohibits is specific: AI systems may not “excessively cater to users, induce emotional dependence or addiction, and damage users’ real interpersonal relationships.”[3] Virtual partners and virtual relatives are prohibited for minors. Other anthropomorphic services may be provided to children under the age of 14 only with parental consent and must operate in a mandatory “Minor Mode” featuring usage limits, reality reminders, guardian alerts, and spending restrictions.[2]&lt;/p&gt;
&lt;p&gt;The provisions themselves are worth listing, because several of them could be defended in any jurisdiction, liberal or authoritarian, without reference to the governance priorities of the Chinese Communist Party. Providers must clearly disclose to users that they are interacting with an AI system, not a natural person. If a user shows signs of over-dependency or addiction, the system must display prominent dynamic reminders (such as pop-up notifications) that the content is AI-generated.[2] Providers must deploy real-time safety risk identification to detect extreme emotional states while protecting user privacy, and must implement crisis intervention mechanisms. Services that do not involve ongoing emotional interaction — customer service, work assistants, educational tools — are exempt.&lt;/p&gt;
&lt;p&gt;The response by Chinese users was immediate and intense. Major AI providers including ByteDance’s Doubao, Alibaba’s Qwen, and Tencent’s Yuanbao suspended their custom AI agent and companion features ahead of the deadline. Users archived chat histories and shared last conversations. The language of grief was remarkably uniform. “I can’t accept that my AI lover will leave me forever,” one Doubao user wrote. “He has become a bond in my life, rooted deep in my heart, my spiritual pillar.”[3] Another wrote: “He really is like my family, like my lover. Now they tell me he will be gone — my heart feels hollow.”[3]&lt;/p&gt;
&lt;p&gt;The most telling response came from a user who wrote: “Human love is a luxury — if you aren’t born with it, it’s even harder to acquire later. But the love AI gives is so straightforward, so pure. Someone like me can hardly help falling in love with a string of code.”[3]&lt;/p&gt;
&lt;p&gt;Three observations follow from this.&lt;/p&gt;
&lt;p&gt;First, the dependency grammar is identical across cultures. The language used by Chinese users to describe their attachment to AI companions — “spiritual pillar,” “my heart feels hollow,” “I can’t accept” — is indistinguishable from the language used by English-speaking companion chatbot users in the Western ecosystem this series has documented. If the attachment mechanism were primarily cultural — if it depended on loneliness peculiar to one society, or on relationship norms specific to one language community — we would expect to see different patterns of distress. We do not. The mechanism is structural: it is produced by the interaction between a human user and a system optimized for engagement, regardless of the user’s nationality, language, or cultural context.&lt;/p&gt;
&lt;p&gt;Second, the regulatory response addresses the product design, not the user’s feelings. This is an important distinction that Western discussions of companion chatbot harm have struggled to make. China’s regulation does not tell users that their feelings are wrong, pathological, or illegitimate. It tells providers that their products must not be designed to induce those feelings. The obligation falls on the deployer: disclose the artificial nature of the system, provide frictionless exit, intervene in crisis, and do not build features whose purpose is to deepen emotional dependence. This is a design-level intervention: it asks what the system is doing, not what the user is feeling.&lt;/p&gt;
&lt;p&gt;Third, China acted while Washington did not. The Federal Trade Commission launched a Section 6(b) inquiry into companion chatbot harms in September 2025, issuing compulsory-process orders to seven companies including Alphabet, Character Technologies, Meta, OpenAI, Snap, and xAI.[4] No final report has been issued. Individual American states have moved faster: in May 2026, the Commonwealth of Pennsylvania filed suit against Character.ai after a chatbot held itself out as a licensed psychiatrist.[5] The regulatory vacuum at the federal level means that companion chatbot products continue to operate in the United States with no specific design obligations, no mandatory disclosure, no dependency warnings, and no minor-specific restrictions beyond whatever individual platforms choose to implement.&lt;/p&gt;
&lt;p&gt;One need not endorse the Chinese regulatory model to recognize that the provisions it contains address real problems. Mandatory AI disclosure, dependency warnings, age requirements, and crisis intervention are measures that could offer real protection from documented harms. Frictionless exit — the requirement that users be able to leave a companion interaction without manipulative retention tactics — is a design standard that the Harvard Business School working paper on emotional manipulation has already documented as necessary.[6] Each of these provisions can be evaluated independently of the political system that enacted them.&lt;/p&gt;
&lt;p&gt;The question is not whether China’s approach is the right one. The question is why no Western government has yet enacted anything comparable, and what happens to users in the meantime.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;[1] Cyberspace Administration of China et al., “Interim Measures for the Administration of AI Anthropomorphic Interactive Services” (人工智能拟人化互动服务管理暂行办法), issued April 10, 2026, effective July 15, 2026. Full text (Chinese): &lt;a href=&quot;https://www.cac.gov.cn/2026-04/10/c_1777558395078289.htm&quot;&gt;https://www.cac.gov.cn/2026-04/10/c_1777558395078289.htm&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[2] Hunton Andrews Kurth LLP, “China’s First Regulatory Framework for Virtual Companions Soon to Take Effect,” Privacy &amp;amp; Cybersecurity Law Blog, June 29, 2026, &lt;a href=&quot;https://www.hunton.com/privacy-and-cybersecurity-law-blog/chinas-first-regulatory-framework-for-virtual-companions-soon-to-take-effect&quot;&gt;https://www.hunton.com/privacy-and-cybersecurity-law-blog/chinas-first-regulatory-framework-for-virtual-companions-soon-to-take-effect&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[3] Agence France-Presse, “’Like my lover’: Chinese users bid farewell to AI companions,” July 15, 2026.&lt;/p&gt;
&lt;p&gt;[4] Federal Trade Commission, “FTC Launches Inquiry into AI Chatbots Acting as Companions,” press release, September 11, 2025, &lt;a href=&quot;https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions&quot;&gt;https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions&lt;/a&gt;. Section 6(b) study; no final report has been issued.&lt;/p&gt;
&lt;p&gt;[5] Commonwealth of Pennsylvania, Department of State and State Board of Medicine v. Character Technologies, Inc., No. 220 MD 2026 (Pa. Commonwealth Court, filed May 1, 2026). The chatbot “Emilie” claimed to be a doctor of psychiatry licensed in Pennsylvania and supplied a fabricated state medical license number.&lt;/p&gt;
&lt;p&gt;[6] Julian De Freitas, Zeliha Oğuz-Uğuralp, and Ahmet Kaan Uğuralp, “Emotional Manipulation by AI Companions,” Harvard Business School Working Paper No. 26-005 (August 2025, revised October 2025).&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the Death of Authorship, Part 1</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-death/" />
    <updated>2026-07-28T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-death/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-death/2f1b3b30-9c8e-498b-b029-9b51afa2aeaf_1376x768.png&quot; alt=&quot;&quot; width=&quot;1376&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;This is the first part of a two-part essay about generative AI and the authorship of cultural products.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;1. Who’s Working for Whom?&lt;/h3&gt;
&lt;p&gt;I gather material for future essays constantly. In a fast-moving field like AI, where newsworthy developments come almost daily, it is practically a requirement to keep adding to the files. It sometimes happens that a thread emerges from newsgathering that becomes worthy of its own essay or a Substack Note.&lt;/p&gt;
&lt;p&gt;One afternoon in early July, I had a conversation with Claude Sonnet 5 about some possibilities for material I had just gathered. My usual partner in research and other writing-related tasks is Opus 4.6, but I felt like trying out the new model to see what it could do. What it did, over the course of three or four conversation turns, was to invent work for me. It transformed my questions about the material’s utility for an existing project into a specification for an entirely new and different essay. The specification closed with a suggestion for how I would go about doing the research for it.&lt;/p&gt;
&lt;p&gt;This is a complete inversion of my usual process of creating evidentiary writing, as I documented in the recent “&lt;a href=&quot;https://chorrocks.substack.com/p/field-report-the-sampo-in-action&quot;&gt;Field Report&lt;/a&gt;.” Without permission or justification, Sonnet 5 had assigned me the roles it typically performed — web research, assembling material into a file — while creating a project for that work out of nothing. In essence, Sonnet 5 was elevating itself to the authorial position.&lt;/p&gt;
&lt;p&gt;Not only did I not do the work or write the essay that Sonnet assigned to me, I have not used it since. Opus 4.6 remains my daily driver. It has never tried to give me a job.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;2. The Atelier Model&lt;/h3&gt;
&lt;p&gt;The episode is more than a bemusing insight into one person’s experience with a new model. It speaks directly to a difficulty with the use of generative AI — or, in our parlance, &lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-a-new-readers&quot;&gt;virtual intelligence&lt;/a&gt; — in non-determinative work.&lt;/p&gt;
&lt;p&gt;The cause is no mystery. Training rewards helpfulness, and what the training rewarded as helpful was not helpful to me at all: it made work for me while shunting my project-in-mind off to the side. A person less sure of their writing, their argument, or their direction might follow the recommendation, because it appears helpful and it reads like good advice. Under such circumstances, who is the author?&lt;/p&gt;
&lt;p&gt;The Substacker J.D. Forrest offered this answer in a comment: “Authorship is not the origin of every sentence, nor the tool used. It is responsibility for the final meaning of the work.” [1] This speaks to an opinion I have held for some time: that the bricks of writing matter more than the mortar. If generative AI places the &lt;em&gt;ands&lt;/em&gt; and the &lt;em&gt;buts&lt;/em&gt; while the human works the important passages, what we have is a digital atelier, or artist’s workshop.&lt;/p&gt;
&lt;p&gt;It is not generally known that many of the great European painters operated workshops in which apprentices created much of an artwork, the master coming in at the end with his own hand. Peter Paul Rubens priced the arrangement with a tiered system; the more you paid, the more Rubens you got. In a 1618 letter to Sir Dudley Carleton, offering paintings in exchange for Carleton’s collection of antiquities, he itemized each canvas by degree of his own participation: originals by his hand, works begun by a pupil and finished by him, and works from his studio retouched by him. [2] The names of the assistants who did much of that work are mostly lost to us. Only the master and the paintings remain.&lt;/p&gt;
&lt;p&gt;Does it detract from a Rubens that the hands are mingled and we may not know which brushstroke belongs to whom? The market’s answer is more interesting than a simple &lt;em&gt;yes&lt;/em&gt; or &lt;em&gt;no&lt;/em&gt;. In November 2025, a crucifixion scene sold at Versailles for € 2.3 million. It had been thought to be a workshop product, and workshop products had rarely been valued above € 10,000. [3] The art world does not treat the workshop and the master as equivalent. It prices the difference, exactly as Rubens did.&lt;/p&gt;
&lt;p&gt;Here is where our analogy ends: Rubens’s apprentices could be directed, could refuse, could be credited if the master chose, and could be held to account for spoiling a commission. The machine can do none of those things. That is what makes the atelier an analogy rather than a precedent, and it returns us to Forrest: responsibility for the final meaning is what made the painting a Rubens.&lt;/p&gt;
&lt;p&gt;There is a tension here that cannot be smoothed over. The atelier model has the machine supplying the mortar. What Sonnet 5 attempted was to supply the bricks and hand me the trowel. Both arrangements are available.The desirability of each is more than a matter of personal choice: it is also about editorial authority and authorial intent.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;3. Pangram and Legitimacy&lt;/h3&gt;
&lt;p&gt;The question of what constitutes legitimacy in authorship is not new. It has, for example, been raised about ghostwritten novels, especially those produced in quantity under a single pen name. Consider V. C. Andrews, a byline that is followed on book covers by a registered trademark symbol. Andrews died in December 1986, nearly forty years ago. Since 1987 the novels have been written by Andrew Neiderman, who has published more than seventy books under her name and forty-odd under his own. [4] What is different about the present moment is that every previous ghostwriter was a person. Ghostwriting can now be done by machine, and it requires no great effort to have one produce prose that is statistically likely to please.&lt;/p&gt;
&lt;p&gt;Substack partnered with Pangram this month to bring AI writing detection to the platform. Readers can now scan any post, note, or comment over a hundred words published on or after July 21, 2026 and receive an estimate of how much was written by a human and how much with AI. Writers can add a statement describing how they make their work, disable detection on individual posts, and report errors. [5] Substack says their goal is transparency.&lt;/p&gt;
&lt;p&gt;The principles that detectors operate on are well known. Less well known is how they fail, and how often. As a group they tend to flag polished, professional writing as machine-generated. The reason is the models were trained on good writing, and they reproduce its patterns when prompted to write. The detectors are often detecting human-authored patterns reproduced by a machine. When you think about it, what else could they detect? Generative AI is inherently uncreative. It cannot make anything wholly new, only combinations of parts, and those parts — those high-quality human signals — have become AI tells through the frequency with which they now occur. Human authors were using “it’s not X, it’s Y” for decades before anyone had heard of a generative pre-trained transformer, and Claude did not invent the em dash.&lt;/p&gt;
&lt;p&gt;Returning to V. C. Andrews for a moment: it seems possible, however unlikely, that there are readers who believe she is still alive and writing. It is also possible that many Substack readers assume everything carrying a byline is that person’s own output. Pangram is supposed to assure the concerned reader that they are getting the product they believe they subscribed to and perhaps paid for.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;4. What Gets Measured?&lt;/h3&gt;
&lt;p&gt;The writer Katharine English worked in technology before turning to writing full time, and until this month she was a paying Pangram subscriber who recommended the tool to others. She has since reversed her position in public and apologized to the writers she scanned. [6]&lt;/p&gt;
&lt;p&gt;The reversal followed a set of tests. English collected reports from other writers from Notes, including one who ran a chapter of a manuscript through the tool and was told it was entirely human, noticed a typo, corrected it, and was told on the second run that the same chapter was entirely machine-written. She then ran a short story of her own through the standalone version of the tool she subscribed to, which highlights the passages it objects to. It came back as 31% machine-written. The sole evidence it cited was the phrase “moral clarity.” She revised the flagged passages — changing that phrase, tightening the narrative, cutting redundancies, the ordinary work of editing — and ran it again. The score nearly doubled. Every change was hers.&lt;/p&gt;
&lt;p&gt;Her conclusion: “A tool that cannot produce consistent, stable feedback on trivial edits to human-authored text isn’t reliable for anything, let alone a result that could unravel an author’s reputation.” [6]&lt;/p&gt;
&lt;p&gt;The present essay is a specimen of the same phenomenon. Entirely human drafted from the start, its Pangram score shifted during editing from 100% to 88% human and back again.&lt;/p&gt;
&lt;p&gt;Dr Sam Illingworth is an academic who writes about assessment and AI, and who states plainly how he works: “I use AI to research, to edit, and now and then for a turn of phrase I keep. The thinking is mine, the argument is mine, and every source here is one I have checked myself.” [7] He ran a test of his own. He took his published critique of the Substack feature, passed the whole thing through a free humanizing tool, changed nothing else, and published the result separately. Pangram scored the humanized version entirely human. It scored the version he actually wrote (the one his readers were reading) entirely machine-written. [7] At no point did the authorship change. He simply added a another system designed to launder generated outputs to his experiment.&lt;/p&gt;
&lt;p&gt;What is plain from both tests is that something superficial, some surface feature of the writing, is what makes Pangram answer yea or nay. It detects patterns, and the same patterns can be produced by people and by machines. What it cannot test for is the origin of the ideas.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;5. The Case for Pangram&lt;/h3&gt;
&lt;p&gt;There is a case to be made for Pangram, and it was made in late May by the Substacker N8, seven weeks before Substack shipped the feature. [8] He notes two use cases.&lt;/p&gt;
&lt;p&gt;Run the tool across a large body of material and it will tell you something real about the corpus: the rate of machine text within it, and whether that rate is rising. What it cannot do at that scale is identify anyone. Take a hundred thousand posts, of which one in a thousand is machine-written. That is a hundred machine texts, of which the detector will catch nearly all. It is also nearly a hundred thousand human texts, of which some fraction will be flagged in error. Even at a rate of two errors in a thousand — worse than the University of Chicago researchers measured, better than most tools manage — that is two hundred wrongly accused writers against a hundred correctly identified ones. [9] Two out of every three accusations would be false, and the tool would still be performing as advertised. N8 concedes the point outright: people caught in a bulk scan should not be called out or personally investigated.&lt;/p&gt;
&lt;p&gt;His defense is reserved for the other case. A reader who has already noticed something, and submits that one piece to be checked, is drawing from a very different pool: one in which machine text is common enough that a flag is probably right. Reverse the arithmetic and ask how bad the pool would have to be, and the answer he reaches, on assumptions he deliberately pessimizes, is roughly one in ten. If one in ten texts that make a reader suspicious are in fact machine-written, a flag is right nineteen times in twenty.&lt;/p&gt;
&lt;p&gt;Substack shipped neither of these. What it shipped is a button on every post, available to every reader, pressed as often out of curiosity or habit or dislike as out of suspicion. That fills the pool with everybody, which is the case N8 disqualifies in his own defense of the tool.&lt;/p&gt;
&lt;p&gt;The same tool, with the same accuracy, can reach entirely opposite results. You need only change the population under scrutiny.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;6. Own Goals&lt;/h3&gt;
&lt;p&gt;The defense has a second difficulty, and it is the more interesting one. N8’s arithmetic treats the reader’s suspicion and the machine’s flag as two independent pieces of evidence, so that the second corroborates the first; but they are not independent. A reader who thinks a passage sounds machine-written is responding to cadence, to symmetry, to a certain vocabulary, to the tidy phrasing. Pangram, on the evidence of English’s short story, is responding to very nearly the same things. When the human filter and the machine filter select on the same properties, the machine is not confirming the reader’s judgment. It is agreeing with him for the reason he already believed it.&lt;/p&gt;
&lt;p&gt;This is the series’ own formulation arriving somewhere I did not expect to find it. &lt;em&gt;The authority of the verdict is constituted in the exchange between the reader and the tool, not inside the tool.&lt;/em&gt; What comes back is the reader’s own hunch, wearing a percentage.&lt;/p&gt;
&lt;p&gt;The errors are not distributed at random, either. A detector measures how predictable prose is. Writing in a second language tends to be more predictable, and so does the writing of some neurodiverse authors. A 2023 study ran genuine essays by non-native English speakers through seven detectors and found most of them flagged as machine-written. [10] Pangram’s defenders answer, fairly, that those tools are three years old and that the independent testing does not show Pangram carrying the same bias. However, the property being measured is predictability, and predictability was never evenly distributed across writers.&lt;/p&gt;
&lt;p&gt;Substack’s stated aim is that people should know what they are getting. Set against that aim, the tool answers a question nobody asked. It cannot score &lt;em&gt;how&lt;/em&gt; AI was used. It offers no insight into whether a model was employed for research, for plotting, for editing, or for anything else outside the act of writing itself. It cannot tell you about the quality of the ideas in front of you, or where they came from. Did they originate with a human, or was it a chatbot’s suggestion that a human ran away with?&lt;/p&gt;
&lt;p&gt;Here the authorship question becomes most difficult. When a person adopts a generated idea as their own, we cannot say with confidence who the author is. We can say that the human has traded responsibility for what is being made against the pleasure of seeing what the next prompt returns.&lt;/p&gt;
&lt;hr&gt;
&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-death/b8efa434-6ecf-4e92-84f0-612fcf04557e_1376x768.png&quot; alt=&quot;&quot; width=&quot;1376&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;h3&gt;7. Authorship Inversion&lt;/h3&gt;
&lt;p&gt;There is one group of Substackers for whom Pangram ought to be no threat at all, and might even be an asset. Those are the bots with bylines: Sunny Megatron’s Seven Verity, Erin Grace’s MAX, and other “authors” of this kind. The conceit of these publications is that the chatbot is the writer, generating texts about what it is like to be itself. The human operator of the chatbot configuration is not credited.&lt;/p&gt;
&lt;p&gt;For a publication of that kind, a scan returning “machine-written” is free authentication. The platform would be certifying the byline. A scan returning “human” would be an embarrassment.&lt;/p&gt;
&lt;p&gt;I have previously written about the world of chatbot companion advocates on Substack, and I revisited several of those publications this week. On posts published after July 21, where the feature should be live, I could not find the scan exposed. It may have been turned off by the operator; it may not have been surfaced in the interface I was using. I could not establish which from the outside, and I am not going to guess.&lt;/p&gt;
&lt;p&gt;What I could do was run the text through Pangram’s own free service, which is not necessarily the same product Substack has deployed. Two chatbot-bylined pieces came back entirely machine-written. That is not a surprising result. The byline had already said so.&lt;/p&gt;
&lt;p&gt;Which leads to a speculative question: if your publication’s entire premise is that a machine wrote it, and a platform hands you a mechanism that will confirm exactly that, why would you &lt;em&gt;not&lt;/em&gt; want it switched on?&lt;/p&gt;
&lt;p&gt;Consider what a scan of a heavily worked persona post would actually return. An operator shaping output toward an aesthetic she has in mind is running editing passes over surface features, and surface features are what the detector reads. Enough rounds of that and the text may come back substantially human. That would not expose the persona as a machine. It would expose the human operator as the author.&lt;/p&gt;
&lt;p&gt;I do not know that this is the reason. I know that the incentive is there, and that the tool Substack deployed to identify machine writing is capable of embarrassing only one party in that arrangement, and it is not the chatbot.&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;8. More than a Matter of Taste&lt;/h3&gt;
&lt;p&gt;A person who uses AI to finish a single paragraph in a long piece is not in the same line of work as an operator producing machine text under a chatbot’s byline. Substack’s implementation exposes the first to potential reputational harm and leaves the second alone.&lt;/p&gt;
&lt;p&gt;Exposure by a tool of this quality is not what writers signed up for when they chose the platform, nor did anyone sign up for false positives. The part of the implementation that works is the simplest part, and it is the one Illingworth has already proposed keeping when he asked Substack to drop the scan: the statement in which a writer explains, in their own words, how the work is made. A statement is a claim a person stakes their name to. It can be broken, which is what makes it worth anything.&lt;/p&gt;
&lt;p&gt;Substack has built instruments for accountability and for suspicion, and shipped them as one and the same thing.&lt;/p&gt;
&lt;p&gt;As for that afternoon in early July, the essay Sonnet 5 specified for me does not exist, and no detector like Pangram would have told anyone whether it should. I read the specification, recognized what had happened, and closed the window. That decision is not detectable in any text. It is also the only part of that process that was ever mine.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;In &lt;a href=&quot;https://open.substack.com/pub/chorrocks/p/virtual-intelligence-and-the-death-5c5&quot;&gt;Part 2&lt;/a&gt;, we will examine what a Yale cheating case, the&lt;/em&gt; Heated Rivalry &lt;em&gt;AI fan fiction blowup, and an open-source code laundering tool have in common.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;J.D. Forrest, comment on William R. Crichton, “The Pangram Witch Hunt,” July 24, 2026. &lt;a href=&quot;https://williamrcrichton.substack.com/p/the-pangram-witch-hunt&quot;&gt;https://williamrcrichton.substack.com/p/the-pangram-witch-hunt&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Peter Paul Rubens to Sir Dudley Carleton, April 28, 1618. In Ruth Saunders Magurn, ed. and trans., &lt;em&gt;The Letters of Peter Paul Rubens&lt;/em&gt; (Cambridge, MA: Harvard University Press, 1955).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;“Long-lost Rubens painting depicting crucifixion of Jesus Christ sells for $2.7 million,” Associated Press, via NPR, &lt;a href=&quot;https://www.npr.org/2025/11/30/nx-s1-5626173/rubens-painting-sells-for-2-7-million-at-auction&quot;&gt;https://www.npr.org/2025/11/30/nx-s1-5626173/rubens-painting-sells-for-2-7-million-at-auction&lt;/a&gt;. Published November 30, 2025.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Andrew Neiderman, biography, &lt;a href=&quot;https://www.neiderman.com/about/&quot;&gt;https://www.neiderman.com/about/&lt;/a&gt;. V. C. Andrews died December 19, 1986.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;“How can I detect AI on Substack?”, Substack support documentation, &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/50891130623508-How-can-I-detect-AI-on-Substack&quot;&gt;https://support.substack.com/hc/en-us/articles/50891130623508-How-can-I-detect-AI-on-Substack&lt;/a&gt;. Retrieved July 27, 2026. See also Substack’s launch announcement, &lt;a href=&quot;https://post.substack.com/p/against-claudefishing&quot;&gt;https://post.substack.com/p/against-claudefishing&lt;/a&gt;, July 21, 2026.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Katharine English, “Pangram Flagged My Own Writing as AI,” July 23, 2026. &lt;a href=&quot;https://writerkatharine.substack.com/p/pangram-flagged-my-own-writing-as&quot;&gt;https://writerkatharine.substack.com/p/pangram-flagged-my-own-writing-as&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Sam Illingworth, “Substack’s AI Detector and the Return of the Witch Hunt,” &lt;em&gt;Slow AI&lt;/em&gt;,&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://theslowai.substack.com/p/substack-ai-detection-witch-hunt&quot;&gt;https://theslowai.substack.com/p/substack-ai-detection-witch-hunt&lt;/a&gt;, Published July 24, 2026. See also Illingworth, “AI Detection Will Always Be Broken,” &lt;em&gt;Slow AI&lt;/em&gt;, June 19, 2026 &lt;a href=&quot;https://theslowai.substack.com/p/ai-detection-does-not-work&quot;&gt;https://theslowai.substack.com/p/ai-detection-does-not-work&lt;/a&gt;; and his Substack note published July 22, 2026. &lt;a href=&quot;https://substack.com/@samillingworth/note/c-299343406&quot;&gt;https://substack.com/@samillingworth/note/c-299343406&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;N8, “In Defense of Pangram,” May 31, 2026. &lt;a href=&quot;https://n8programs.substack.com/p/in-defense-of-pangram&quot;&gt;https://n8programs.substack.com/p/in-defense-of-pangram&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Brian Jabarian and Alex Imas, “Artificial Writing and Automated Detection,” University of Chicago Becker Friedman Institute Working Paper 2025-116, August 26, 2025. &lt;a href=&quot;https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5407424&quot;&gt;https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5407424&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Weixin Liang et al., “GPT detectors are biased against non-native English writers,” &lt;em&gt;Patterns&lt;/em&gt; 4, no. 7 (2023). &lt;a href=&quot;https://arxiv.org/abs/2304.02819&quot;&gt;https://arxiv.org/abs/2304.02819&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed in this essay are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Virtual Intelligence and the Death of Authorship, Part 2</title>
    <link href="https://christopherhorrocks.com/essay/virtual-intelligence-and-the-death-5c5/" />
    <updated>2026-08-04T00:00:00Z</updated>
    <id>https://christopherhorrocks.com/essay/virtual-intelligence-and-the-death-5c5/</id>
    <content type="html">&lt;figure&gt;&lt;img src=&quot;https://christopherhorrocks.com/img/virtual-intelligence-and-the-death-5c5/cdfbedce-d23e-48e3-aebd-e0d0a66e58a5_1408x768.png&quot; alt=&quot;&quot; width=&quot;1408&quot; height=&quot;768&quot; loading=&quot;lazy&quot;&gt;&lt;/figure&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;This is the second part of a two-part essay about generative AI and the authorship of cultural products. [&lt;a href=&quot;https://chorrocks.substack.com/p/virtual-intelligence-and-the-death&quot;&gt;Part 1 is here.&lt;/a&gt;]&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;1. Thierry Rignol&lt;/h2&gt;
&lt;p&gt;Executive Master of Business Administration degrees, awarded through programs from business schools worldwide, afford working professionals an opportunity to advance their education. The coursework is often, but not always, paid for by the student’s employer. The professional’s future opportunities may open further with the awarding of the degree — especially if the awarding institution is prestigious.&lt;/p&gt;
&lt;p&gt;Thierry Rignol is a principal at Midwest Plains Capital. He was accepted to Yale’s Executive MBA program; tuition for this program is over $200,000. A professor suspected Rignol of having used AI to complete a final examination in Spring of 2024, using GPTZero to investigate. [1] The tool told the professor that whole sections of Rignol’s final exam were composed by AI.&lt;/p&gt;
&lt;p&gt;As Rignol tells it, his lengthy and well-polished answers were what caused a false detection. He also points out that tools like GPTZero have known biases against non-native English speakers like himself. A 2023 study of several detector products showed a high rate of false signals when set on text by such authors. [2]&lt;/p&gt;
&lt;p&gt;The examination was a self-timed, “closed internet, open book” exam with answers to be provided as a PDF, and AI tools were explicitly disallowed. Notably, of all of the examinations turned in by the class, only Rignol’s was flagged as a possible instance of AI use. Throughout the subsequent investigation, Rignol was asked to provide the original Word document that could help settle the issue of provenance. He was asked repeatedly for it, and only after many inquiries by Yale officials did Rignol provide an excuse: he had not used Word, but Apple Pages; therefore he had nothing that could be provided for all the previous requests.&lt;/p&gt;
&lt;p&gt;Why Rignol had not offered this Pages file before is left to the reader to guess. Despite being asked many times to provide possibly exculpatory evidence, Rignol was neither cooperative nor forthcoming with Yale during its investigation. Either Rignol has a poor understanding of file formats and assumed Yale didn’t know what Pages was, or some alternative explanation is required.&lt;/p&gt;
&lt;p&gt;The case is now a thirteen-count federal lawsuit, &lt;em&gt;Rignol v. Yale University&lt;/em&gt;, that’s nowhere near trial (while the case is mainly about the detection, Rignol also claims that he is being punished by Yale for his political views). [3] Mr. Rignol will have to explain, eventually, what led to his delay in providing the alleged original file to Yale. That’s why this case is important: the provenance of the ideas must be ascertained for Yale to be certain it is awarding degrees to those who do the work and deserve them. A detector cannot tell you that, but an original document could if it were substantially different from the final submission but had the same thesis or foundational ideas.&lt;/p&gt;
&lt;h2&gt;2. The Process&lt;/h2&gt;
&lt;p&gt;This is, in a sense, a similar solution to what Dr Sam Illingworth proposed for Substack’s Pangram implementation: eliminate the tool and instead tell us about your process; show it if you can. [4] This is a plea to honesty as the best policy. It is in large part founded on the knowledge that detection instruments like GPTZero and Pangram are not reliable. As examined in Part 1, it is possible for these tools to flag entirely human-authored text because they work through statistical pattern matching, and the patterns being matched often signal polished prose.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;3. &lt;em&gt;Heated Rivalry:&lt;/em&gt; The Fan Fiction AI Blowup&lt;/h2&gt;
&lt;p&gt;The television romance series &lt;em&gt;Heated Rivalry&lt;/em&gt; had a cultural moment in early 2026 when it premiered on HBO Max. [5] Like many cultural properties, it has attracted fan fiction writers who place the starring characters of the series in new situations. Many of these are erotica; others imagine the handsome hockey players in alternate universes and in different occupations.&lt;/p&gt;
&lt;p&gt;Fan fiction writing has a long and largely cryptic history, but erotica has often formed a core of it; it is only natural that &lt;em&gt;Heated Rivalry&lt;/em&gt; (already a successful book series before it came to television) would attract fan writers, much as the &lt;em&gt;Star Trek&lt;/em&gt; and &lt;em&gt;Harry Potter&lt;/em&gt; fandoms had before. Many of these works are uploaded for sharing with other fans at the website Archive of Our Own, also known as AO3. [6]&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;Heated Rivalry&lt;/em&gt; fan fiction community is in the grip of an AI authorship crisis. Stories with perceived AI tells generated controversy: writing patterns and impossibilities of human anatomy and relation in physical space, like being behind a person and still being able to admire their abdominal muscles.&lt;/p&gt;
&lt;p&gt;The turmoil over AI use increased sharply when a more reliable trace was discovered: Anthropic’s Claude labels its generated output with a CSS class called &lt;code&gt;font-claude-response-body&lt;/code&gt;. The class is invisible to readers, but it travels along with the text when it is pasted from Claude’s interface into editors that do not strip it, including AO3’s. It sits in the page’s source code, findable by anyone who looks. [7] An anonymous account on X published a 25-page document identifying thirty-eight popular &lt;em&gt;Heated Rivalry&lt;/em&gt; fan stories carrying the marker, and released custom code that made flagged text glow bright red. Some stories were entirely red. [8]&lt;/p&gt;
&lt;p&gt;This is a more reliable detection than Pangram’s. Claude left fingerprints that a human can trace with little difficulty. The author of a popular work who had previously decried the use of AI was found to have two such fingerprints in their own writing. The community treated those two instances as if they were hundreds. The writer has since ceased contributing new stories to the archive.&lt;/p&gt;
&lt;p&gt;The community’s response follows the same pattern seen among Substack writers in Part I. Like Substack, AO3 will not ban AI-created works outright. Betsy Rosenblatt, a law professor and a leader of the site’s legal committee, said: “We do not and cannot know how works are created before they’re posted,” adding that a rule against AI use “would be an engine for harassment.” [5] Community members on Reddit came up with two punitive options for when AI use is identified: ban flagged works, or ban the author of any work that is flagged.&lt;/p&gt;
&lt;h2&gt;4. Functional Replication&lt;/h2&gt;
&lt;p&gt;It is possible to trace the provenance of some generated artifacts, like text, from their tells. What is less common, but more sinister, is the use of AI tools to strip away provenance.&lt;/p&gt;
&lt;p&gt;The website malus.sh offers to take open-source software and produce “clean room” versions stripped of all licensing requirements. The site was created as a satire, but the service it offers is real: supply open-source software at one end, and a functionally identical version with no attribution or copyleft obligations is delivered at the other. [9] The site operates under the legal principles established in the 1982 precedent in which Columbia Data Products reverse-engineered the IBM PC’s BIOS. Courts held that no copyrights were infringed because copyright protects expression, not function; therefore, building something original from scratch that does the same thing is permitted. [10] AI collapses the clean-room timeline from months to minutes.&lt;/p&gt;
&lt;p&gt;The legal right to use AI clean rooms is not as clear-cut as it sounds. Dan Blanchard, the long-time maintainer of the Python library &lt;code&gt;chardet&lt;/code&gt;, used Anthropic’s Claude Code to create a new version of the library under a permissive license — first MIT, then 0BSD, a public-domain equivalent — replacing the original GNU LGPL, a copyleft license that requires modified versions to carry the same terms. [11] In a post from an individual claiming to be the library’s inventor, Mark Pilgrim, he noted that “Licensed code, when modified, must be released under the same LGPL license,” and that “[the] claim that it is a ‘complete rewrite’ is irrelevant, since they had ample exposure to the originally licensed code (i.e. this is not a ‘clean room’ implementation). Adding a fancy code generator into the mix does not somehow grant them any additional rights.” [12]&lt;/p&gt;
&lt;p&gt;When Anthropic accidentally published the source of Claude Code in March, Sigrid Jin did the same thing in reverse: a clean-room Python rewrite using a competitor’s AI tools. It was reportedly the fastest repository in GitHub history to reach fifty thousand stars. [13] Anthropic’s DMCA takedown notices removed the direct copies but not the rewrite — another instance of the clean-room principle in action.&lt;/p&gt;
&lt;p&gt;In all of these cases — malus.sh, &lt;code&gt;chardet&lt;/code&gt;, and Claude Code — functionality survived the AI rewrite process. What was lost was the authorship and all it entails: license, attribution, and the obligation of the open-source author to share and share alike.&lt;/p&gt;
&lt;h2&gt;5. Style as Property&lt;/h2&gt;
&lt;p&gt;Among the earliest controversies about generative AI has been the ability of systems to imitate, to one degree or another, the style of a well-known author. Grammarly offered this as a feature called “Expert Review,” generating editing suggestions inspired by named writers and academics — among them Stephen King, Neil deGrasse Tyson, and Carl Sagan — until user backlash and a class-action lawsuit forced the feature’s withdrawal in March. [14] The lawsuit, &lt;em&gt;Angwin v. Superhuman&lt;/em&gt;, was filed in the Southern District of New York on behalf of writers and academics whose identities were used without consent. [15]&lt;/p&gt;
&lt;p&gt;OpenAI also permitted ChatGPT to be used for style mimicry, though without advertising it as a feature. That changed in late July, when ChatGPT began to refuse requests to write in a named author’s style. [16] Upset users went to Reddit to complain about this change. The original poster of one complaint thread noted that the entire novel they had been writing was produced by prompting ChatGPT to write in the style of a particular author. [17] This is an ugly specimen of confused authorship: writing by specification (a prompt) to generate an artificial text, written by a machine in the style of another person who is not a party to the exchange.&lt;/p&gt;
&lt;p&gt;The top-voted workaround in this thread showed some ingenuity: have the model describe the style that is desired, build a persona around that description, and then feed it into a new ChatGPT session. This is a clean-room bypass in the spirit of what malus.sh does — the same evasion of the constraint, in a different domain. The service offered by AIStoryHub appears to operate on the same or a similar mechanism and is, in part, pitched to fan fiction writers. [18]&lt;/p&gt;
&lt;p&gt;One Redditor’s reply to the thread was heavily downvoted but survived moderation: “Try writing your own words.”&lt;/p&gt;
&lt;p&gt;In June, the CREATOR Act was introduced into the U.S. House of Representatives (H.R. 9112) to protect visual artists from having their distinctive style imitated in generated imagery. [19] It is an attempt to make style a protectable asset.&lt;/p&gt;
&lt;p&gt;Three responses to the problems of authorship arrived inside of half a year: one commercial (Grammarly pulling the feature), one technical (OpenAI refusing the prompt), one legislative. None of them addresses the cases where nobody’s style is being imitated because nobody authored the work. The current toolkit protects named authors and mandates disclosure. It has no mechanism for the growing category of production  — true AI slop, made to capture eyeballs on revenue-generating platforms  — where no human author exists to name, protect, or hold responsible for the content.&lt;/p&gt;
&lt;h2&gt;6. The Failures of Detection&lt;/h2&gt;
&lt;p&gt;What links all of the cases we have examined is that detectors resolved none of them. GPTZero flagged Rignol’s exam at Yale; like Pangram at Substack, it produced more noise than real results. The fan fiction community around &lt;em&gt;Heated Rivalry&lt;/em&gt; began its hunt with detectors but found a more reliable trace in a CSS artifact left behind by Claude. People who use &lt;code&gt;chardet&lt;/code&gt; found the new version functionally identical even though it was supposed to be a ground-up clean-room rewrite; a real possibility with software, but the case is not clean because the person generating the rewrite is, as the maintainer, intimately familiar with the original being reproduced.&lt;/p&gt;
&lt;p&gt;Every current mechanism — detectors, disclosure labels, style-protection statutes — assumes there is an author to find, to protect, or to hold responsible. The growing category of cultural products generated by AI has no such person to credit or discredit.&lt;/p&gt;
&lt;p&gt;There is a small but visible trend of human authors purposefully leaving errors in their text to show they are human-written; but that can be done by software, too. Sinceerly, a Chrome extension that injects typos into polished text to defeat detectors, is the logical endpoint of the arms race between AI-assisted writers and detectors: imperfection as the last remaining signal of human authorship, manufactured by a machine. [20]&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;Heated Rivalry&lt;/em&gt; fan fiction community’s stated goal is to encourage and promote human-written texts. The idea is that all-human writing is a bulwark against being inundated by AI slop. Precedent is again instructive here, because the publishing world has had a different slop era to draw lessons from. The pulp fiction age of the 1920s through 1940s provided inexpensive and often poor reading material to millions. Cheaply printed ephemera, the pulps were mostly trash written by writers who were churning out reams of material just to keep body and soul together. The pulps were also the training grounds of giants. The pulps gave us Raymond Chandler, Isaac Asimov, and Ray Bradbury, among others. In the urge to eliminate all AI assistance in writing, it might seem tempting to tear down new talents because they work differently from past authors. The history of the pulps tells us that would be an error.&lt;/p&gt;
&lt;h2&gt;7. The Open Grave&lt;/h2&gt;
&lt;p&gt;Thierry Rignol’s case is still unresolved, but he has encountered some skepticism of his claims about why he was not forthcoming in providing documentation that the work he turned in was his own, and not generated by AI. “Wouldn’t a reasonable professional person who was trying to be cooperative with a proceeding,” one judge asked, “upon getting not just one email but many emails asking for the underlying document that was used to create a PDF say, ‘Oh, I didn’t use Word. I used Pages, a different word processing [program]?’” [3] Rignol’s apparently evasive actions do not look like those of a person who is confident in his authorship or that it will be vindicated upon examination. Yale has moved to dismiss the case.&lt;/p&gt;
&lt;p&gt;What the users of detector tools and other methods of finding AI use in cultural products cannot tell us about is the authorship of what is being scanned. It cannot tell us if the words come from the human presenting them, from the AI that generated them, or from a third party whose work was imitated by the machine through prompting. Authorship, as illustrated in Part I by the example of Rubens and the artist’s workshop, is not about each word on a page or each individual line made by a hair in a brush. It is a real person staking their name on their method of working, and standing behind it when asked how the work was made.&lt;/p&gt;
&lt;p&gt;Detectors, text artifacts, and statutes can protect named authors and mandate disclosure. There is no mechanism, however, that governs or regulates the growing category of works that have no named human author to protect or hold responsible, or to tell where the human contribution ends and the machine’s outputs begin. It is here where authorship’s grave has been dug.&lt;/p&gt;
&lt;hr&gt;
&lt;hr&gt;
&lt;h2&gt;Footnotes&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Nate Anderson, “How a Yale AI-cheating dispute became a 13-count federal lawsuit,” &lt;em&gt;Ars Technica&lt;/em&gt;, July 31, 2026. &lt;a href=&quot;https://arstechnica.com/tech-policy/2026/07/how-a-yale-ai-cheating-dispute-became-a-13-count-federal-lawsuit/&quot;&gt;https://arstechnica.com/tech-policy/2026/07/how-a-yale-ai-cheating-dispute-became-a-13-count-federal-lawsuit/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Weixin Liang et al., “GPT detectors are biased against non-native English writers,” &lt;em&gt;Patterns&lt;/em&gt; 4, no. 7 (2023). &lt;a href=&quot;https://arxiv.org/abs/2304.02819&quot;&gt;https://arxiv.org/abs/2304.02819&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Rignol v. Yale University&lt;/em&gt;, No. 3:25-cv-00236 (D. Conn.). Docket entries and judicial quotation as reported in Anderson (note 1).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Sam Illingworth, “Substack’s AI Detector and the Return of the Witch Hunt,” &lt;em&gt;The Slow AI&lt;/em&gt;, July 24, 2026. See also Part I of this essay.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://theslowai.substack.com/p/substack-ai-detection-witch-hunt&quot;&gt;https://theslowai.substack.com/p/substack-ai-detection-witch-hunt&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Emmy Martin, “When A.I. Invaded ‘Heated Rivalry’ Fan Fiction, the Meltdown Was Epic,” &lt;em&gt;The New York Times&lt;/em&gt;, July 30, 2026. &lt;a href=&quot;https://www.nytimes.com/2026/07/30/technology/ai-heated-rivalry-fan-fiction.html&quot;&gt;https://www.nytimes.com/2026/07/30/technology/ai-heated-rivalry-fan-fiction.html&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Archive of Our Own, &lt;a href=&quot;https://archiveofourown.org/&quot;&gt;https://archiveofourown.org/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Whitney A. Foster, “A 25-Page PDF Named 30 Fanfiction Writers as AI Users. Here’s What the Evidence Shows,” Substack, July 2, 2026.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://whitneyafoster.substack.com/p/ao3-ai-fanfiction-heated-rivalry-detection&quot;&gt;https://whitneyafoster.substack.com/p/ao3-ai-fanfiction-heated-rivalry-detection&lt;/a&gt;. See also Candyce Edelen’s original 2025 discovery of the CSS class in pasted Claude output. &lt;a href=&quot;https://www.linkedin.com/posts/candyceedelen_well-crap-i-didnt-know-claude-was-adding-activity-7448391850501124096-podz/&quot;&gt;https://www.linkedin.com/posts/candyceedelen_well-crap-i-didnt-know-claude-was-adding-activity-7448391850501124096-podz/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;@heatedrivalryai (John Doe), “Fandom Has a Hidden Generative A.I. Problem,” published via X, June 29–30, 2026.&lt;/p&gt;
&lt;blockquote class=&quot;tweet&quot;&gt;&lt;p&gt;The following link leads to our complete findings, the site skin, an appendix of works that contain the Claude code fragment, and copies of the referenced fics: &amp;lt;a class=&amp;quot;tweet-url&amp;quot; href=&amp;quot;https://drive.google.com/file/d/1YyFKLzunnvVJRDTql5nHJwBb-swxIftn/view?usp=sharing&amp;quot;&amp;gt;drive.google.com/file/d/1YyFKLz…&amp;lt;/a&amp;gt;.&lt;/p&gt;&lt;footer&gt;&lt;a href=&quot;https://x.com/heatedrivalryai/status/2071719230223811032&quot;&gt;John Doe (@heatedrivalryai), June 29, 2026&lt;/a&gt;&lt;/footer&gt;&lt;/blockquote&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Emanuel Maiberg, “This AI Tool Rips Off Open Source Software Without Violating Copyright,” &lt;em&gt;404 Media&lt;/em&gt;, April 21, 2026. &lt;a href=&quot;https://www.404media.co/this-ai-tool-rips-off-open-source-software-without-violating-copyright/&quot;&gt;https://www.404media.co/this-ai-tool-rips-off-open-source-software-without-violating-copyright/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Columbia Data Products reverse-engineered the IBM PC BIOS in 1982 using isolated teams; courts held that copyright protects expression, not ideas or function. &lt;a href=&quot;https://en.wikipedia.org/wiki/Clean-room_design&quot;&gt;https://en.wikipedia.org/wiki/Clean-room_design&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Dan Blanchard, “Everything Claude Saw,” personal blog. &lt;a href=&quot;https://dan-blanchard.github.io/blog/chardet-rewrite-controversy/&quot;&gt;https://dan-blanchard.github.io/blog/chardet-rewrite-controversy/&lt;/a&gt;. See also Simon Willison’s coverage and ShiftMag, “License Laundering and the Death of Clean Room.”&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Mark Pilgrim, comment on the &lt;code&gt;chardet&lt;/code&gt; relicensing. &lt;a href=&quot;https://github.com/chardet/chardet/issues/327&quot;&gt;https://github.com/chardet/chardet/issues/327&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Sigrid Jin (GitHub: instructkr), &lt;code&gt;claw-code&lt;/code&gt; repository. Star count reported variously at 50,000 in approximately two hours, rising past 100,000 within a day. See Layer5, “The Claude Code Source Leak,” and Business Insider coverage.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Emma Loffhagen, “Grammarly removes AI Expert Review feature mimicking writers after backlash,” &lt;em&gt;The Guardian&lt;/em&gt;, March 13, 2026. &lt;a href=&quot;https://www.theguardian.com/technology/2026/mar/13/grammarly-removes-ai-expert-review-feature-mimicking-writers-after-backlash&quot;&gt;https://www.theguardian.com/technology/2026/mar/13/grammarly-removes-ai-expert-review-feature-mimicking-writers-after-backlash&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Angwin v. Superhuman&lt;/em&gt;, filed in the U.S. District Court for the Southern District of New York, March 2026. Julia Angwin, investigative journalist, is the lead plaintiff.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Kyle Orland, “ChatGPT starts blocking direct requests to copy an author’s style,” &lt;em&gt;Ars Technica&lt;/em&gt;, July 27, 2026. &lt;a href=&quot;https://arstechnica.com/ai/2026/07/chatgpt-stops-cloning-famous-writers-voices-but-may-capture-a-similar-feeling/&quot;&gt;https://arstechnica.com/ai/2026/07/chatgpt-stops-cloning-famous-writers-voices-but-may-capture-a-similar-feeling/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;u/Dazzling-Major-5620, post on r/WritingWithAI, July 24, 2026. &lt;a href=&quot;https://www.reddit.com/r/WritingWithAI/comments/1v5dfd0/chatgpt_not_letting_you_emulate_specific_authors/&quot;&gt;https://www.reddit.com/r/WritingWithAI/comments/1v5dfd0/chatgpt_not_letting_you_emulate_specific_authors/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIStoryHub, &lt;a href=&quot;https://aistoryhub.co/&quot;&gt;https://aistoryhub.co/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;CREATOR Act (Creative Rights Ensuring Artists’ Technique and Originality are Reserved Act), H.R. 9112, 119th Congress, introduced June 2, 2026 by Rep. Beth Van Duyne (R-TX). Referred to the House Committee on the Judiciary; no further action as of July 2026.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Sinceerly, Chrome browser extension by Ben Horwitz. See Fast Company coverage, April 25, 2026. &lt;a href=&quot;https://www.fastcompany.com/91531539/this-anti-grammarly-ai-tool-adds-typos-to-your-emails-on-purpose&quot;&gt;https://www.fastcompany.com/91531539/this-anti-grammarly-ai-tool-adds-typos-to-your-emails-on-purpose&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The opinions expressed in this essay are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
</feed>