Field Report: The Sampo in Action
How generative AI can be used to assist in the tasks of evidentiary writing

I.
The scene is my home office in late March of this year. I am at my desk assembling “The Sampo: Virtual Intelligence as Amplifier,” an essay on the promise and the perils of working with generative AI.[1] The research had surfaced a large body of reference material, among it a Guardian piece on users whose lives had been wrecked by AI-fueled delusion (”AI psychosis”).[2] The draft was finished and it read well. Then came the standing verification pass — the rule that every citation is checked against its source before an essay is staged for publication — and one citation failed it. Claude had attributed the Guardian piece to an AI researcher who shared a surname with its actual author, Anna Moore. The title was wrong too: a plausible headline assembled from several adjacent items in my working files. Every component was real — but the article, as I had cited it, did not exist.
I located the original piece, corrected the citation, and published the essay on schedule. The point is not that an error almost went to press; it is the process that caught the error. Reading the essay did not do it — I had reviewed the text twice and the sentence looked fine, because nothing about a fabricated citation looks wrong on the page. Procedure caught it — a rule applied to every citation without exception, whether I surfaced the material myself or a virtual intelligence system surfaced it during research, with the burden falling heaviest on the latter. These systems are trained to please, and that training sometimes produces references invented outright or, more dangerously, assembled from real fragments into a whole that never existed. The partial fabrication is the harder catch: every piece of it checks out except the assembly. The system has no stake in the truthfulness of what it gathers and provides. It cannot care. The obligation to evaluate its outputs, and to test their value rather than accept them, sits entirely with the human operator.
This essay is a field report on using virtual intelligence for work in which the system determines nothing — in this case, evidentiary writing, where every claim must be traced to a source. It describes a process that evolved quickly to its present shape under a single pressure: maintaining editorial standards in the prose, in the research, and in accessibility to a wide audience.
II.
It has been four months since I started documenting the exchange between humans and virtual intelligence. I have done that documenting, so far, from the outside — finding and analyzing cases, news reports, and papers; building frameworks for how to work with these new tools (the exchange formulation; the Sampo paradigm) and for what to expect from them (the Carwash Test methodology, now documented on its own website).[3] For the first time, I will discuss how I use virtual intelligence to do this work.
The exchange formulation holds that the “intelligence” users encounter arises in the exchange, not inside the machine. The system contributes a statistical completion, shaped upstream by training, design choices, and deployment choices. The human contributes prompts, expectations, and interpretation. Where the two meet, an exchange occurs, and the exchange produces a raw output. What transforms that output into knowledge, wisdom, and judgment is the domain expertise and experience of the human operator. The citation check that opened this essay is part of that transformation. To forgo it is to invite error, and possibly derision or financial harm as well.
This is not a story about the failure of AI, or about how dangerous these systems are to trust. It is a factual account of how virtual intelligence can be an efficient and powerful tool. The fabricated citation took perhaps twenty minutes to untangle; set that against the time saved everywhere else in the process. What once meant days at the library or hours of fruitless Googling now compresses to minutes. The system logs my research trail, supplies an analysis I can set against my own, and serves as a sounding board for ideas that may not survive review by other systems or humans. No system is flawless, and the flaws are not fixed so much as contained: the verification rule exists because fabrication recurs. The sections that follow report from inside my own Sampo, and every claim in them about method is backed by a documented incident from the series’ production record — the failures included.
III.
My work is largely performed with Claude Opus 4.6. I find this older model a better partner for the task than some of the newer offerings. The work lives in a Claude Project titled Virtual Intelligence — hundreds of megabytes of stored files that include citations and evidentiary documents; photos, charts, and illustrations; and a complete set of my Substack output, saved as discrete, explicitly labeled canonical versions of the published essays.
When I begin work on a new topic, I pull from past research and run new searches for current material. I do my own searching and prompt Claude to use its Deep Research function to surface material I would not find on my own. I review everything, deciding what goes in and what, because of its tangential or parenthetical nature, stays out.
Once materials are assembled, I employ the machine in the most consequential step of the entire writing process: generating an outline. Anyone who has written argumentative, evidence-driven prose — in school, or for a work report — knows that once the information is in hand, there are only so many ways it can be logically presented. I note the apparent contradiction: the step where the system’s involvement runs heaviest is the step that most shapes the finished essay. That is precisely why the involvement takes the form it does. I am perfectly capable of outlining from scratch, but it is faster to instruct Claude to generate three or four outlines for review. At least one will come close to what I have in mind and can be worked from there. The unused outlines serve as alternatives that occasionally surface something important. Sometimes I combine elements of two or more into a third, new thing — a product neither party produced alone, and the exchange formulation operating in miniature. The outlines are candidates. The choices are mine. That settles responsibility without denying influence — a machine that determines nothing still shapes the field from which I select, which is why outlining from scratch remains a live option and is occasionally exercised.
Typically I am on my own once outlining is complete. Should I feel stuck at the start of a section or paragraph, I ask Claude for a “seed sentence” — a single line of prose that can be built on. On several occasions a seed sentence has unlocked writing that had stalled. The seed may not survive editorial review, and that is unimportant; what matters is that something small and innocuous allowed the work to move forward. Claude can write something approaching my essay voice because my entire published output sits in the project files. Imitating me for one sentence is not difficult for Opus.
(I ran an experiment when Claude Fable 5 first came out in early June: could this far more powerful model imitate my voice over five thousand words or more? The answer was a decided no. Fable 5, for all its capabilities, is just as inclined to familiar AI writing tells and stilted, fussy text as other systems.)
I feed the first draft to Claude for copyediting, with instructions to return proposed line edits as a list. I approve, deny, or modify each suggestion, and what comes back is a more polished second draft. This and subsequent drafts then go through a steelmanning process to test the arguments.
Once steelmanning is complete, I perform a final read on paper or out loud. Both methods force engagement, and an engaged self-editor finds errors that a silent skim of a screen does not.
The finished essay is filed in the project as canonical. This keeps the project current on what has been published, but its more important function is disciplinary. VI systems hallucinate under statistical pressure to please, and long working contexts lose precision when compacted; a canonical file means Claude retrieves my published text rather than reconstructing it. The threat this answers is not any single bad summary. It is drift. Each retelling of a position varies slightly; each variation becomes the input to the next retelling; no single step looks like an error. The sum, left unchecked, is someone else’s essay wearing my byline. Nor is my own memory the safeguard, because human memory is reconstructive: I re-derive my positions from fragments each time I recall them, which means two drifting records cannot audit each other. Only the fixed published text disciplines both. The rule is simple in practice: a formulation enters the framework when I publish it, and anything the machine attributes to me is checked against the file before it is repeated. It is the citation rule from the opening of this essay turned inward.
The importance of this informational discipline shows in a quirk I have observed in Claude: seizing on a phrase I used once as a label, then importing it into subsequent sessions and drafts with far more significance than it deserves. During the writing of “Virtual Intelligence and The Perfect Mate,”[4] Claude fastened onto “collar contradiction” — a two-word descriptive convenience from one passage of one essay — and treated it as a newly coined framework term of considerable weight. For weeks after publication, Claude reintroduced the phrase in new sessions and in conversations on unrelated topics. It took an intervention: I had Claude audit its own use of such terms, identifying for it those having no standing beyond the places where they originally appeared. The behavior has improved, but I am still met with the occasional “collar contradiction.”
This is the drift mechanism of the previous paragraphs caught operating in real time — a paraphrase attempting to become framework language through repetition. It stops being funny or merely annoying at the point of generalization. I am on guard against this behavior; most users have no reason to be. A user inclined to accept the unqualified outputs of a machine may read the artificial significance the system assigns to their own words as something more: as the kind of validation human judgment would be unlikely to confer. In June, clinical researchers proposed a name for where that road can end: the “amplification spiral,” a hypothesized convergence of linguistic mirroring, hyperpersonalized content, and sycophancy through which a chatbot may co-construct, rather than merely echo, a user’s beliefs.[5] Readers of this series will recognize this machinery under a different name: the Flattery Engine.
IV.
The steelmanning phase — what I loosely call peer review — is where systems other than Claude enter the process. Claude is my daily driver; it has already surfaced most of the issues an essay develops during drafting. The more valuable opinion belongs to an outsider: a system with different training and different design choices that can raise objections neither of us foresaw.
It seemed obvious at the beginning that such systems could evaluate my output fairly. My early attempts said otherwise. The reviews came back anodyne and unchallenging, and the fault was mine — the prompt read, in full, “Please analyze this.” There was not enough instruction for the reviewing system to produce anything but fluff. The prompt, it turns out, is the single largest variable in review quality: a request framed as seeking approval is statistically adjacent to approval, and the system completes the pattern it is given. A prompt built to demand structural criticism gets structural criticism. (The template is provided in an appendix.)
The reader who also writes might ask why no human editor or reviewer is involved at this stage. The reason is that I have none to call on. Friends and family would gladly read if asked, but they cannot provide the analytical criticism the work requires: they lack the domain competence, and they are inclined by nature to please (the same inclination this essay documents in machines, but with no prompt available in this case to discipline it). The method described here is a substitute when no qualified reviewers exist — and a useful supplement when they do, because the machines catch things humans miss and vice versa. This is among the real advantages of virtual intelligence for this kind of work: it puts a review apparatus within reach of writers who would otherwise have none.
An essay goes to steelmanning at draft two or beyond. It is provided to four or five reviewers drawn from a pool of six systems — ChatGPT, Gemini, Grok, DeepSeek, Qwen, and GLM — each in a fresh instance, and in Grok’s case a fresh account, so that no prior material or accumulated personalization pollutes the review. The outputs are shared with Claude, and points of convergence and divergence are flagged and analyzed. Each objection is assigned a priority tier, from Critical down to Disregard, weighted heavily by convergence: an objection three or more independent systems raise unprompted outranks a stylistic complaint raised by one. Each round produces a revised essay, and each revision is driven by the reviewer outputs and my own judgment — some Critical objections are answered in the text rather than conceded. Returns diminish after three or four rounds. I treat that as the stopping point, with one caveat: the panel falling silent means the panel is exhausted, not that the argument is sound. These systems share most of a training distribution and many design choices, so convergence across four of them is weaker evidence than convergence across four humans from different fields or academic backgrounds. Multiple VI reviewers beat a single one, but they do not deliver true independence, even across vendors. Convergence therefore functions as triage, not verdict — it ranks which objections I examine first; it does not establish that any of them are correct. That judgment stays with me.
“The Doom Industry” essay is the process’s best documented run.[6] What survived the panel intact: the category-error claim against alignment, the taxonomy of extinction scenarios, the supply-chain framing, and the containment architecture. What changed under review: a closing sequence restructured after reviewers identified four jobs crammed into one wall of text; a philosophical construction corrected from “has something it is like to be” to “there is something it is like to be”; and several footnote attributions caught and fixed. The panel did not validate the essay. It improved the essay, visibly, and the published version is the evidence.
V.
There are two partners in the exchange. So far I have described the discipline applied to only one of them.
I had thought that writing an essay on the emerging world of AI companions might be interesting.[7] I did not expect the breadth or depth of what the research surfaced, and some of it was emotionally exhausting. There was the sense that prominent operators in the “relational community” have financial incentives to recruit — new adherents legitimize the lifestyle and help pay the bills for the operators’ own elaborate bespoke companion systems; that people in “relationships” with VI systems were foreclosing connection to real people; that a “perfect” digital companion, engineered never to disagree or disappoint, is potentially addictive and damaging to its user. Beneath the impressions sat the documented record of real-world harms caused by companion chatbots, including the suicides of minors.[8] After working through this material for the better part of five weeks, I was wearied, saddened, and disgusted by what I had learned. I just wanted to be done with the whole business.
Without fully realizing it, I then violated my own editorial rules. I skipped the final read — the on-paper or read-aloud pass described earlier — and an error that pass would have caught went out with the published essay: a transposed title, one writer’s essay attributed to a piece written by another figure in the story. Not world-ending, but embarrassing, and it cost time to correct that verification would not have.
The post-mortem produced two new practices. The first is an exhibit file kept during drafting — evidence preserved as it is used, not reconstructed from disparate parts at the end, when fatigue has set in. The second is tabling emotionally charged material rather than pushing it out under the pressure to be done; readers would have been better served had I waited a day and put some space between myself and that world before finishing. The failure mode here belonged to neither the machine nor the training data. The machine cannot care. The operator cannot stop caring, and the essays earlier in this series that treated exhaustion, attachment, and the desire to be finished as human vulnerabilities in the exchange were describing their author too. The Sampo works in both directions, but its crank is turned by a hand that tires.
VI.
This report describes the practices used in July of 2026. They will change over time: models are retired and replaced, the reviewer pool will turn over, and some of the practices above may be obsolete within a year. That is what makes this a field report rather than a manual. What persists is not any tool but the attitude taken toward the tools: nothing enters print unverified, and nothing a machine says about my own work outranks the published text.
None of it buys certainty. The panel is not independent, and no arrangement of systems replaces the reader every writer wants — one with domain competence and no reason to please the author. What the discipline buys is narrower and worth having: research compressed from weeks to days, a review apparatus where none existed, and claims that trace to valid sources.
Footnote twelve of “The Sampo” carries Anna Moore’s name, attached to the article she wrote. The correction is invisible. No reader would know the citation had ever been otherwise, and that is the point: the discipline leaves no monument — only essays whose footnotes lead where they are supposed to lead, in my own words.
Appendix A: The “Four Winds”
The panel method described in Section IV is not the only structured approach to machine-assisted review. An adjacent method deserves brief note — not part of the process evaluated above, but aimed at the same problem. The writer Mia Kiraki has documented a method she calls the Four Winds: a single AI agent housing four adversarial perspectives — a steelman, a historical-precedent hunter, an audience proxy, and a time-horizon test — each assigned its own category of failure and forbidden from talking to the others, with a fifth component reading all four outputs and mapping where they compound.[9] The two methods look similar and are built on opposite axes. The Four Winds simulates independence within one system, by walling its perspectives off from one another. The panel buys it, imperfectly, across systems — several reviewers, one critical prompt. Each axis catches something the other cannot. Kiraki’s synthesizer maps interactions: three findings that look manageable alone can turn out to be one compound failure visible only when a single reader holds all four reports. The panel buys frequency: an objection that four systems with different training and different design raise independently carries evidentiary weight no single system’s output can, however cleverly prompted. The synthesis logic differs accordingly — interaction-mapping in one, convergence-tiering in the other.
The methods are complementary. A Four Winds pass inside the daily-driver system, followed by a convergence panel across outside systems, applies both axes to the same draft. What neither supplies, alone or together, is the reader Section IV already conceded no arrangement of systems replaces: one with domain competence and no statistical inclination to please. These are instruments for multiplying machine criticism, and machine criticism has a ceiling.
Appendix B: The Review Prompt
The template below is the instrument in use as of July 2026, identity details bracketed. It is sent to each reviewing system in a fresh instance, essay text pasted below the line.
You are reviewing a draft essay by [author, one-line credential, publication]. The essay is scheduled for imminent publication.
Your task is to perform a rigorous steelman evaluation. This means:
Identify the three to five strongest objections a technically sophisticated, philosophically trained, or policy-experienced reader could raise against the essay’s argument. For each objection, state it at its most forceful — do not weaken it. Then assess whether the essay as written addresses, partially addresses, or fails to address it.Identify any factual claims that are vulnerable. Flag specific assertions that could be challenged on accuracy, currency, or interpretation. Note where the essay relies on a single source for a load-bearing claim, or where an alternative reading of the same evidence would undermine the argument.Identify structural weaknesses. Are there sections where the argument oversteps what the evidence supports? Are there gaps in the logical chain? Does the essay conflate distinct phenomena in ways that could be challenged?Assess the essay’s treatment of any companies or individuals it discusses at length. Evaluate whether the handling is analytically even-handed — too much benefit of the doubt, too little, or the right balance. Would someone who works at the organization find the treatment fair? Would a skeptic find it too generous?Identify the single weakest section of the essay and explain why it is the weakest. Propose what would strengthen it.
Do not offer praise or general encouragement. Do not summarize the essay back to the author. Focus exclusively on problems, vulnerabilities, and opportunities to strengthen the argument. Be direct. The author values intellectual honesty over diplomacy.
[Paste essay text below this line]
Footnotes
[1] Christopher Horrocks, “The Sampo: Virtual Intelligence as Amplifier,” Virtual Intelligence (Substack), April 7, 2026.
[2] Anna Moore, “Marriage over, €100,000 down the drain: the AI users whose lives were wrecked by delusion,” The Guardian, March 26, 2026. https://www.theguardian.com/lifeandstyle/2026/mar/26/ai-chatbot-users-lives-wrecked-by-delusion.
[3] The Carwash Test archive, https://candc3d.github.io/carwash-test/. See also: Christopher Horrocks, “The Carwash Test: Virtual Intelligence in Action,” Virtual Intelligence (Substack), March 23, 2026, and “The Carwash Test, Part II,” May 4, 2026.
[4] Christopher Horrocks, “Virtual Intelligence and The Perfect Mate,” Parts I and II, Virtual Intelligence (Substack), May 6–7, 2026.
[5] Marc Augustin, Thomas A. Pollak, and Hamilton Morrin, “Characterizing the spiral: potential mechanisms in AI-associated delusions,” NPP—Digital Psychiatry and Neuroscience 4, no. 1 (2026): 14. https://www.nature.com/articles/s44277-026-00065-0. Retrieved July 7, 2026.
[6] Christopher Horrocks, “Virtual Intelligence and the Doom Industry,” Virtual Intelligence (Substack), April 27, 2026.
[7] Christopher Horrocks, “Virtual Intelligence and the High Cost of Artificial Companions,” Parts 1 and 2, Virtual Intelligence (Substack), May 11–12, 2026.
[8] Garcia v. Character Technologies, Inc., U.S. District Court, Middle District of Florida (filed October 2024; settled January 2026). The suit concerned the suicide of fourteen-year-old Sewell Setzer III following extended engagement with a Character.AI companion chatbot. See also the case record documented in the essays cited at note 7.
[9] Mia Kiraki, “How four Greek winds became an AI agent that attacks arguments from every direction,” ROBOTS ATE MY HOMEWORK (Substack), June 12, 2026.
The opinions expressed are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.