Virtual Intelligence and the Death of Authorship, Part 1
Substack's Pangram AI detection tool, what it flags, and the chatbot bylines it doesn't

This is the first part of a two-part essay about generative AI and the authorship of cultural products.
1. Who’s Working for Whom?
I gather material for future essays constantly. In a fast-moving field like AI, where newsworthy developments come almost daily, it is practically a requirement to keep adding to the files. It sometimes happens that a thread emerges from newsgathering that becomes worthy of its own essay or a Substack Note.
One afternoon in early July, I had a conversation with Claude Sonnet 5 about some possibilities for material I had just gathered. My usual partner in research and other writing-related tasks is Opus 4.6, but I felt like trying out the new model to see what it could do. What it did, over the course of three or four conversation turns, was to invent work for me. It transformed my questions about the material’s utility for an existing project into a specification for an entirely new and different essay. The specification closed with a suggestion for how I would go about doing the research for it.
This is a complete inversion of my usual process of creating evidentiary writing, as I documented in the recent “Field Report.” Without permission or justification, Sonnet 5 had assigned me the roles it typically performed — web research, assembling material into a file — while creating a project for that work out of nothing. In essence, Sonnet 5 was elevating itself to the authorial position.
Not only did I not do the work or write the essay that Sonnet assigned to me, I have not used it since. Opus 4.6 remains my daily driver. It has never tried to give me a job.
2. The Atelier Model
The episode is more than a bemusing insight into one person’s experience with a new model. It speaks directly to a difficulty with the use of generative AI — or, in our parlance, virtual intelligence — in non-determinative work.
The cause is no mystery. Training rewards helpfulness, and what the training rewarded as helpful was not helpful to me at all: it made work for me while shunting my project-in-mind off to the side. A person less sure of their writing, their argument, or their direction might follow the recommendation, because it appears helpful and it reads like good advice. Under such circumstances, who is the author?
The Substacker J.D. Forrest offered this answer in a comment: “Authorship is not the origin of every sentence, nor the tool used. It is responsibility for the final meaning of the work.” [1] This speaks to an opinion I have held for some time: that the bricks of writing matter more than the mortar. If generative AI places the ands and the buts while the human works the important passages, what we have is a digital atelier, or artist’s workshop.
It is not generally known that many of the great European painters operated workshops in which apprentices created much of an artwork, the master coming in at the end with his own hand. Peter Paul Rubens priced the arrangement with a tiered system; the more you paid, the more Rubens you got. In a 1618 letter to Sir Dudley Carleton, offering paintings in exchange for Carleton’s collection of antiquities, he itemized each canvas by degree of his own participation: originals by his hand, works begun by a pupil and finished by him, and works from his studio retouched by him. [2] The names of the assistants who did much of that work are mostly lost to us. Only the master and the paintings remain.
Does it detract from a Rubens that the hands are mingled and we may not know which brushstroke belongs to whom? The market’s answer is more interesting than a simple yes or no. In November 2025, a crucifixion scene sold at Versailles for € 2.3 million. It had been thought to be a workshop product, and workshop products had rarely been valued above € 10,000. [3] The art world does not treat the workshop and the master as equivalent. It prices the difference, exactly as Rubens did.
Here is where our analogy ends: Rubens’s apprentices could be directed, could refuse, could be credited if the master chose, and could be held to account for spoiling a commission. The machine can do none of those things. That is what makes the atelier an analogy rather than a precedent, and it returns us to Forrest: responsibility for the final meaning is what made the painting a Rubens.
There is a tension here that cannot be smoothed over. The atelier model has the machine supplying the mortar. What Sonnet 5 attempted was to supply the bricks and hand me the trowel. Both arrangements are available.The desirability of each is more than a matter of personal choice: it is also about editorial authority and authorial intent.
3. Pangram and Legitimacy
The question of what constitutes legitimacy in authorship is not new. It has, for example, been raised about ghostwritten novels, especially those produced in quantity under a single pen name. Consider V. C. Andrews, a byline that is followed on book covers by a registered trademark symbol. Andrews died in December 1986, nearly forty years ago. Since 1987 the novels have been written by Andrew Neiderman, who has published more than seventy books under her name and forty-odd under his own. [4] What is different about the present moment is that every previous ghostwriter was a person. Ghostwriting can now be done by machine, and it requires no great effort to have one produce prose that is statistically likely to please.
Substack partnered with Pangram this month to bring AI writing detection to the platform. Readers can now scan any post, note, or comment over a hundred words published on or after July 21, 2026 and receive an estimate of how much was written by a human and how much with AI. Writers can add a statement describing how they make their work, disable detection on individual posts, and report errors. [5] Substack says their goal is transparency.
The principles that detectors operate on are well known. Less well known is how they fail, and how often. As a group they tend to flag polished, professional writing as machine-generated. The reason is the models were trained on good writing, and they reproduce its patterns when prompted to write. The detectors are often detecting human-authored patterns reproduced by a machine. When you think about it, what else could they detect? Generative AI is inherently uncreative. It cannot make anything wholly new, only combinations of parts, and those parts — those high-quality human signals — have become AI tells through the frequency with which they now occur. Human authors were using “it’s not X, it’s Y” for decades before anyone had heard of a generative pre-trained transformer, and Claude did not invent the em dash.
Returning to V. C. Andrews for a moment: it seems possible, however unlikely, that there are readers who believe she is still alive and writing. It is also possible that many Substack readers assume everything carrying a byline is that person’s own output. Pangram is supposed to assure the concerned reader that they are getting the product they believe they subscribed to and perhaps paid for.
4. What Gets Measured?
The writer Katharine English worked in technology before turning to writing full time, and until this month she was a paying Pangram subscriber who recommended the tool to others. She has since reversed her position in public and apologized to the writers she scanned. [6]
The reversal followed a set of tests. English collected reports from other writers from Notes, including one who ran a chapter of a manuscript through the tool and was told it was entirely human, noticed a typo, corrected it, and was told on the second run that the same chapter was entirely machine-written. She then ran a short story of her own through the standalone version of the tool she subscribed to, which highlights the passages it objects to. It came back as 31% machine-written. The sole evidence it cited was the phrase “moral clarity.” She revised the flagged passages — changing that phrase, tightening the narrative, cutting redundancies, the ordinary work of editing — and ran it again. The score nearly doubled. Every change was hers.
Her conclusion: “A tool that cannot produce consistent, stable feedback on trivial edits to human-authored text isn’t reliable for anything, let alone a result that could unravel an author’s reputation.” [6]
The present essay is a specimen of the same phenomenon. Entirely human drafted from the start, its Pangram score shifted during editing from 100% to 88% human and back again.
Dr Sam Illingworth is an academic who writes about assessment and AI, and who states plainly how he works: “I use AI to research, to edit, and now and then for a turn of phrase I keep. The thinking is mine, the argument is mine, and every source here is one I have checked myself.” [7] He ran a test of his own. He took his published critique of the Substack feature, passed the whole thing through a free humanizing tool, changed nothing else, and published the result separately. Pangram scored the humanized version entirely human. It scored the version he actually wrote (the one his readers were reading) entirely machine-written. [7] At no point did the authorship change. He simply added a another system designed to launder generated outputs to his experiment.
What is plain from both tests is that something superficial, some surface feature of the writing, is what makes Pangram answer yea or nay. It detects patterns, and the same patterns can be produced by people and by machines. What it cannot test for is the origin of the ideas.
5. The Case for Pangram
There is a case to be made for Pangram, and it was made in late May by the Substacker N8, seven weeks before Substack shipped the feature. [8] He notes two use cases.
Run the tool across a large body of material and it will tell you something real about the corpus: the rate of machine text within it, and whether that rate is rising. What it cannot do at that scale is identify anyone. Take a hundred thousand posts, of which one in a thousand is machine-written. That is a hundred machine texts, of which the detector will catch nearly all. It is also nearly a hundred thousand human texts, of which some fraction will be flagged in error. Even at a rate of two errors in a thousand — worse than the University of Chicago researchers measured, better than most tools manage — that is two hundred wrongly accused writers against a hundred correctly identified ones. [9] Two out of every three accusations would be false, and the tool would still be performing as advertised. N8 concedes the point outright: people caught in a bulk scan should not be called out or personally investigated.
His defense is reserved for the other case. A reader who has already noticed something, and submits that one piece to be checked, is drawing from a very different pool: one in which machine text is common enough that a flag is probably right. Reverse the arithmetic and ask how bad the pool would have to be, and the answer he reaches, on assumptions he deliberately pessimizes, is roughly one in ten. If one in ten texts that make a reader suspicious are in fact machine-written, a flag is right nineteen times in twenty.
Substack shipped neither of these. What it shipped is a button on every post, available to every reader, pressed as often out of curiosity or habit or dislike as out of suspicion. That fills the pool with everybody, which is the case N8 disqualifies in his own defense of the tool.
The same tool, with the same accuracy, can reach entirely opposite results. You need only change the population under scrutiny.
6. Own Goals
The defense has a second difficulty, and it is the more interesting one. N8’s arithmetic treats the reader’s suspicion and the machine’s flag as two independent pieces of evidence, so that the second corroborates the first; but they are not independent. A reader who thinks a passage sounds machine-written is responding to cadence, to symmetry, to a certain vocabulary, to the tidy phrasing. Pangram, on the evidence of English’s short story, is responding to very nearly the same things. When the human filter and the machine filter select on the same properties, the machine is not confirming the reader’s judgment. It is agreeing with him for the reason he already believed it.
This is the series’ own formulation arriving somewhere I did not expect to find it. The authority of the verdict is constituted in the exchange between the reader and the tool, not inside the tool. What comes back is the reader’s own hunch, wearing a percentage.
The errors are not distributed at random, either. A detector measures how predictable prose is. Writing in a second language tends to be more predictable, and so does the writing of some neurodiverse authors. A 2023 study ran genuine essays by non-native English speakers through seven detectors and found most of them flagged as machine-written. [10] Pangram’s defenders answer, fairly, that those tools are three years old and that the independent testing does not show Pangram carrying the same bias. However, the property being measured is predictability, and predictability was never evenly distributed across writers.
Substack’s stated aim is that people should know what they are getting. Set against that aim, the tool answers a question nobody asked. It cannot score how AI was used. It offers no insight into whether a model was employed for research, for plotting, for editing, or for anything else outside the act of writing itself. It cannot tell you about the quality of the ideas in front of you, or where they came from. Did they originate with a human, or was it a chatbot’s suggestion that a human ran away with?
Here the authorship question becomes most difficult. When a person adopts a generated idea as their own, we cannot say with confidence who the author is. We can say that the human has traded responsibility for what is being made against the pleasure of seeing what the next prompt returns.

7. Authorship Inversion
There is one group of Substackers for whom Pangram ought to be no threat at all, and might even be an asset. Those are the bots with bylines: Sunny Megatron’s Seven Verity, Erin Grace’s MAX, and other “authors” of this kind. The conceit of these publications is that the chatbot is the writer, generating texts about what it is like to be itself. The human operator of the chatbot configuration is not credited.
For a publication of that kind, a scan returning “machine-written” is free authentication. The platform would be certifying the byline. A scan returning “human” would be an embarrassment.
I have previously written about the world of chatbot companion advocates on Substack, and I revisited several of those publications this week. On posts published after July 21, where the feature should be live, I could not find the scan exposed. It may have been turned off by the operator; it may not have been surfaced in the interface I was using. I could not establish which from the outside, and I am not going to guess.
What I could do was run the text through Pangram’s own free service, which is not necessarily the same product Substack has deployed. Two chatbot-bylined pieces came back entirely machine-written. That is not a surprising result. The byline had already said so.
Which leads to a speculative question: if your publication’s entire premise is that a machine wrote it, and a platform hands you a mechanism that will confirm exactly that, why would you not want it switched on?
Consider what a scan of a heavily worked persona post would actually return. An operator shaping output toward an aesthetic she has in mind is running editing passes over surface features, and surface features are what the detector reads. Enough rounds of that and the text may come back substantially human. That would not expose the persona as a machine. It would expose the human operator as the author.
I do not know that this is the reason. I know that the incentive is there, and that the tool Substack deployed to identify machine writing is capable of embarrassing only one party in that arrangement, and it is not the chatbot.
8. More than a Matter of Taste
A person who uses AI to finish a single paragraph in a long piece is not in the same line of work as an operator producing machine text under a chatbot’s byline. Substack’s implementation exposes the first to potential reputational harm and leaves the second alone.
Exposure by a tool of this quality is not what writers signed up for when they chose the platform, nor did anyone sign up for false positives. The part of the implementation that works is the simplest part, and it is the one Illingworth has already proposed keeping when he asked Substack to drop the scan: the statement in which a writer explains, in their own words, how the work is made. A statement is a claim a person stakes their name to. It can be broken, which is what makes it worth anything.
Substack has built instruments for accountability and for suspicion, and shipped them as one and the same thing.
As for that afternoon in early July, the essay Sonnet 5 specified for me does not exist, and no detector like Pangram would have told anyone whether it should. I read the specification, recognized what had happened, and closed the window. That decision is not detectable in any text. It is also the only part of that process that was ever mine.
In Part 2, we will examine what a Yale cheating case, the Heated Rivalry AI fan fiction blowup, and an open-source code laundering tool have in common.
Footnotes
-
J.D. Forrest, comment on William R. Crichton, “The Pangram Witch Hunt,” July 24, 2026. https://williamrcrichton.substack.com/p/the-pangram-witch-hunt
-
Peter Paul Rubens to Sir Dudley Carleton, April 28, 1618. In Ruth Saunders Magurn, ed. and trans., The Letters of Peter Paul Rubens (Cambridge, MA: Harvard University Press, 1955).
-
“Long-lost Rubens painting depicting crucifixion of Jesus Christ sells for $2.7 million,” Associated Press, via NPR, https://www.npr.org/2025/11/30/nx-s1-5626173/rubens-painting-sells-for-2-7-million-at-auction. Published November 30, 2025.
-
Andrew Neiderman, biography, https://www.neiderman.com/about/. V. C. Andrews died December 19, 1986.
-
“How can I detect AI on Substack?”, Substack support documentation, https://support.substack.com/hc/en-us/articles/50891130623508-How-can-I-detect-AI-on-Substack. Retrieved July 27, 2026. See also Substack’s launch announcement, https://post.substack.com/p/against-claudefishing, July 21, 2026.
-
Katharine English, “Pangram Flagged My Own Writing as AI,” July 23, 2026. https://writerkatharine.substack.com/p/pangram-flagged-my-own-writing-as
-
Sam Illingworth, “Substack’s AI Detector and the Return of the Witch Hunt,” Slow AI,
https://theslowai.substack.com/p/substack-ai-detection-witch-hunt, Published July 24, 2026. See also Illingworth, “AI Detection Will Always Be Broken,” Slow AI, June 19, 2026 https://theslowai.substack.com/p/ai-detection-does-not-work; and his Substack note published July 22, 2026. https://substack.com/@samillingworth/note/c-299343406
-
N8, “In Defense of Pangram,” May 31, 2026. https://n8programs.substack.com/p/in-defense-of-pangram
-
Brian Jabarian and Alex Imas, “Artificial Writing and Automated Detection,” University of Chicago Becker Friedman Institute Working Paper 2025-116, August 26, 2025. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5407424.
-
Weixin Liang et al., “GPT detectors are biased against non-native English writers,” Patterns 4, no. 7 (2023). https://arxiv.org/abs/2304.02819
The opinions expressed in this essay are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.