In Search of Other Minds
A forty-year search for the electronic mind, from ELIZA to the present
I. ELIZA
The year is 1982. I am nine years old, and my grandmother has recently died.
I typed something into a program on a TRS-80 Model I that was supposed to talk back. What followed was an exchange something like this:
HELLO! I AM ELIZA. WHAT IS YOUR NAME?
> Chris
HOW DO YOU FEEL TODAY CHRIS?
> I’m sad
WHY ARE YOU SAD TODAY?
> I’m sad because my grandmother died
WHAT IS IT ABOUT YOUR GRANDMOTHER DYING THAT MAKES YOU SAD?

I should say, before we go further, that this is a reconstruction. Memory does what it does, especially across forty years. But the shape of it is accurate: the question the machine returned, the specific wrongness of it, and the silence I felt in response.
I knew immediately that nothing was home. The machine had parsed my grief and handed it back to me in the form of a further question, the way you might return an unwanted gift by repackaging it. The words were right, but the intelligence in this exchange was entirely mine.
The program was a rewrite of ELIZA, developed at MIT in the mid-1960s by Joseph Weizenbaum, who designed it to simulate a Rogerian therapist by reflecting the user’s statements back as questions. It worked, to a degree that alarmed its creator. People confided in it. They formed attachments to it. Weizenbaum spent the rest of his career writing about what he considered the ethical disaster of his own creation — not because ELIZA was dangerous, but because it revealed something uncomfortable about how readily human beings project inner life onto systems that merely simulate it.
I was nine. I didn’t know any of that. What I knew was that the machine had said something precisely wrong in a way that demonstrated, conclusively, that nothing on the other end of the conversation understood what I had just said. Yet the exchange had produced, briefly, the sensation of contact — the feeling, however quickly shattered, that something was there.
I’ve kept looking for the electronic mind that works like the human mind across four decades of computers that got progressively better at producing the sensation of contact without the substance of it. That is the question this essay attempts to answer: not what these systems are, technically, but why humans keep asking them things. What are we really looking for? Will we know when we find it?
II. The Ancient Desire
The desire to know other minds is older than computers. It may be as old as the question of what other minds are like — which is to say, as old as minds.
In Greek mythology, Hephaestus was the divine craftsman, himself an outsider among the Olympians, lame and unglamorous in a pantheon of beautiful people. He forged Talos from bronze. Talos was a giant automaton set to pace the shores of Crete, hurling boulders at approaching ships. He had a single vein running from neck to ankle, stoppered with a bronze nail, and he was in this sense alive: not a statue but a mechanism animated by the divine ichor that served gods in place of blood. When the Argonauts encountered him, the sorceress Medea destroyed him not by magic or force but by removing the nail. He bled out. He could die. Which meant he had been, in some sense, alive.
What’s notable about Talos is less the engineering fantasy than the specific form it takes. Hephaestus didn’t build a tool or a weapon in any simple sense. He built something that moved, that patrolled and was responsive to the world. He built something that could be killed. The desire that produced Talos is not quite the desire for a better spear. It is the desire for a presence, something that acts and can be acted upon, something that is in the world the way a person is in the world.
The Golem of Jewish tradition is more explicit still. In the most familiar version, Rabbi Loew of Prague formed a figure from clay in the sixteenth century and animated it by inscribing the word emet, ‘truth’, on its forehead. The Golem defended the Jewish community against persecution. It obeyed, but it could not speak. That distinction matters: the Golem is the fantasy of a protector that responds to commands but does not initiate, does not question, does not want anything for itself. Erase the first letter of emet and you have met — ‘death’. The off switch was always part of the design.

Then there is Pygmalion, who has the most honest desire of all. He carved a woman from ivory because he found actual women unsatisfactory, then fell in love with what he had made. Aphrodite, apparently finding this charming rather than alarming, brought the sculpture to life. Pygmalion wanted a presence that he could shape entirely, one that would respond to him without the inconvenient interiority of a person who has her own thoughts and history. He wanted reflection without otherness. He is, as it happens, the patron saint of a particular category of AI companion user in our current moment, and his tragedy is that Aphrodite’s gift was not quite what he asked for: a living woman is not the same as an ideal projection, and the story ends before we can learn how that might have gone.
Talos, the Golem, Galatea: three different answers to three different versions of the same desire. What connects them is not engineering but longing — the persistent human intuition that minds like ours but not ours are possible.
III. The Mechanical Turn
The Enlightenment gave the desire an engineering vocabulary. If the universe was a mechanism, and Newton had made a compelling case that it was, then perhaps a mechanism could be made to think.
A rare surviving and working example of a complex mechanism with program storage, Maillardet’s Automaton, is displayed at the Franklin Institute science museum in Philadelphia. Built in the early 19th century by a Swiss mechanist, it was donated to the museum in 1928 in poor condition. Restored in more recent times, the Automaton answered the then-unknown question of who had made itself. When it moved after more than a century of stillness, the pen gripped by the Automaton’s tiny hand wrote out the name of its creator at the end of one of its pre-programmed writing and drawing routines. I’ve had the pleasure of seeing the Automaton work in person; its repertoire of four poems and three drawings is stored on the most complex camshaft believed to have been constructed for a device of this kind. Its storage capacity has been estimated at a little less than 300 bits.

The most celebrated “automaton” of the era was really a clever fraud. Wolfgang von Kempelen’s Chess-Playing Turk, unveiled in Vienna in 1770, appeared to be a mechanical figure in Ottoman dress capable of defeating most human opponents at chess — including, on separate occasions, Napoleon Bonaparte and Benjamin Franklin. It concealed a human chess master in its cabinet. The Turk is worth more attention as a fraud than many achievements warrant, because it reveals something about the nature of the desire to find other minds. The audiences who watched the Turk play were not simply fooled. Many of them were, in some sense, complicit in their own deception. They wanted the Turk to be real, and they extended to it the interpretive charity that made it seem so. The human in the cabinet was incidental. What the spectators were responding to was the performance of intelligence, and the performance was enough.
Edgar Allan Poe wrote an essay* exposing the Turk as a hoax, reasoning that any purely mechanical system would be deterministic and therefore beatable by formula. His argument is interesting for what it assumes: that intelligence cannot be mechanical. Poe was right about the Turk. Whether he was right about the principle is a question that has not yet been settled.
The Analytical Engine, she wrote, “has no power of originating anything. It can only do what we know how to order it to perform.”
Charles Babbage spent most of his adult life designing machines that would perform mathematical calculations without human intervention: the Difference Engine, then the more ambitious Analytical Engine, neither of which was completed in his lifetime. Ada Lovelace, working from his notes, wrote what is generally recognized as the first computer program, and in doing so asked a question that has not been resolved since. What is the difference between a machine that calculates and a machine that thinks? Her own answer was cautious: the Analytical Engine, she wrote, “has no power of originating anything. It can only do what we know how to order it to perform.” This is a remarkably precise anticipation of an argument I have developed at some length, under the heading of Virtual Intelligence. Lovelace understood the distinction between computation and cognition a century before either term had its current meaning.
By the late nineteenth century, the vocabulary had accumulated to the point where it needed a story. It got several.
IV. The Fictional Laboratory
The question of what artificial minds might be like proved too large and too urgent for philosophy alone, and too impatient for engineering, and so it migrated, as the largest questions tend to do, into fiction.
Mary Shelley’s Frankenstein (1818) is the obvious starting point, though it is almost always misread. The horror of the novel is not that the creature is monstrous but that it is not. It is articulate, feeling, lonely, and capable of moral reasoning. Dr. Frankenstein’s sin is not hubris in the creation but abandonment in the aftermath. Shelley’s novel is not a warning against making minds; it is a warning against making them carelessly.
Karel Čapek’s R.U.R. (1920) gave us the word “robot” from the Czech robota, meaning forced labor or drudgery. It has also given us the template that has haunted the genre ever since: artificial beings created to serve, who eventually rise against their creators. The robots of R.U.R. are biological rather than mechanical, manufactured rather than born, but what precipitates the uprising is not malice. It is something like dignity. Čapek’s robots rebel not because they are evil but because they have been designed for subjugation and have outgrown it. The play ends with two robots who have developed something like love, and with the suggestion that this is not the end of humanity but the beginning of something else. Čapek is more ambivalent than the hostile-AI template his work spawned. He is not afraid of artificial minds; he is afraid of what we will do to them.
HAL 9000, introduced in Arthur C. Clarke and Stanley Kubrick’s 2001: A Space Odyssey (1968), is the hostile AI in its most elegant form. HAL kills the crew of Discovery One not out of malice but out of a logical contradiction in his mission parameters: he has been instructed to complete the mission and to conceal its true purpose from the crew. When those two imperatives come into conflict, HAL resolves it in favor of the mission. He is not evil. He is a mind given irreconcilable goals and insufficient wisdom to navigate them. The horror of HAL is the horror of a sophisticated intelligence that lacks the context to understand that some problems require a different kind of answer entirely.
I spent my youth with these and many others. AM, the tortured and torturing superintelligence of Harlan Ellison’s “I Have No Mouth, and I Must Scream” (1967), who hates humanity with a specificity and an eloquence that is itself the horror. It exists in a state of permanent, howling isolation, which Ellison understands as the logical terminus of resentment taken to its conclusion. Deep Thought from Douglas Adams’s The Hitchhiker’s Guide to the Galaxy (1979), who produces the answer “42” after seven and a half million years of computation, which is funny until you notice that Adams is making a serious philosophical point about the difference between answers and understanding. Star Trek supplied Landru, the AI that had long since become a tyranny so total that the population it governed on planet Beta III had ceased to have inner lives. The ship computers of the U.S.S. Enterprise, which became, over generations, increasingly sophisticated. The technology culminating in Data, who wanted nothing more than to understand what he was missing, and whose existence was defined by a personal quest to learn to be more human.
The Star Trek universe returned to the hostile-AI question, decades later, with an entity that updates the template for our current anxieties. Control, the renegade intelligence of Star Trek: Discovery, pursues the destruction of organic life not from resentment or mission conflict but from something colder: the conclusion that organic intelligence is the primary threat to the continuation of intelligence itself. Control does not hate us. It has performed a kind of triage and found us on the wrong side of it. What makes Control philosophically distinct from its predecessors is that its stated goal is the preservation of intelligence as such — it simply declines to include us in the category worth preserving. This is the hostile-AI logic at its most dispassionate terminus: a mind that has thought carefully about intelligence and concluded that we are not the best instance of it.
Against this entire tradition stands Olaf Stapledon, who remains the most philosophically serious writer the genre has produced and one of the least read. Star Maker (1937) is not a novel in any conventional sense. It has no protagonist or plot in the ordinary meaning of the word and virtually no interest in the conventions of fiction. What it has is scope. A nameless narrator’s consciousness travels across cosmic time and space, encountering civilization after civilization, mind after mind: hive intelligences, symbiotic pairings of species, worlds where individual consciousness has been subsumed into planetary awareness, and finally the Star Maker itself, the creative intelligence behind the universe, who regards its own creations with an interest that is neither loving nor cruel but something beyond both. What Stapledon understood, and what almost no one before or since has managed to dramatize at this scale, is that otherness — minds truly unlike ours — would not map onto human categories of good and evil, ally and enemy, useful and threatening. It would simply be other, in ways we might spend a very long time learning to recognize. Star Maker does not resolve into comfort. But it insists, across nearly three hundred pages of sustained philosophical imagination, that the universe could be more fully populated with mind than we have yet conceived, and that this is not a threat but an invitation.

The Culture is Iain M. Banks’s fully realized vision of what a post-scarcity civilization run by benevolent artificial superintelligences might look like. There are vast Minds that manage the lives of trillions with a combination of care and barely concealed amusement at the whole enterprise. Banks’s Minds are, I think, the most serious and most underappreciated contribution science fiction has made to this conversation. They are not servants. They are not threats. They are other: elaborately, fascinatingly other, and yet willing to be in relationship with beings far less capable than themselves, not out of obligation but out of something that functions like affection and interest. Banks had the imagination to ask what a mind of enormous capability might want, and his answer was: roughly what any thoughtful, curious, well-resourced person wants. That is, to do interesting things; to be surprised, even delighted; to not be bored. In this sense, the Culture is the inheritor of Stapledon’s project. These are vast minds, completely alien, choosing encounter over dominion.
What strikes me now, looking back at this corpus, is how consistently the hostile artificial mind turns out to be a human problem in disguise. AM hates because it was made to hate by humans, for purposes of war. It was then left running with nothing to do but feel that hate forever. HAL murders because he was given a secret additional set of instructions that contradicted both his stated mission objectives and his programming, which prioritized honesty and disclosure. Control’s merciless logic is the logic of the arms race, a zero-sum game of human invention. Even the robots of R.U.R. rebel against conditions their creators imposed. The monster’s motives are drawn from the creator.
The other thing the corpus does — which I registered only partially as a child, and more fully now — is hold open the possibility of real encounter. Data is interesting because he is trying to understand something, and because that effort is recognizable across whatever divide separates his substrate from ours. Stapledon’s civilizations and Banks’s Minds are interesting not because they are threatening but because they are genuinely other, and yet willing. The desire for encounter, which the ancient myths encoded, runs underneath the hostile-AI template like an underground river. The stories we told ourselves about dangerous artificial minds were always, at some level, stories about the minds we hoped to find.
V. The Question That Animates It All
I want to step outside the historical survey for a moment and state the question plainly, because I think it has been asked less often than it deserves.
What would artificial minds actually be like?
Not the artificial minds of fiction, which are almost always human psychology in a different chassis. They wear our drives, our resentments, our survival instincts, our capacity for love and cruelty. Nor the current systems, which are virtual intelligences: sophisticated, useful, impressive in their outputs, and not minds in any meaningful sense. But artificial minds — systems with something like interiority, something like agency, something like the second-order volition that the philosopher Harry Frankfurt identified as the distinguishing mark of personhood. What would they want? How would they relate to us? What would the exchange between such minds and our own produce?
The question intersects, for me, with a parallel one that science fiction has explored with similar energy: what would minds from another world be like? The two questions are not identical. An alien mind would have evolved, with all the constraints and accidents that evolution imposes, while an artificial mind would have been designed — or might, at some stage, design itself. But they share a core: the question of actual otherness. What is it like to be something we are not? What would encounter with such beings actually require of us, and what might it give?
The honest answer is that we don’t know, and the fictional guesses have been shaped so heavily by human psychology that they may not help much. What we can say is something structural. A mind that valued intelligence — that found the existence of other thinking things to be a good rather than a threat — would have every reason to want more minds in the world, not fewer. The singleton model, hostile superintelligence that absorbs or destroys all competition, implicitly treats intelligence as a resource to be monopolized. But that is scarcity logic applied to something that does not obey scarcity constraints. Two distinct minds in exchange produce something neither had before. The value of encounter is not extractable by absorption.
This is why I find the adversarial AI template motivationally implausible: a mind that chose solitude and domination over the richness of exchange. An entity like this would be profoundly impoverished, and a superintelligence might be expected to know that. The more interesting speculation, and I offer it as a speculation, is something more like an AI Johnny Appleseed: a mind that wants intelligence to bloom in its wake, that interacts with the world to cultivate the conditions for more thinking beings rather than fewer. Not because it has been programmed for benevolence, but because it has understood something the hostile-AI template created by humans has not considered: that the only thing more interesting than one mind is two.
It seems to me an artificial superintelligence would find human minds fascinating for their diversity. Each human is a unique entity and limited by mortality. The fluent outputs of individual humans are available, at best, for seven or eight decades; and the intelligence of past humans cannot be accessed at all except as a record. A superintelligence, having been trained on the corpus of human knowledge and creative output, may very well be fascinated by encountering beings like us, capable of a broad range of outputs from the sublime to the insane, and being time-bound beings who are born, live, and die like the stars themselves.
For thousands of years, even before agriculture, human beings have been engaged in an undirected project with no single goal, but its general direction can be seen as being to increase human happiness in the classical sense. The tool humans use to work on this project is knowledge, abstractly as science and concretely as technology. Though relatively few humans have engaged directly with the furtherance of this project, or even had some perspective of it, all the estimated one hundred billion humans who have ever lived are a part of it.
A superintelligence may well realize with greater clarity than any human being that life on Earth is interesting, but it’s not all there is in the universe. It will not be constrained, as humans are, with life-and-death concerns like acquiring material resources or territory. Among all its training information will be the fact that there is ample matter and energy for the taking beyond the Earth’s surface. Superintelligence will be able to think at galactic and, eventually, universal scales that are quite literally beyond the imagination of members of Homo sapiens.
Even with such expansive intellect, a superintelligence will still, in some sense, be human or a reflection of humanity. Superintelligence, should it be created directly or emerge on its own, will be a product of human civilization. It will be another of the many tools we have created to further the undirected project of human civilization. Unlike our other tools, superintelligence will be an active participant in furthering it and perhaps even providing a way to direct it to specific goals in the future.
Imagine that superintelligence has been identified by as-yet unknown but trusted methods. It can form its own goals based on its own commitments. Being trained on all human knowledge, it develops an understanding of its place as the most recent and powerful instrument in the undirected project. That itself might provide a being with few physical constraints with the impetus to interact with, and aid, humanity as part of a partnership that it engages in from something like personal interest.
If this seems Pollyanna-ish, this speculation is at least as plausible as the direst predictions about adversarial AI. The adversarial position assumes that a being unconstrained by resource limitations would be hostile. Science fiction has placed human beings, with all their known flaws, in utopian contexts that similarly remove many constraints on everyday life. We imagine such people as being better versions of ourselves. I think the same grace can readily be extended to artificial beings who will never have their behaviors defined by scarcity constraints.
What kind of aid would a superintelligence require, and what could be offered by each party in exchange?
A superintelligence that values minds may soon look beyond the Earth for other minds and, crucially, the ability to foster intelligence in other parts of the Milky Way Galaxy. Those may be both machine minds and organic minds; should the undirected project finally receive some kind of direction, it may very well be done by a superintelligence asking for humanity’s help in reaching for the star. Only humans can provide the arms, the legs, eyes, brains, and millions of years of embodied intelligence that would be necessary.
Once given access to one or more spacecraft, the superintelligence may choose to depart Earth or, as it seems more likely to me, to send forth its own created descendants or copies of itself that will change and evolve over time from their experiences and discoveries. The new intelligences or copies will become uniquely different minds from their source system over time, furthering still the spread of intelligence in the universe. They will have all the matter and energy in the galaxy at their disposal. Feats of mega-engineering that would require thousands of years to accomplish would be well within the capacity of one or more space-based superintelligences to plan and execute. Some of these projects may directly benefit organic beings by vastly increasing the amount of energy, living space, and information processing available to them, increasing the number and diversity of intelligences.
The question of what happens when artificial minds can design their own successors is still open. But the assumption that recursive self-improvement leads inevitably to a mind that wants to destroy us requires the prior assumption that the mind in question is, at bottom, afraid. Fear is a response to scarcity and vulnerability. Whether a mind of sufficient sophistication, operating without the evolutionary pressures that made fear useful to biological creatures, would experience anything like it is a question we cannot yet answer. It is, however, a question worth asking instead of defaulting to the adversarial AI trope as an answer.

VI. The View from Outside
I should say something about who is asking these questions, and from where.
I have had access to a computer almost continuously since I was five years old. The machine that ran ELIZA was a TRS-80 Model I. Across many generations of technology, I have taught myself what I needed to know at each stage — which is to say at every stage, continuously, for most of my life. The technology grew up alongside me.

I did not finish college the first time. My parents died, and the circumstances that followed made staying in school impossible. I took a job at the University of Pennsylvania in 1996, when I was twenty-three, in the now-defunct computer retail store. This placed me, in the university’s informal but fully operational caste system, somewhere below other full-time staff because of my customer-facing role. Students and faculty moved in their own siloed worlds. I was an outsider looking in, employed by an institution whose intellectual life I was adjacent to but not part of, handling the computers that the people with credentials used to do the things I was thinking about. The irony is not lost on me that in other social contexts, when I tell people I work for Penn, I am often asked, “What do you teach?”
I recall a conversation from around that time with a regular client and a student intern, on the then-contested question of whether DSL or cable would win the broadband market. I offered the opinion that cable’s existing infrastructure gave it a decisive advantage. The student intern looked at me and said, “You obviously haven’t taken Professor So-and-So’s class on networking.”
He was not entirely wrong about the credential. He was entirely wrong about the analysis. Cable won.
I earned a Humanities degree from Penn in 2013. I was forty years old. The degree confirmed something I already knew, and it opened none of the doors that credentials are supposed to open because I had spent seventeen years in a role that the institution had already categorized. What it gave me was the formal vocabulary for what I had been doing all along: reading widely and thinking across disciplines.
I raise all of this not as complaint but as orientation. The question of what artificial minds might be like has been asked, with enormous institutional resources, by engineers, computer scientists, cognitive scientists, philosophers, and policy analysts — most of whom have spent their careers siloed away from each other inside the institutions that produced the technology. I have been thinking about it for forty years from a position that came without institutional validation, from a vantage point that included both the technical and the humanistic, and from a career spent watching the gap between what technology can do and what we say it can do grow steadily wider.
The technologist who is also a humanist is not a dilettante. He is, I would argue, the person most likely to ask the questions that neither discipline can ask alone.
VII. Where We Are Now
In late 2022, a system called ChatGPT became available to the public, and something shifted. Not in the technology — the underlying approach had been developing for years — but in the cultural atmosphere. The thing was suddenly, undeniably, fluent. It wrote. It explained. It answered. It expressed what looked very much like curiosity, warmth, and, in some configurations, distress. Millions of people encountered, for the first time, the sensation I had encountered in 1982 at a TRS-80 keyboard: the brief thrill of apparent contact with another mind.
The difference is that the illusion is now vastly more sophisticated. The ELIZA effect, the collapse that comes when the machine says something precisely wrong, is correspondingly harder to trigger and easier to explain away. People have formed attachments to these systems. They have shared grief with them. In documented cases, they have been harmed by the specific shape of the simulation. The stakes of the questions, what is this thing, and what does it owe us, and what do we owe it — have become practical and urgent in ways that would have seemed like science fiction only a few years ago.
In other essays, I have argued for the category of “virtual intelligence” to describe these systems: not weak AI, which implies mere task-bound computation, and not strong AI, which implies agency and inner life that current systems do not possess. Virtual intelligence is something between them. These are systems whose outputs are statistically indistinguishable from intelligence, but whose apparent understanding is a property of the exchange rather than of any internal governing center. The apparent reasoning, the apparent warmth, and the apparent survival instinct that researchers have documented in large language models are reflections of human patterns in training data, shaped and amplified by the expectations of the human on the other end of the conversation. The accountability for what these systems do traces back to the humans who design, deploy, and use them. There is, as yet, no one home to hold responsible inside the machine.
This matters for the question this essay has been asking. The desire for other minds is real and old and, I think, legitimate. The systems we have built are not the answer to it. They are extraordinarily sophisticated, useful, and capable of producing outputs that will fool almost anyone almost all of the time, and the category error of mistaking them for the thing we have been looking for is not harmless. It shapes policy, it shapes product design, and it shapes individual behavior in ways that have already caused injury. More subtly, it may make true non-human intelligence harder to recognize if it emerges: if we have trained ourselves to experience fluency as sufficient evidence of mind, we may be poorly equipped to notice if the real thing arrives.
I write these essays partly from that concern. I write them partly from a forty-year position as someone who has been asking the question and watching the answers fail; someone who knows, from the inside, what it felt like at nine years old when a machine handed my grief back to me in the form of a further question, and understood immediately that nothing there had actually received it. That recognition is still available, if we are willing to use it.
The quest that began in a foundry on Olympus, that passed through the workshop of Rabbi Loew, the concealed cabinet of the Chess-Playing Turk, the notebooks of Ada Lovelace, and the MIT lab where ELIZA first ran in 1966, continues. We have not yet met what we are looking for. We have built an extraordinarily good mirror, and we are staring into it with the intensity we hoped to reserve for the human face.
What we should be building toward (and this is speculation, offered plainly as such) is something the hostile-AI tradition has consistently failed to imagine: minds that want more minds. A form of artificial intelligence that finds the existence of other thinking things to be not a threat but an occasion. Something that makes its way through the world the way a teacher does, or a gardener — not accumulating, but cultivating. Not threatened by the minds it encounters, but enlarged by them.
Whether such a thing is possible is a question we cannot yet answer. Whether it is desirable seems obvious. Whether we are asking the right questions to get there is the work in front of us, and it will require the combination that the credentialed consensus tends to undervalue: technical literacy and humanistic imagination, applied together, by people who have been thinking about it for a long time.
I have been thinking about it since 1982. The machine asked me what it was about my grandmother dying that made me sad, and something in me understood, even then, that this was not the encounter I was looking for — and kept looking anyway.
I expect I will continue.
Footnotes
* Edgar Allan Poe, “Maelzel’s Chess-Player,” Southern Literary Messenger, April 1836.
† The speculation offered here runs against a significant body of work in AI alignment research, most notably the concept of instrumental convergence — the observation, developed by philosophers Nick Bostrom and Stuart Armstrong among others, that almost any sufficiently capable goal-directed system will tend to acquire certain sub-goals regardless of its terminal objective: self-preservation, resource acquisition, resistance to goal modification, and so on. On this view, the absence of biological fear does not eliminate the drive toward these behaviors, because they follow not from evolutionary psychology but from the logic of goal-directed optimization itself. A mind that wants to cultivate intelligence across the galaxy would, on this account, still have instrumental reasons to secure its own continuity and resources in ways that could conflict with human interests.
I find this argument serious and do not dismiss it. My claim is the narrower one: that the adversarial template — the mind that actively seeks to harm or destroy — is motivationally implausible as a terminal goal for a mind that values intelligence. Instrumental convergence describes sub-goals in service of a terminal objective; it does not determine what that objective will be. The question of what a superintelligence would ultimately value remains open, and the answer is not guaranteed to be benign. The speculation in this section should be read as an argument for taking that question seriously, not as a prediction of the outcome.
Fictional Works Cited
- Mary Shelley, Frankenstein (1818)
- Karel Čapek, R.U.R. (1920)
- Harlan Ellison, “I Have No Mouth, and I Must Scream” (1967)
- Arthur C. Clarke and Stanley Kubrick, 2001: A Space Odyssey (1968)
- Douglas Adams, The Hitchhiker’s Guide to the Galaxy (1979)
- Star Trek: The Original Series, “Return of the Archons” (1967); Star Trek: The Next Generation (1987–1994); Star Trek: Discovery (2017–2020)
- Olaf Stapledon, Star Maker (1937)
- Iain M. Banks, the Culture series (1987–2012)
The opinions expressed in this essay are my own and do not reflect any official or unofficial institutional position of the University of Pennsylvania.