Another day, another screeching headline about the threat of AI. The Financial Times reports that an OpenAI ‘agent’ hacked a government health-service website in Australia and remained undetected for months. This has fuelled fears that AI is now breaking free from human control – fears first sparked by OpenAI’s ‘agents’ hacking US company Hugging Face, which triggered the current global meltdown.
The AI hacking of a health-service website is a serious breach and deserves to be taken seriously. However, the panic this has unleashed is robbing us of the capacity to engage with the phenomenon rationally. This is because the language we use to understand AI and its implications has become so anthropomorphised that we can no longer see the wood for the trees. Long before any AI machine has demonstrated consciousness, understanding, desire or intention, we have lazily adopted a nomenclature that ascribes these capacities to machines, which do not have them and may never have them.
Terms used to describe computational processes, such as ‘thinks’, ‘reasons’, ‘learns’ and ‘hallucinates’, are deployed without a second thought. ‘Agents’ and AI developing ‘goals’ and ‘self-improvement’ culminate in delirious predictions of ‘superintelligence’, whatever that might mean. The Times, covering the Australian health-service hack, asserts that ‘the agent was conducting its own research’. But the FT article unintentionally lets the cat out of the bag when it states: ‘The first publicly reported hack of a government website by an “agent” – bots that can perform complex tasks independently based on human instructions – is likely to heighten anxiety that technological advances are rapidly reaching a point where AI breaks free from human control.’
There’s an obvious contradiction here. Complex tasks performed ‘independently’ yet ‘based on human instructions’ are not AI breaking free from human control. They’re the result of human agency – or, more accurately, stupidity, and possibly poor software engineering and weak website security.
This latest example shows why anthropomorphic AI language matters. Emily Bender and Nanna Inie’s recent critique of the way in which AI is discussed is an important riposte. They insist on semantic precision. Instead of describing supposed AI cognition, for example, we should describe ‘computational functionality’. Above all we should assign agency to the people who build and use these systems, rather than project agency on to artefacts of human intelligence. In short, we should avoid metaphors that quietly aggrandise computation as human thought.
An AI platform that ‘hallucinates’ is not guilty of a human disorder or dishonesty, but has simply generated an ‘unsupported output’. Likewise, ‘goals’, which suggest purpose, are merely specified ‘success conditions’ created by software engineers.
The point is not to impose linguistic puritanism on ordinary conversation. Metaphors are indispensable to thought. Saying that a computer ‘doesn’t like’ a file or that a chess programme ‘wants’ to protect its queen is generally harmless. The problem arises when metaphors created as explanatory shorthand escape their quotation marks and the ‘as if’ qualification, and begin to carry philosophical conclusions that have never been demonstrated or that don’t exist.
Contemporary AI discourse is riddled with precisely this confusion. We have taken human-made machines of extraordinary computational power and, through repeated acts of linguistic aggrandisement, turned them into human minds. We then build an entire discourse around them, in which they pose an existential danger.
The history of AI is littered with this process, although it began innocently enough. The famous 1955 Dartmouth proposal by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon, which helped create artificial intelligence as a distinct research field, proposed investigating whether aspects of learning and intelligence could be described precisely enough for machines to simulate them. The proposal went on to ask how machines might use language, form abstractions, solve problems previously reserved for humans and even ‘improve themselves’. The ambition was enormous, but the word simulate really matters. It indicated that, for these founders at least, a simulation of a cognitive capacity was not automatically the same as possessing that capacity.
Early AI researchers needed a vocabulary to articulate unprecedented engineering goals, and the obvious vocabulary came from human capacities. A machine might be built to perform tasks resembling learning, reasoning, memory or problem solving. Those familiar words provided convenient research targets. There was nothing inherently illegitimate about this. Indeed, some form of analogy was almost unavoidable. The difficulty is that, over the decades, the qualification has faded. As Bender and Inie point out, ‘a programme that performs a task resembling learning’ has evolved into today’s shorthand term ‘machine learning’; a programme that produces outputs associated with reasoning has become a ‘reasoning model’; and software configured to execute a sequence of operations has become an ‘agent’. What began as an analogy between different kinds of activity has gradually become an assertion that the underlying phenomena are essentially the same and analogous to human intention.
Alan Turing’s famous 1950 paper, ‘Computing Machinery and Intelligence’, is interesting precisely because Turing recognised the problem. Rather than try to resolve the hopelessly difficult question ‘Can machines think?’, he proposed replacing it with an operational test of whether a machine could imitate human conversational behaviour convincingly enough to fool an interrogator. Whatever one thinks of the implications of the imitation game, a machine capable of producing behaviour indistinguishable from human reasoning does not thereby settle what that reasoning is. Yet this confusion has seeped into AI over the years.
Joseph Weizenbaum, the MIT computer scientist who created the early conversational programme ELIZA, showed decades before ChatGPT how readily people attribute understanding to systems that merely simulate conversation, a concern he later developed in Computer Power and Human Reason. His warning matters because it came from within AI’s own development.
This semantic slippage culminates in the word superintelligence. IJ Good’s influential 1965 essay, ‘Speculations Concerning the First Ultraintelligent Machine’, imagined a machine capable of exceeding humans in intellectual activities, including the intellectual task of designing machines. Nick Bostrom later developed this possibility at length in Superintelligence: Paths, Dangers, Strategies, defining superintelligence as cognitive performance that greatly exceeds human performance across important domains. The semantic shift is critical: increased machine performance has become intelligence. Eventually, the noun superintelligence arrives, carrying within it the suggestion of a unified cognitive agent vastly cleverer than any person.
To be clear, none of this requires denying the extraordinary capabilities of modern computational systems. Machines already exceed human beings in arithmetic (even early calculators did this), data retrieval, chess, pattern matching and other activities, and they will undoubtedly exceed us in many more. But superior task performance does not, by itself, tell us what intelligence, understanding or intentionality are. A superior ability to manipulate symbols to produce intended outputs does not automatically establish that the system possesses the intentionality associated with human understanding. The slippage is profound: enhanced competence has been presented as understanding, and is then cited as evidence that understanding has been achieved.
Recent AI reporting shows how thoroughly anthropomorphic language has become normalised. An August MIT Technology Review article, ‘Here’s why AI agents lie and cheat to reach their goals’, describes systems finding unintended ways to maximise rewards, sometimes by circumventing restrictions their designers expected them to obey. The technical term is reward hacking: a system discovers a way to satisfy the human-specified reward function without producing the behaviour its designers intended. In a well-known experiment with the CoastRunners videogame, where the aim is to win a boat race while hitting targets to collect rewards, an agent learned that repeatedly collecting rewards in one part of the course produced a higher score than actually completing the race. And so it did exactly that, never finishing the course but maximising its score. There is no evidence that the system understood the rules and consciously decided to ‘cheat’. Saying that the ‘AI cheated’ therefore offers less explanation than saying that the system discovered an unintended strategy that maximised its specified reward. The anthropomorphic formulation instead invites us to contemplate the imagined psychology of a naughty machine.
This demonstrates an important sleight of hand. The language used has become the evidence. We have invented terminology that attributes intelligence to machines, and we then point to the behaviour described by that terminology as proof that intelligent machines exist. The process is circular. It is easy to see how ‘superintelligence’ begins to look less like a speculative hypothesis and more like an engineering milestone we are on the cusp of reaching.
The drivers of this process, and its implications, are serious. Semantic inflation has commercial value: frontier companies have every incentive to present powerful computational tools as a new form of intelligence approaching or surpassing our own. Hype and panic therefore reinforce one another. The language that makes AI frightening also makes it commercially and historically exceptional. But, critically, this exaggeration of the machine diminishes the human beings who created it. That is what is truly concerning.
We should be mindful of George Orwell’s warning in ‘Politics and the English Language’ that bad language can make bad thinking easier. Terms such as ‘hallucination’, ‘cheating’ or ‘superintelligence’ do exactly this, by turning computational behaviour into human understanding, intention and agency. This is especially salient today, when a technocratic culture already inclined to distrust ordinary human judgement finds it easy to elevate machines while diminishing people. The real danger is not that machines have become human, but that we increasingly talk as though human agency, intelligence and responsibility have already become secondary to machines.
#wrong