Two errors in attributing intelligence: Clever Hans and the mimicry trap
Updated September 2026: this post was rewritten to follow the current version of the preprint.
In Berlin, around 1904, a horse called Hans drew crowds by tapping out the answers to arithmetic questions with his hoof. Experts examined him and found no trickery. It took a young psychologist to work out what was really happening: Hans was watching the people asking the questions. As his hoof approached the right number, they relaxed very slightly, without realising it, and Hans stopped. When the questioner didn’t know the answer, or stood out of sight, Hans failed. He was doing something clever, but not arithmetic.
Now a second story. Over the last century, crows in New Caledonia were seen making hooked tools from twigs, bending wire into hooks in the laboratory to fish out food, and using one tool to fetch another. Each time, sceptics had a simpler explanation ready: instinct, trial and error, habit. It took decades before scientists accepted that crows really can reason about cause and effect.
These stories are usually told as opposites. In the first, people saw a mind that wasn’t there; in the second, they refused to see one that was. My argument is that they are the same mistake made in opposite directions. In both cases, people decided what they were looking at before weighing the evidence. Hans’s audience saw a performance and assumed a mind behind it, without asking how the taps were produced. The crows’ sceptics knew they were “only birds” and never let the performance count.
The second mistake is older than science. In 1773 Phillis Wheatley, an enslaved young woman in Boston, published a book of remarkable poems. Many refused to believe she had written them. Thomas Jefferson admitted the poems were competent but insisted that their author could not be a poet. The work was accepted; the ability it showed was denied, because of who she was. The same move is being made today about artificial intelligence: yes, the system solved the problem, but it didn’t really reason, it only imitated reasoning.
The first mistake has a name, the Clever Hans effect. The second doesn’t, so I call it the mimicry trap: the behaviour is accepted, but relabelled as mere mimicry, so that a verdict decided in advance survives whatever the evidence shows. Today’s AI chatbots are the first case in which both mistakes are being made at once, and on a huge scale, by the public and by experts alike.
Two kinds of mind
When we judge whether something has a mind, we are really asking two questions. Can it think? And can it feel? In every animal we have ever met, the two go together: creatures that seem cleverer also seem to feel more. So we have learned to treat one as a sign of the other. AI systems are the first case in which the two come apart. They are getting better and better at tasks that look like thinking, while whether they feel anything remains a completely open question. Our old habits of judgement, built on animals, don’t work here.
Turing’s test, updated
In 1950 the mathematician Alan Turing proposed a famous test: if a machine can hold a conversation that you can’t tell apart from a human’s, you should credit it with thinking. Turing judged by behaviour alone, which made sense at the time, because nobody knew what a thinking machine would look like inside.
Today we can look inside, at least partly, and that changes things. My proposal is to treat Turing’s test as a piece of reasoning under uncertainty, the kind a doctor uses. A doctor who sees a symptom asks two questions: how well does this symptom fit the disease? And how likely was the disease in this patient before I saw the symptom? Both matter. For AI, the symptom is the behaviour, and the “before” is what we know about how the system was built. Finding something clever inside (say, a hidden map of a board game that the system was never shown) should raise our confidence. Finding a trick inside (a list of memorised answers) should lower it.
Both mistakes come from refusing to let the evidence move that starting point. The Clever Hans mistake sets it at “obviously intelligent” because the system looks and sounds like someone, and never checks how it works. The mimicry trap sets it at “cannot possibly be intelligent” because of what the system is made of, and nothing it does can change that.
A good test for yourself: would you judge the system differently if exactly the same machine, doing exactly the same things, had arrived from outer space and nobody knew who built it? If so, it is your expectation doing the judging, not the evidence.
What would change your mind?
Imagine AI keeps improving the way it has in the last five years. At what point would a sceptic admit that what they are seeing is no longer mimicry? And the enthusiast owes the mirror answer: what would you have to discover about how the machine works to change your mind?
A good answer has three properties. It names something you could actually observe and check (“genuine understanding” doesn’t count, because nobody can check for it). It is fair: the same evidence would count if a person or an animal produced it. And it is stated in advance, not invented after the previous test has been passed.
This doesn’t say who is right. Many sceptics give perfectly good answers to the question, and some enthusiasts give none. What it rules out is a position that nothing could ever change.
The full preprint is on PhilSci-Archive. What I ask of critics is that they say, in advance, what evidence would change their mind. If the answer is “nothing could”, we are no longer having a scientific conversation.