Psychology
The Tower of Hanoi as a Cognitive Test
A Victorian toy became one of the most-used problem-solving tasks in psychology — and then became the standard example of why a good puzzle can make a poor test.
On this page
This page is not a test
Nothing here scores you, and no result from playing the puzzle online means anything about your health. Tower tasks are administered and interpreted by clinicians alongside a history, an interview and other measures — and the research below is largely about how little a single tower score says on its own. If you are worried about your memory, attention or planning, talk to a doctor rather than to a puzzle.
What the puzzle is used to measure
The Tower of Hanoi arrived in psychology because it has a property that is rare in laboratory tasks: it is genuinely hard, it is completely specified, and every step of a person's attempt can be written down. A researcher gets a full record of what somebody did, move by move, against a solution that is known to be optimal — 2ⁿ − 1 moves, no argument.
Herbert Simon used it that way in 1975, in a paper that catalogued the different strategies people actually use — working backwards from the goal, recognising a repeating perceptual pattern, or applying the recursive rule — and showed that several quite different methods produce the same moves. That is the finding the whole later literature keeps rediscovering: the same score can come from very different heads.
The abilities the task is taken to draw on are usually listed as three:
- Planning — looking several moves ahead before touching anything, and having a route in mind rather than a next move.
- Working memory — holding the sub-goals in mind. "Get disk 4 to C" spawns "clear disks 1 to 3 onto B", which spawns another, and the stack of intentions has to be kept somewhere.
- Inhibition — being able to make a move that looks like the wrong one. This is the interesting one, and it has a section of its own below.
Those three are the usual shorthand for executive function, which is why the puzzle is so often described as an executive-function or "frontal lobe" test. That description is a reasonable summary and an unreasonable conclusion, for reasons the last section gets to.
Planning, and the move that goes backwards
Here is what makes this puzzle worth a psychologist's attention rather than any other puzzle.
To finish the game, the largest disk must end on the goal peg. To get it there, every smaller disk must first be piled on the third peg — neither the one it starts on nor the goal. So at the exact moment you are closest to solving the puzzle structurally, the board looks furthest from solved: a tall stack sitting on the wrong peg. A person who is playing by "make the board look more finished" cannot get past that point, and a person who is playing by "hold the plan and tolerate the detour" can.
Psychologists call this a goal–subgoal conflict, and it is the property clinical versions of the task are built around. The number of moves somebody takes matters less than where they get stuck, and a person who repeatedly undoes their own last move is showing something a total score would hide.
If you want to feel the effect rather than read about it, the random-start board is the version that produces it most reliably: from a scrambled position there is no memorised opening to fall back on, and the first move is a decision every time.
Tower of Hanoi vs Tower of London
These two are constantly confused, including in published work, and they are not the same instrument. The Tower of London was introduced by Tim Shallice in 1982, explicitly to study planning impairments in patients with frontal lobe damage — and explicitly because the Tower of Hanoi did not offer enough variety of problems at graded difficulty.
Tower of HanoiThe disks are the constraint
Three identical pegs, each able to hold every disk. What you may not do is put a larger disk on a smaller one — so the restriction travels with the pieces, and it applies wherever they are.
Tower of LondonThe pegs are the constraint
Three balls of one size, and three pegs that hold three, two and one. Nothing about a ball stops it going anywhere — the peg it is going to either has room or it does not, and the task is to reach a pictured arrangement in a stated number of moves.
The difference that matters is where the rule lives. In the Tower of Hanoi the constraint belongs to the pieces, so it applies everywhere on the board and a player must carry it in their head. In the Tower of London the constraint belongs to the apparatus, so it is visible: you can see that the short peg is full. That makes the London task purer as a planning measure — there is less to remember — and it is why the two tasks correlate far less well than their names suggest.
The other difference is the goal. Hanoi has one: everything on the far peg. London hands the patient a picture of an arrangement and a move budget, which lets the examiner build dozens of problems of calibrated difficulty out of one set of equipment.
| Task | Introduced | Pieces | What restricts a move | Goal |
|---|---|---|---|---|
| Tower of Hanoi | Lucas, 1883 — as a toy | Disks of graded size | No larger disk on a smaller one | Move the whole stack to another peg |
| Tower of London | Shallice, 1982 | Balls of one size, three colours | Pegs hold three, two and one | Match a pictured arrangement in a set number of moves |
| Tower of Toronto | Saint-Cyr and colleagues, 1988 | Disks of one size, four colours | No darker disk on a lighter one | Move the whole stack to another peg |
| Stockings of Cambridge | CANTAB, computerised | Coloured balls in hanging socks | Sock capacity | Match a pictured arrangement, minimum moves |
The Tower of Toronto is worth a second look in that list. Saint-Cyr and colleagues introduced it in 1988 to study Parkinson's and Huntington's disease, and its trick is to replace size with colour: four identically sized disks, and the rule is that no darker disk may sit on a lighter one. The puzzle is mathematically the same and perceptually much harder, because the ordering is no longer something your eye does for you.
How it is used in practice
A clinician administering a tower task does not record "solved" or "failed". The measures that carry the information are:
- Moves taken against the known minimum — the efficiency of the route, not whether it finished.
- Time to first move, taken as an index of how much planning happened before acting. A person who moves instantly and a person who thinks for forty seconds are doing different things even if both finish.
- Rule violations — attempting to place a larger disk on a smaller one, or moving two disks at once. These are treated separately from inefficiency, because they indicate the rule was not held rather than that the route was poor.
- Where the attempt breaks down, which is the qualitative observation the numbers exist to support.
Standardised versions exist — the Tower Test in the Delis–Kaplan Executive Function System, and the computerised Stockings of Cambridge in the CANTAB battery — and they come with normative data for the specific version and administration. Those norms do not transfer between versions, and they certainly do not transfer to a browser.
The amnesia experiments
The puzzle's most cited appearance in neuropsychology is not about planning at all. It is about memory, and it is an argument that is still not settled.
In 1985, Cohen, Eichenbaum, Deacedo and Corkin reported that severely amnesic patients — including H.M., the most studied patient in the history of the field — improved at the Tower of Hanoi across sessions while having no memory of ever having seen it. That is a striking result, and it was taken as strong evidence for a separation between declarative memory, which is knowing that, and procedural memory, which is knowing how. You can learn to do this puzzle, the argument went, without being able to remember doing it.
Then the finding stopped replicating. Xu and Corkin returned to it in 2001 in a paper titled, with some ceremony, H.M. revisits the Tower of Hanoi Puzzle, and found that neither H.M. nor a second amnesic patient mastered the recursive strategy under conditions where healthy controls did. The same year, a case study titled Why the Tower of Hanoi Has Fallen Down described an amnesic patient who solved the puzzle by spontaneously using the odd–even rule — a rule the patient knew, not a skill learned by hand. The current reading is that the puzzle is not a clean measure of procedural learning: it has too much declarative content, because learning it well means learning a rule, and a rule is a thing you know rather than a thing your hands do.
That is a genuinely interesting failure. The task was chosen for this work precisely because it looked like a motor skill — you move disks with your hands — and what the research since has established is that most of it happens somewhere else.
Why it is a weak test on its own
Four problems, all well documented, and all of them apply to any version of the task including the one on this site.
Task impurity. Every executive-function task is also a task. To do this one you need vision, spatial reasoning, motor control, an understanding of the instructions and some willingness to keep going — so a low score has many possible causes and "poor planning" is only one of them. This is the standard objection to executive-function measures generally and the Tower of Hanoi is one of its stock examples.
Practice destroys it. The task measures planning only for somebody meeting it for the first time. Once a person knows the recursive method, the same puzzle measures recall of a method, and their score improves for reasons that have nothing to do with the ability being assessed. That rules out using it as a repeated measure, and it is why someone who has studied the method — here or anywhere else — is no longer a fair subject for the test.
Modest reliability. Test–retest correlations for tower tasks are often lower than a clinician would want from a measure used to make decisions, partly because of the practice problem and partly because a single run has so few scored events in it.
Several strategies, one score. Simon's 1975 point again: a perfect run can come from having planned it, from having memorised it, or from a perceptual pattern that needs no plan at all. The number of moves does not distinguish them, which is why the qualitative observations above are the part clinicians actually rely on.
None of that makes the puzzle useless. It makes it one measure among several, interpreted by somebody who watched the attempt — which is what a neuropsychological assessment is, and what a web page cannot be. What a web page can do is explain the puzzle properly, which is what the rest of this site is for: the recursive algorithm and why recursion works if you want the method, the history if you want the Victorian toy it started as, and the teaching materials if you are running it with a class.
Common questions
What does the Tower of Hanoi test measure?
In psychology it is used as a problem-solving task that draws on planning, working memory and inhibitory control — looking several moves ahead, holding sub-goals in mind, and being willing to make a move that appears to go backwards. It is often described as a test of executive function or of frontal lobe function, but that description is stronger than the evidence supports: performance also depends on visuospatial ability, processing speed, motivation and whether the person has met the puzzle before.
Is the Tower of Hanoi an IQ test?
No. It is not a measure of general intelligence and it is not scored like one. It appears in neuropsychological batteries as one task among many, where a clinician interprets it alongside an interview, a history and other tests. On its own a Tower of Hanoi score says very little about anybody, and no result from playing it online means anything diagnostically.
What is the difference between the Tower of Hanoi and the Tower of London?
The Tower of Hanoi uses disks of different sizes on three identical pegs, and the rule is about the disks: a larger disk may never sit on a smaller one. The Tower of London, introduced by Tim Shallice in 1982, uses balls of the same size on pegs that hold three, two and one — so the rule is about the pegs — and the task is to reach a pictured arrangement in a stated number of moves rather than to move a whole stack. Shallice designed it to offer a wider range of problems at graded difficulty than the Tower of Hanoi provided.
Which part of the brain does the Tower of Hanoi use?
Tower tasks are associated with the prefrontal cortex, and the Tower of London was created specifically to study patients with frontal lobe damage. Imaging studies also implicate parietal regions and the basal ganglia, which is what you would expect from a task that is spatial and sequential as well as strategic. No single brain area 'does' the puzzle, and a poor score does not localise anything.
Can practising the Tower of Hanoi improve executive function?
It will certainly make you better at the Tower of Hanoi. Whether that transfers to anything else is the open question in all cognitive-training research, and the general finding is that improvements stay close to the trained task. Practice effects are also why the puzzle is a weak repeated measure: once somebody has learned the recursive method, the task stops measuring planning and starts measuring recall of the method.
How many moves should a person take on the Tower of Hanoi?
The minimum is 2ⁿ − 1: 7 moves for three disks, 15 for four, 31 for five. Clinical versions typically use three or four disks and record the number of moves taken, the number of rule violations, and the time to the first move, rather than a single pass or fail. There is no universal cut-off, and normative data depend on the specific version and administration.