The move that looked wrong
In March 2016, I watched a program I had helped build place a stone on the fifth line of a Go board in Seoul. The move looked so implausible that some commentators suspected a software error. AlphaGo won the game, then the match against Lee Sedol, one of the greatest professional Go players, by 4:1.
Lee later said he had thought of AlphaGo as a machine that simply calculated probabilities. Move 37 changed his mind: it seemed to him that the system could be creative.
The move was unusual, but it was not a product of intuition alone. AlphaGo had two parts. Its move-selection network learned to predict what a strong human player would choose; it assigned the move only about a one-in-ten-thousand chance. A search mechanism then explored thousands of possible continuations, weighing the move’s consequences rather than just its plausibility in the current position.
The contrast with chess helps explain why that mattered. When Deep Blue beat Garry Kasparov in 1997, it searched six to eight moves ahead for each player and evaluated 200 million positions per second. Human-defined rules made that kind of calculation possible. Go’s positions depend on how groups of stones and territory develop over dozens of moves; calculating even a small share of possible outcomes would take a supercomputer billions of years. AlphaGo needed both a fast sense of promising moves and a way to examine what might happen next.
Fluency is not a reasoning system
Large language models work differently. They repeatedly predict the next token, producing fluent continuations of familiar patterns. Chain-of-thought methods ask them to lay out intermediate steps before answering, and those methods have improved performance, particularly in mathematics and programming.
But a longer sequence of predictions is not the same as a separate reasoning mechanism. Unlike AlphaGo’s search, a chain of thought does not give the model an explicit process for testing possibilities against a maintained state of knowledge. There are three gaps:
That distinction is consequential in medicine, engineering and research. When a diagnosis or treatment choice goes wrong, the answer alone is not enough. A practitioner needs to know whether the system reasoned badly, relied on weak evidence or began from a false assumption.
A state of knowledge, not just a stream of words
I recently left Google DeepMind because I believe machine reasoning needs a different architecture, built on an idea AlphaGo already demonstrated: keep a representation of the problem and update it as the system explores. AlphaGo’s game tree holds the positions and continuations it has considered, along with evaluations from its neural networks. It uses that accumulated information to choose a move.
For broader tasks, a system would need a comparable record of what it regards as established, what remains uncertain, what it has ruled out and which questions are still open. Reasoning would then mean changing that record through explicit steps: drawing implications, breaking a problem into parts, and choosing the next question, calculation or experiment.
The open world is harder than a board game. Its state is only partly known, possible actions are numerous and changing, and their consequences may be uncertain. Language models can still help: they can propose approaches based on known facts and available resources, use tools through APIs or code, and help check claims against evidence. But an independent component would need to assess whether each step actually reduces uncertainty. Beliefs should change only when evidence supports the change.
I think the key question is not how convincingly a model can explain an answer, but whether its intermediate judgments can be checked and revised. Scaling up the fast, associative part may make its intuitions more accurate; it does not turn intuition into deliberation. Move 37 mattered because AlphaGo held the position in view, assessed possible continuations and chose a move its own intuitive network was unlikely to favor. In science and medicine, the equivalent will require systems that can show how evidence changed what they believe—not merely tell a persuasive story after the fact.
Tore Grepel is a professor of machine learning at University College London. He was one of the key members of the AlphaGo team at DeepMind and works on making AI beneficial to people.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X