Don’t be fooled—LLMs don’t reason

AI-rewritten: This is a summary of an article from MIT Technology Review, rewritten by AI (Qwen, running locally) to make it easier to read. The facts come from the original article – read it for the full story.

MIT Technology Review • Thore Graepel • October 2, 2026

In March 2016, AlphaGo played move 37 in a match against Lee Sedol, a professional Go player who initially believed the computer relied solely on probability. Lee later admitted that seeing this move convinced him AlphaGo was creative. Unlike Deep Blue’s chess victory in 1997, which used hard-coded rules and brute-force calculation, AlphaGo faced a vastly more complex game where computing all outcomes would take billions of years. The author argues that move 37 resulted from genuine reasoning powers that today’s AI lacks, rather than pure machine intuition.

AlphaGo combined two systems: a policy network that guessed human-like moves and a search machinery that weighed future consequences by constructing a game tree with thousands of branches. This mirrors Daniel Kahneman’s distinction between fast, gut-level System 1 thought and slow, deliberative System 2 thinking. In contrast, large language models like ChatGPT operate primarily as System 1, predicting the next token based on patterns without a separate reasoning mechanism. While techniques like "chain of thought" allow models to generate intermediate steps, these are still produced by the same prediction process rather than a distinct logical structure.

The author identifies three shortcomings in current chatbots: they lack an explicit record of their hypotheses and evidence, knowledge is interwoven with reasoning in neural weights rather than represented separately, and they often fabricate reasoning paths after reaching an answer. In high-stakes fields like medicine, it is crucial to understand how a system arrives at a conclusion to identify errors. The author recently left Google DeepMind to advocate for a new approach inspired by AlphaGo’s architecture.

This proposed system would maintain an "epistemic state," a clear record of what is known, doubted, or ruled out, allowing reasoning to systematically update this state. Unlike board games with fixed rules, real-world reasoning must handle partial knowledge and uncertain outcomes. The author suggests that future systems should use independent parts to evaluate moves based on whether they resolve uncertainty, updating beliefs only when backed by evidence. This approach aims to create trustworthy intelligence capable of producing novel insights in areas like drug discovery and climate science through auditable sequences of logic.

Source: MIT Technology Review • Thore Graepel • October 2, 2026

Read the original article at MIT Technology Review →

Leave Comment

Your email address will not be published. Required fields are marked *