The paper that says models can prove, but not guess
Tom Zahavy’s “LLMs can’t jump” argues that models have induction and are getting deduction, and lack abduction: the leap to a hypothesis nobody has written down yet. It is a personal position paper, not DeepMind’s view, whatever the coverage said.
Published in January and still being argued about, which is a reasonable definition of a paper worth reading. Tom Zahavy, a DeepMind researcher and a co-author on AlphaProof, sets out a distinction with three parts. Induction, finding the pattern in the data, models are extremely good at. Deduction, proving a result from established premises, they are getting good at fast. Abduction, inventing the premise in the first place, is the one he argues they have no mechanism for.
His case study is general relativity, chosen because it defeats the popular account of creativity as compression. There was very little observational data for Einstein to compress. The premise came first and the evidence came afterwards.
One thing to get right before repeating any of this. Most of the coverage rendered it as DeepMind pouring cold water on AI for science. Zahavy has said in public that it is a personal position paper, not the company’s view, and not even a claim that models can never make discoveries. The distance between what the paper argues and what the headlines said is itself a small lesson in reading sources.
Why it belongs under interpretability rather than under research. If a system can execute a proof but cannot originate the premise, then the premise came from somewhere, and in an enterprise deployment that somewhere is a person. Knowing which half of the work the model did is not a philosophical question when a decision has to be defended. It decides who signs.