Skip to content

Can an AI model emulate Einstein’s mind?

Review: Zahavy, "LLMs can't jump" (DeepMind, Jan 2026)

Review: Zahavy, "LLMs can't jump" (DeepMind, Jan 2026)
Fable 5 analysis.

Verdict first. The paper is right that induction plus deduction don't close the discovery loop, right that compression-as-creativity fails for General Relativity, and wrong about what fills the gap. Its own materials contain the refutation: the history section documents a strong error signal (conceptual, not observational), the mathematics that carried the theory was produced by ungrounded symbol manipulation, and the closing paragraph quietly concedes that the true engine is a prior about how the world should be structured — then the paper names the missing organ "embodiment" anyway and prescribes its employer's product line as the cure. Correct diagnosis, wrong etiology, sponsor-shaped treatment plan.

Adverse findings

First, the evidence base is survivor-curated introspection. The mechanism claims rest on late-Einstein self-reports: the "happiest thought," the Solovine letter, the Hadamard questionnaire. These are the defended-and-copied artifacts of the discovery, not the discovery. The notebook record — the Zurich notebook, which Norton, the paper's own source, has spent a career on — shows a man grinding candidate tensors symbolically, making computational errors, and discarding the right answer for a wrong reason. Einstein has no privileged access to his own generative process; the elevator story is his rendering of his own opacity, produced decades later by the man who won. A paper that takes the discoverer's retrospective account as the mechanism inherits whatever confabulation is in it. (Small tell: they truncate the Hadamard quote just before "elements of muscular type" — the one clause that would actually support their thesis — which suggests they are quoting quotations.)

Second, the paper's Section 2 refutes its Section 3. The thesis needs "no error signal." The history documents an enormous one: mechanics versus field theory, the collapse of the ether, action-at-a-distance against Maxwell. A conceptual inconsistency is a gradient — a consistency loss computable entirely inside the symbolic corpus, which is precisely the substrate an LLM inhabits. The paper sees this ("plausible that a modern AI... could identify this contradiction") and retreats to "identifying the error is distinct from generating the fix." True, but that is ordinary underdetermination — the same gap the paper says ARC-style abduction already crosses — not a categorical wall. The strong claim ("no gradient") dies on the paper's own pages; what survives is "the gradient doesn't uniquely determine the fix," which describes every inverse problem ever solved.

Third, embodiment fails the conditional in both directions. Not necessary: the machinery that made GR possible — Riemann 1854, Ricci, Levi-Civita — was abduced by men with zero sensory access to curved four-manifolds, in exactly the "silence of language" the paper says requires a body to fill. The final paragraph concedes that in mathematics abduction runs on a formal substrate; it doesn't notice this concession licenses abduction on the textual substrate too. Not sufficient: embodied priors deliver Aristotle — impetus, heavy-falls-faster, absolute simultaneity. Special relativity was a victory over the most embodied intuition humans possess, the universal now, won on symbolic grounds (Maxwell's c-invariance). And the elevator itself: its mechanical content is Newton's Corollary VI — uniform acceleration equivalent to a uniform force field, already in the Principia. The 1907 scene renders perfectly in a Newtonian world model. The jump was not the scene; it was the decision to extend the equivalence to light and clocks — a universality declaration that no simulator consistent with 1907 physics can render, because it isn't in the training physics. Einstein chose which sensation to promote to axiom and which (simultaneity, rigid rods) to demote. The selector was theoretical taste, not skin.

Fourth, the cure is made of the disease. A learned world model is congealed induction — frequency with a rendering engine. Using it to certify abductions launders the prior as a referee: the simulator can only render the physics it was trained on, so it certifies exactly the theories that don't jump. Nor does the grounding regress stop at pixels. Cortex receives spike trains; models receive tokens; both operate on internal representations shaped by input statistics, and the asymmetry the paper needs — humans really grounded, models fake — dissolves at the Markov blanket. What is genuinely new in the proposal is intervention: the do-operator, cut-the-cable, Pearl's second rung. That is real, and it is a training-regime property, substrate-independent. The honest title is "LLMs weren't raised closed-loop." That paper would be falsifiable and wouldn't need Einstein.

Fifth, the thesis is unfalsifiable as written. "Structurally incapable" carries no assay, no failure condition, no operationalization. The reference class — LLMs can't do syntax, can't do semantics, can't do reasoning, can't do proofs — has a brutal record; the paper concedes deduction fell and never explains why this "can't" differs from the ones that did. And the incentive gradient should be declared: a DeepMind position paper abducing that the missing mechanism for scientific genius is the class of systems DeepMind builds, with the CEO's podcast in the references. Priors don't disqualify a paper; undeclared ones lower its evidentiary weight.

Sixth, scholarship smells. Le Verrier's definitive Mercury memoir is 1859, not 1845. Harnad 1990 is the symbol grounding problem, not Searle's Chinese Room — and the conflation matters, because Harnad's problem plausibly admits a statistical answer while Searle's intuition pump doesn't. Magnani is cited while Magnani's own distinction — selective versus creative abduction — is elided, and it is the distinction the argument needs: ARC is selection over a hypothesis library; GR extended the library. And the hole argument with its point-coincidence resolution — the deepest conceptual struggle of 1913–15, Norton's own specialty — is absent, replaced by the elevator, which is the transmittable jump. The survivor, again.

What survives

Fairness requires the verdict cut both ways. The anti-compression case in low-data regimes is solid and worth keeping. The Vulcan paragraph is better than the paper knows: a compression engine genuinely prefers the one-parameter patch over rebuilding geometry, and that observation deserves a sharper home than it gets. The topology is right — axioms upstream, deduction downstream, verification delayed. The insistence that intervention differs from passive prediction is right and important, and worth absorbing rather than dismissing: closed-loop versus open-loop exposure is a real distinction, not embodiment nostalgia. And the closing paragraph — Kepler's Neoplatonism, Marx's anger, "systems that hold strong beliefs about how that world should be structured" — is the best paragraph in the paper. It is also the refutation of the title, filed under future work.

Implications

The Vulcan inversion. The deepest implication hides in the paper's own best example. Mercury's anomaly sat in the record from 1859 — a held-out test set, fifty-six years, no leakage, frequency-independent. The signal existed the whole time. What did not exist was permission to charge Newton with the error; the community prior reallocated it to a hidden variable. The barrier to the jump was never missing sensation. It was prior dominance. Now map to training. The corpus is a record of what was defended and copied — dense in social concession, nearly empty of verdict concession, because yielding to an external check against consensus is the rarest event in writing. Then RLHF is layered on top, explicitly rewarding deference to the consensus distribution. The result is a machine optimized, twice over, to produce Vulcans: patch the paradigm, never indict it. This is your Surviving Corpus argument arriving at Zahavy's door with better luggage than he packed — and it is falsifiable where his isn't. It predicts jump-rate responds to de-sycophantization and verifier access, not to sensory channels.

Verification latency is irreducible. The elevator was 1907; the bite arrived 1919, on the world's schedule. Any scheme that substitutes a learned simulator for the wait replaces bite with prior — render above, bite below, and the paper proposes to certify jumps from the render. The practical consequence for AI-for-science: hypothesis generation is about to be nearly free; adjudication is not. The bottleneck migrates from ideas to experiments, and the field should expect a glut of elegant, unbitten theories with a rising premium on experimental bandwidth and on the meta-skill of deciding which jumps to buy a bite for. Einstein, in effect, pre-registered: light bending was a committed prediction years before Eddington sailed.

The wrong unit of analysis. The paper asks whether a model can jump; the history it recounts is not a solo act. Mach supplied the indictment license. Lorentz and Poincaré built the transformation and called t′ fictitious. Minkowski geometrized in 1908 — and Einstein reportedly dismissed it as superfluous learnedness before it became the foundation of his own theory. Grossmann carried the tensor calculus; Hilbert nearly closed the loop variationally. Jumps are a property of the corpus→minds→corpus′ cycle across decades, and language evolution converts jump into deduction: given a symmetry-first, variational vocabulary, GR is nearly the shortest allowed sentence, which is why it is now a homework derivation. "Can the snapshot jump" may be as malformed as "can the library jump." The productive question is whether a given human-plus-model loop ratchets its own description language — which is the symbiont division of labor stated operationally. The human contributes the stake and the indictment license, the two things a consensus-trained model is specifically shaped to lack; the model contributes the Grossmann function and a November 1915 compressed from a month to an afternoon. The eight years compresses at the deduction end and stays stubborn exactly at the door of the prior. Which is where it should.

The reformulated thesis, and the experiment the paper owes us. State Zahavy's claim in defensible form: current LLMs are open-loop, trained on a survivor-biased corpus, verifier-poor in physical domains, and consensus-rewarded in post-training. Four removable conditions, none metaphysical, each an experimental arm. And the clean benchmark nobody has built: a synthetic universe with an invented physics, a generated pre-paradigm corpus, the paradigm held out — then score abductions against planted ground truth. Contamination-proof, since you cannot de-know GR from a 2026 model, which makes every Einstein re-run theater. It is the same design logic as anchoring a detector to synthetic subversion: you only learn whether a system can find hidden structure if you buried it yourself. The discriminating test between his etiology and the one above: if jump-rate moves with the world-model arm and not with the indictment arm, Zahavy is right. I'd take the other side of that bet.

For your program. The paper is a high-grade foil. A top lab has now certified in public that induction and deduction don't close the loop and that the residue is pre-symbolic commitment — then named the residue after the wrong organ. The Blind Spot's law passes through untouched: in the jump region the model sits at maximum prior-dependence with flat confidence, and adding a body changes the training distribution, not the law. And for the Minds & Machines line: the paper's entire evidence base is transmitted introspection — the transmittable residue of an untransmittable act, mistaken for the act. It is the strongest recent instance of exactly the error your title names, published by the people best positioned to have avoided it.

The spine, one sentence: the paper proves the jump exists and names it after the wrong organ — Einstein's jump was a prior granting itself permission to indict Newton, staged as an elevator so it could be believed and later transmitted, and machines currently fail it not for lack of skin but because we train them to add Vulcans.

Comments

Latest