PNAS Study Finds LLM Surprisal Misses Human Rereading

A PNAS study published August 7 compared eye movements from 368 adults with 409 language-model surprisal estimates. The estimates tracked early forward-reading difficulty but consistently missed the extra time and backward movements caused by syntactic reanalysis, indicating that next-word prediction alone is not a complete model of human sentence processing.
Researchers from New York University, the University of Massachusetts Amherst, and collaborating institutions published a large eye-tracking study in PNAS on August 7. The team analyzed 368 adult readers and compared their behavior with 409 language-model-based estimates of word surprisal, a measure derived from how unexpected a word is given the text before it.
The central result is a split between the first pass through a sentence and what happens when comprehension breaks down. Language-model surprisal approximated the direction and magnitude of difficulty during forward reading. It did not explain the additional processing cost visible when readers moved their eyes backward and reread earlier words after encountering syntactic ambiguity.
What the study measured
Participants read controlled sentences designed to create several kinds of syntactic difficulty, alongside matched controls and naturalistic filler text. The researchers separated eye movements into four measures, including forward reading time, whether a reader made a regression, the time spent before that regression, and the time spent rereading before moving ahead again.
That separation matters because a single total reading-time measure can blend word recognition with later repair. Across the tested estimators and model families, surprisal was much better at accounting for uninterrupted forward reading than for regression-based measures. In garden-path sentences, the models often captured the early slowdown but underpredicted the rereading cost by a large margin.
The authors interpret this as evidence for a multistage account of reading. Prediction helps explain routine word recognition and successful integration, but it does not fully capture the error detection and structural reanalysis that occur when an initial interpretation fails. Readers also tended to return to words useful for repairing the sentence structure, suggesting that rereading is targeted rather than random.
What the result means
The study does not show that language models are generally poor models of human language behavior. It identifies a narrower boundary: next-word probability can model an early component of reading while missing important later processes. The experiment also focuses on English sentence processing under controlled conditions, so it is not a general benchmark of comprehension, reasoning, or model quality.
For researchers building cognitive models or human-centered reading tools, the practical implication is that token probability alone is an incomplete proxy for difficulty. Systems intended to estimate confusion, accessibility, or rereading behavior may need explicit mechanisms for structural integration, error detection, and repair rather than relying only on standard autoregressive surprisal.
Key Points
- 1Across 368 analyzed readers, model surprisal fit forward-reading difficulty better than rereading measures tied to syntactic reanalysis.
- 2The study tested 409 surprisal estimators, making the late-stage mismatch unlikely to be a quirk of one language model.
- 3Next-word probability should not be treated as a complete proxy for human comprehension or reading difficulty.
Scoring Rationale
A large, peer-reviewed comparison across hundreds of model-based estimates establishes a useful boundary on next-word prediction as a cognitive model, with clear implications for language-model evaluation and human-centered systems.
Sources
Primary source and supporting public references used for this report.
Practice with real Logistics & Shipping data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Logistics & Shipping problems

