LDS history archive · ST-001Source-led research edition
The history of statistics and probabilityHow uncertainty became evidence.
Follow the games, population records, experiments, arguments, algorithms, and software that changed how people measure variation and reason from incomplete information.
Every record is anchored to original correspondence, a book, a paper, an archive, a project record, or an official statement. Publication dates, contested credit, later terminology, and limits on what a method can claim are marked.
This history is not a parade of formulas. It begins with fair division in unfinished games, then follows mortality records, distributions, samples, experiments, decisions, simulation, and causal questions. Each step changed what could be learned from data, and each came with assumptions that later users could forget.
Record 01 · Games of chance
Pacioli prints the problem of points
What changed
In Summa de arithmetica, Pacioli asked how players should divide a stake when their game ends early. His answer followed the points already scored rather than each player's remaining chance of winning.
Why it lasted
A practical dispute about fairness became a durable printed problem for later mathematicians to solve.
Record 02 · Mathematical probabilityWritten about 1564, printed 1663
Cardano counts the possible outcomes
What changed
Cardano's Liber de ludo aleae studied dice, cards, fair wagers, and ratios of favorable to possible cases. The manuscript is usually dated to about 1564 or later, but it remained unpublished until his collected works appeared in 1663.
Why it lasted
It is the earliest surviving sustained mathematical treatment of games of chance.
Record 03 · Mathematical probabilityJuly to October 1654
Pascal and Fermat divide an unfinished game
What changed
Across seven surviving letters, Pascal and Fermat compared two ways to divide a stake when play stops early. Pascal reasoned recursively and with combinations, while Fermat enumerated the possible continuations.
Why it lasted
A fair settlement now depended on future possibilities rather than points already scored.
De ratiociniis in ludo aleae set out rules for valuing uncertain prospects and worked through fourteen problems. It appeared in Leiden as an appendix to Frans van Schooten's Exercitationum Mathematicarum.
Why it lasted
Probability became a printed subject that readers could learn through propositions and exercises.
Graunt compared London's weekly bills of mortality across causes, sex, place, and time. He also estimated the city's population and constructed a rough survival schedule from records that did not directly report age at death.
Why it lasted
Administrative counts became evidence for demographic reasoning rather than simple bookkeeping.
Halley used Breslau birth and death records from 1687 to 1691 to estimate survival at successive ages. He then showed how those estimates could guide the pricing of life annuities.
Why it lasted
Observed mortality became a basis for probabilistic financial valuation.
In the posthumous Ars Conjectandi, Bernoulli proved that a sufficiently long sequence of independent trials is highly likely to produce a proportion close to the underlying probability. His nephew Nicolaus prepared the unfinished manuscript for publication.
Why it lasted
Theoretical chance gained a precise connection to the stability of repeated observations.
Record 08 · Normal approximationTextbook 1718, approximation 1733, expanded 1738
De Moivre finds a curve inside the binomial
What changed
The Doctrine of Chances gave English readers a systematic probability text in 1718. In a privately circulated 1733 pamphlet, de Moivre derived a bell-shaped approximation for the symmetric binomial and later included it in the 1738 edition.
Why it lasted
Large binomial calculations became manageable through approximation.
Record 09 · Inverse probabilityRead to the Royal Society in December 1763
Bayes and Price reason backward from observations
What changed
Bayes left a manuscript asking what repeated successes and failures could reveal about an unknown chance. Price prepared it for publication, added a substantial prefatory letter, amendments, an appendix, and numerical material, then communicated it to the Royal Society.
Why it lasted
The paper became a central early record of inference from observed effects to an uncertain cause.
Laplace's Memoire sur la probabilite des causes par les evenements developed a broad method for reasoning from observed events to their possible causes. He applied inverse probability beyond isolated gambling problems.
Why it lasted
Inverse probability became a sustained program for scientific inference.
Record 11 · Estimation and error theoryLegendre 1805, Gauss 1809
Least squares settles conflicting measurements
What changed
Legendre's 1805 comet-orbit book first published the method and the name methode des moindres carres. Gauss's 1809 Theoria Motus linked least squares to a model of observational error and claimed that he had used it since 1795.
Why it lasted
Minimizing squared residuals became a standard way to estimate quantities from inconsistent measurements.
Record 12 · Analytical probabilityCentral-limit work 1810, synthesis 1812
Laplace makes probability an analytical discipline
What changed
Laplace derived a broad approximation for sums of independent variables around 1810. His Theorie analytique des probabilites then assembled generating functions, inverse probability, asymptotic approximation, error theory, and population applications in 1812.
Why it lasted
Probability became a mathematical engine for measurement and inference across several sciences.
Record 13 · Rare-event distributionsMemoir read 1829, printed 1830, book 1837
Poisson gives rare counts their law
What changed
In a memoir on the proportion of female and male births, Poisson wrote the rare-count limiting expression now associated with his name. His 1837 Recherches extended Bernoulli-type trials and introduced the phrase law of large numbers.
Why it lasted
Rare counts and stable aggregate frequencies gained an applied mathematical language.
Quetelet applied averages and frequency patterns to measurements of bodies, crime, marriage, and other social records. He treated the average man as an analytical description of regularity at the population level.
Why it lasted
Statistical reasoning moved decisively into social science and public administration.
Record 15 · Epidemiology and spatial analysisInvestigation 1854, second edition 1855
Snow follows cholera back to the water
What changed
Snow combined household addresses, death records, local interviews, and comparisons between water companies to argue for waterborne transmission. The Broad Street map supported a larger body of evidence presented in the 1855 second edition of On the Mode of Communication of Cholera.
Why it lasted
The investigation became a landmark in observational epidemiology, spatial reasoning, and natural-experiment design.
Record 16 · Statistical graphics and public health
Nightingale makes preventable deaths visible
What changed
Nightingale's report used monthly polar-area diagrams to compare deaths from disease, wounds, and other causes during the Crimean War. Area and color made the scale of preventable mortality legible beyond specialist statistical circles.
Why it lasted
Statistical graphics became a forceful tool for public-health policy and institutional reform.
In Des valeurs moyennes, Chebyshev used a variance bound to prove a law of large numbers for broad classes of independent observations. The proof no longer depended on identical coin-toss-style trials.
Why it lasted
Probability gained a rigorous, distribution-independent way to control deviations from a mean.
Record 18 · Regression and correlationReversion 1877, regression 1886, correlation 1888
Galton traces regression and correlation
What changed
Galton's sweet-pea experiments described reversion in 1877, and his 1886 study of family stature named regression toward mediocrity. His 1888 Royal Society paper then measured co-relation through standardized paired observations.
Why it lasted
Statistics gained practical tools for describing dependence and the behavior of extreme observations.
Record 19 · Mathematical statisticsFrequency curves 1895, correlation treatment 1896
Pearson joins distributions, moments, and correlation
What changed
Pearson's 1895 paper on skew variation developed families of frequency curves and fitted them by moments. His 1896 work on regression, heredity, and panmixia gave a rigorous product-moment treatment of correlation and regression.
Why it lasted
Distribution fitting and association began to operate as a connected mathematical-statistical toolkit.
Das Gesetz der kleinen Zahlen compared Poisson probabilities with several empirical datasets. Its best-known table recorded fatal horse kicks across Prussian army corps from 1875 to 1894.
Why it lasted
The Poisson law became a practical model for rare counts rather than only a limiting formula.
Record 21 · Categorical data analysisPublished July 1, 1900
Observed counts acquire a general goodness-of-fit test
What changed
Pearson compared observed category counts with the counts predicted by a model, then derived a large-sample reference distribution for the resulting discrepancy statistic.
Why it lasted
Categorical observations could now be tested within a reusable inferential procedure. The paper helped turn goodness of fit from a visual judgment into a calculation with a reference distribution.
Markov extends a long-run law to dependent sequences
What changed
Markov showed that long-run regularity could survive a specific form of dependence between successive random quantities. The work became an early foundation for sequences now described as Markov chains.
Why it lasted
Dependence became something probability could model rather than a reason to abandon the mathematics. The idea later supported stochastic processes, queues, genetics, economics, and computation.
Record 23 · Small-sample inferencePublished March 1, 1908
Small samples receive their own theory
What changed
Under a normal population model, Gosset studied a standardized sample mean when the population variance is unknown and must be estimated from the same small sample. He derived the distribution needed to judge the result without relying on a large-sample approximation.
Why it lasted
Experiments with only a few observations gained a defensible inferential method. That mattered in brewing, agriculture, medicine, and any setting where additional measurements were expensive.
Record 24 · Likelihood and estimation theoryPublished April 19, 1922
Likelihood becomes a theory of statistical estimation
What changed
Fisher distinguished likelihood from inverse probability and proposed criteria for comparing estimators. He connected maximum likelihood with consistency and efficiency and developed the ideas of sufficiency and statistical information.
Why it lasted
Estimation gained a shared theoretical vocabulary. Later likelihood methods, standard errors, information matrices, and model-based inference all developed from this program.
Record 25 · Analysis of variancePublished July 1923
An analysis-of-variance table enters experimental science
What changed
Fisher and Mackenzie partitioned variation in a potato experiment into interpretable sources and displayed what historical scholarship identifies as probably the first published analysis-of-variance table.
Why it lasted
Experimental structure could be matched to a compact statistical decomposition. ANOVA soon became a common language for comparing treatments while measuring background variation.
Record 26 · Design-based causal inferenceOriginal Polish paper 1923; English translation 1990
Potential outcomes define a randomized experiment's contrast
What changed
Neyman described the yield each plot could produce under each treatment, then studied treatment averages and their uncertainty under random assignment in agricultural experiments.
Why it lasted
The paper supplied a design-based foundation for average causal effects and randomization variance. It made the unobserved alternative outcome part of the statistical problem.
Record 27 · Statistical process controlInternal memorandum May 16, 1924; public paper April 1930
A chart separates routine variation from process change
What changed
Shewhart's Bell Labs memorandum sketched the recognizable control chart: process measurements plotted against a center line and control limits. His later paper explained how the chart distinguished common variation from evidence of an assignable cause.
Why it lasted
Quality control became a continuing statistical process rather than a final inspection. The chart linked sampling, production decisions, and learning about a process over time.
Randomization becomes the foundation of experimental design
What changed
Fisher connected random allocation, replication, and local control to a defensible estimate of experimental error. The design determined which comparisons the subsequent analysis could support.
Why it lasted
Randomization gave experiments a built-in protection against systematic allocation bias and a basis for testing treatment effects. Study design became part of inference rather than a preliminary logistical step.
Record 29 · Hypothesis testingPublished February 16, 1933
Tests are designed around errors and power
What changed
Neyman and Pearson framed testing as a choice between specified hypotheses, with procedures compared by false-rejection probabilities and power. For simple hypotheses, the likelihood ratio identified a most powerful test at a fixed error rate.
Why it lasted
Alternative hypotheses and power became design criteria rather than afterthoughts. The framework shaped sample-size planning, quality control, clinical trials, and modern test construction.
Probability receives a measure-theoretic foundation
What changed
Kolmogorov represented events as sets and probability as a normalized, countably additive measure. Conditional probability, independence, expectations, and infinite stochastic systems could now be developed inside one framework.
Why it lasted
Finite calculations and continuous probability gained a common rigorous language. The framework remains the standard mathematical foundation for probability and much of theoretical statistics.
Neyman compared purposive selection with stratified random sampling and showed how a probability design supports measurable sampling error, confidence intervals, and efficient allocation of observations across strata.
Why it lasted
A survey's uncertainty could be derived from how its sample was selected. The paper helped move official statistics toward probability samples whose errors could be assessed rather than merely asserted to be representative.
Record 32 · Interval estimationRead March 28, 1935; published August 30, 1937
Confidence intervals receive a repeated-sampling construction
What changed
Neyman defined a procedure that produces intervals covering the fixed but unknown parameter at a stated long-run frequency under repeated sampling. The construction made coverage a property of the method.
Why it lasted
Interval estimation gained a general frequentist foundation. Confidence sets became a standard way to report both an estimate and the uncertainty produced by a sampling procedure.
De Finetti grounds subjective probability in coherence
What changed
De Finetti interpreted a person's probabilities through betting rates that must avoid a sure loss. Exchangeability then connected judgments about repeatable observations with mixtures of independent trials.
Why it lasted
Subjective probability gained a precise consistency criterion and a mathematical bridge from personal uncertainty to predictive distributions. The program later shaped Bayesian decision theory and modeling.
Bayesian inference becomes a broad scientific program
What changed
Jeffreys developed priors guided by invariance and information, posterior estimation, and odds-based comparisons between scientific hypotheses. The book joined philosophical foundations with worked scientific problems.
Why it lasted
Bayesian reasoning gained a durable objective formulation that scientists could apply across disciplines. Jeffreys priors and Bayes-factor ideas remain influential, even where their interpretation is debated.
Record 35 · Sequential analysisPublished June 1945
A test can decide when it has seen enough data
What changed
Wald's sequential probability ratio test evaluated evidence after each observation. Sampling continued while the likelihood ratio remained between two boundaries and stopped when the evidence supported one decision strongly enough.
Why it lasted
Sample size became a data-dependent decision rather than a number fixed before the first observation. Sequential methods reduced expected inspection costs and opened a new theory of monitoring evidence as it arrives.
One paper links information, improved estimators, and geometry
What changed
Rao bounded the accuracy attainable by unbiased estimators using statistical information, showed how conditioning on a sufficient statistic could improve an estimator, and interpreted information as a local metric on a family of distributions.
Why it lasted
The three results became foundations for efficiency theory, the Rao-Blackwell theorem, and information geometry. They connected what a sample contains with how accurately a parameter can be estimated.
Record 37 · Nonparametric inferencePublished December 1, 1945
Ranks provide practical tests without a normal model
What changed
Wilcoxon replaced raw measurement magnitudes with their ordered ranks. His short paper introduced procedures now known as the signed-rank test for paired observations and the rank-sum test for independent samples.
Why it lasted
Useful comparisons no longer depended entirely on normal-theory assumptions or a mean-and-variance description. Rank tests gave applied researchers methods that were simple to calculate and resistant to extreme values.
Record 38 · Information theoryPart I published July 31; Part II published October 1948
Uncertainty becomes an operational quantity
What changed
Shannon quantified the uncertainty in a probability distribution through entropy and the dependence between variables through mutual information. He then connected those quantities to the limits of reliable compression and communication.
Why it lasted
Probability distributions acquired operational measures of uncertainty and shared information. Those measures later became central to statistical learning, experimental design, model comparison, and information geometry.
Estimation and testing become decisions under loss
What changed
Wald treated an estimate or test result as an action whose consequences are measured by a loss function. Procedures could then be compared through their risk across parameter values, including Bayes, minimax, and admissibility criteria.
Why it lasted
Problems that looked different on the surface entered one mathematical framework. Decision theory later shaped shrinkage, classification, machine learning, medical choices, and policy analysis.
Record 40 · Monte Carlo simulationSubmitted March 6; published June 1, 1953
A Markov chain learns to sample a difficult distribution
What changed
The authors proposed trial changes to a physical system and accepted them according to an energy-based probability rule. Repeating the step produced a Markov chain whose long-run states followed the desired equilibrium distribution.
Why it lasted
Complex expectations could be approximated by simulated draws even when direct integration was impractical. The method became the starting point for modern Markov chain Monte Carlo.
Savage started from preferences between possible actions and derived conditions under which those choices can be represented by subjective probabilities and expected utilities. His sure-thing principle became a central consistency condition.
Why it lasted
Bayesian probability gained a behavioral foundation connected directly to decisions. The framework influenced statistics, economics, game theory, and later debates about rational choice.
Record 42 · Empirical BayesSymposium work presented in 1954-1955; published 1956
Many related problems teach each other a prior
What changed
Robbins considered a collection of parallel estimation problems and proposed learning their unknown mixing distribution from the ensemble. The estimated distribution could then support Bayes-style decisions for the individual cases.
Why it lasted
Information could be pooled across repeated problems without fixing a prior entirely in advance. The idea anticipated modern shrinkage, compound decision methods, and large-scale multiple testing.
A survival curve keeps censored observations in the analysis
What changed
Kaplan and Meier estimated a survival function from exact event times while retaining subjects whose observation ended before the event. The curve changes at observed failures and carries censored cases through the risk sets where they remain observed.
Why it lasted
Medical and reliability studies could use incomplete follow-up without treating every censored case as a failure or discarding it. The product-limit curve became a basic descriptive and inferential tool for time-to-event data.
Record 44 · Epidemiologic stratified analysisPublished April 1, 1959
Stratification makes retrospective comparisons more credible
What changed
Mantel and Haenszel combined evidence across strata of a retrospective study, providing a stratified test and a pooled measure of association. Analysts could condition comparisons on measured variables such as age or study center.
Why it lasted
Observational epidemiology gained a practical way to adjust comparisons for measured stratifying factors. The method became a foundation of case-control analysis and modern thinking about confounding.
Several noisy estimates improve through joint shrinkage
What changed
James and Stein showed that when at least three normal means are estimated together under total squared-error loss, shrinking the observed mean vector jointly toward a fixed target, conventionally the origin, can lower total risk everywhere.
Why it lasted
The result overturned the intuition that each unbiased sample mean must be best for its own coordinate. It became a foundation for shrinkage, hierarchical modeling, regularization, and empirical Bayes methods.
An estimator is designed for a model that is only approximately true
What changed
Huber modeled observed data as coming mostly from a reference distribution with a small, unspecified contaminating component. He derived location estimators whose loss is quadratic near the center and linear in the tails, balancing the mean's efficiency with the median's resistance.
Why it lasted
Robustness became a formal optimization problem rather than an informal preference for ignoring outliers. The paper established contamination neighborhoods and M-estimation as enduring tools for procedures that remain useful under modest model failure.
Cox separates relative risk from the baseline hazard
What changed
Cox modeled covariate effects on an event rate while leaving the baseline hazard unspecified and accounting for censored observations.
Why it lasted
The model gave medical and reliability studies a flexible regression method for time-to-event data without requiring a fully specified survival distribution.
Generalized linear models unite separate model families
What changed
Nelder and Wedderburn placed normal, binomial, Poisson, gamma, and related response models inside one framework built from an exponential-family distribution, a linear predictor, and a link function.
Why it lasted
Logistic regression, Poisson regression, analysis of variance, and several other methods could share a common theory and an iterative fitting strategy.
Rubin extends potential outcomes to randomized and observational studies
What changed
Rubin described a causal effect as a comparison between outcomes that the same unit could have under different treatments, then connected inference to the treatment-assignment mechanism.
Why it lasted
The framework made the missing counterfactual explicit and gave later causal methods a precise language for estimands, assignment, and identification assumptions.
Akaike turns model selection into an information problem
What changed
Akaike connected maximum likelihood with information loss and proposed a criterion that rewards fit while penalizing the number of estimated parameters.
Why it lasted
AIC reframed model choice around expected predictive information rather than a sequence of significance tests or fit alone.
Record 53 · Latent-variable estimationSeptember 1977
The EM algorithm exposes the data hiding inside the data
What changed
Dempster, Laird, and Rubin described a general cycle in which the E-step computes the conditional expectation of the complete-data log likelihood given the observed data and current parameter estimate. The M-step maximizes that expected objective.
Why it lasted
Mixtures, missing-data models, latent classes, and many other likelihood problems gained a reusable fitting strategy with a monotonic likelihood property.
Robins studied treatment histories in which a changing covariate can affect later treatment while also carrying the effects of earlier treatment.
Why it lasted
The g-computation framework addressed settings where ordinary adjustment can bias a causal estimate by conditioning on a treatment-affected confounder.
Gelfand and Smith demonstrated how Gibbs sampling, stochastic substitution, and sampling-importance-resampling could recover marginal posterior distributions in structured Bayesian models.
Why it lasted
Bayesian analyses that were blocked by high-dimensional integration became practical across hierarchical and latent-variable models.
Gelman and Rubin compared variation within several simulated chains with variation between chains started from deliberately overdispersed positions.
Why it lasted
The potential scale-reduction factor gave practitioners a practical signal that an iterative simulation might not yet have explored the target distribution adequately.
Causal diagrams make identification assumptions visible
What changed
Pearl used directed acyclic graphs to encode causal assumptions and graphical criteria to decide when an intervention effect could be identified from observational data.
Why it lasted
Causal assumptions could be inspected, challenged, and connected to an estimand before a model was fitted. The framework gave confounding and adjustment a visual and mathematical language.
Ihaka and Gentleman described a free statistical environment that presented an S-like interface while drawing on Scheme for implementation and lexical scoping.
Why it lasted
R lowered the cost of statistical computing and gave researchers a shared language in which new methods, graphics, and reproducible tools could spread quickly.
NUTS teaches Hamiltonian Monte Carlo when to turn back
What changed
Hoffman and Gelman built an adaptive Hamiltonian Monte Carlo method that stops a simulated trajectory before it doubles back and automatically tunes its step size.
Why it lasted
Efficient gradient-based posterior sampling became practical without requiring users to hand-tune the trajectory length for every model.
Record 62 · Reproducibility and research policy26 June 2015
The TOP Guidelines turn openness into policy
What changed
The Transparency and Openness Promotion Guidelines gave journals graded standards for citation, data, code, materials, study design, preregistration, and replication.
Why it lasted
Reproducibility practices moved from individual preference toward policies that journals and funders could adopt, disclose, and enforce.
Record 63 · Statistical interpretation7 March 2016
The ASA draws a boundary around the p-value
What changed
The ASA stated that a p-value does not measure the probability that a hypothesis is true, the size of an effect, or the importance of a result.
Why it lasted
A professional statistical body addressed threshold-driven practice directly and asked researchers to interpret p-values with design, evidence, effect size, and reporting context.
Record 64 · Probabilistic programming11 January 2017
Stan makes a model into an executable probability program
What changed
Stan compiled a probabilistic model into gradient calculations and used Hamiltonian Monte Carlo, including NUTS, to explore its posterior distribution.
Why it lasted
Researchers could specify complex Bayesian models without writing a new sampler for each one, while retaining access to diagnostics and model-generated quantities.
Record 65 · Causal and semiparametric inference16 January 2018
Double machine learning protects inference from flexible nuisance models
What changed
The authors combined Neyman-orthogonal scores with cross-fitting so flexible learners could estimate high-dimensional nuisance functions without overwhelming inference for a lower-dimensional target.
Why it lasted
Prediction tools such as forests, regularized regressions, boosting, and neural networks could enter causal or structural estimation while preserving valid large-sample confidence statements under stated conditions.
Rank normalization repairs a trusted MCMC diagnostic
What changed
Vehtari and colleagues replaced the traditional variance comparison with rank-normalized, split, and folded checks, then paired them with bulk and tail effective-sample-size measures.
Why it lasted
The revised workflow can expose heavy tails, changing scales, poor tail exploration, and other failures that the original R-hat may overlook.
A record is included when it materially changed the history of statistics and probability and the claim can be traced to a reliable source. The timeline separates the original contribution from interpretations that appeared later.
Primary evidence firstOriginal papers, books, standards, archives, and official technical records anchor each milestone.
People named by roleContributors are linked to public profiles and described by the work they performed, not by a vague credit line.
Disputes stay visibleCompeting claims, retrospective labels, and uncertain dates are identified instead of being flattened into one story.
Corrections are welcomeReaders can inspect every cited source and report a factual issue through the public corrections process.
111 links to original correspondence, books, papers, institutional archives, project records, official statements, and historical studies. Dates distinguish a manuscript, meeting, publication, later edition, public release, and retrospective. Named theorems are kept separate from the longer chains of work that produced them.