We are often more certain than the evidence really allows.
More quotesBuying ten copies of the same newspaper does not make the story ten times truer.
Put those copies into a dataset and somebody will admire the sample size.
Volume removes one kind of uncertainty: random wobble among observations. It does absolutely nothing for a method that repeatedly records the wrong thing, excludes the same people or duplicates the same event. At scale, the error becomes beautifully stable.
An active-user survey can become exquisitely precise about people who stayed while remaining blind to everyone who left. A tracking fault can count one action twice across millions of sessions. A model trained on years of narrow decisions can reproduce the old preference faster than any human ever managed. More rows make the result harder to dismiss because confidence has been confused with validity.
Small data deserves no halo. Anecdotes wobble, tiny samples exaggerate chance and hand-picked cases can lie magnificently. Large, well-collected evidence is enormously valuable. The argument concerns origin before quantity.
What process created each observation? Which people or events could never appear? Are repeated records independent? Can a sample be traced back to the thing it claims to measure? Those questions sound dull beside an impressive total. They are also where truth usually lives.
Confidence intervals describe uncertainty inside the arithmetic. They cannot warn that the underlying event was mislabelled or absent. Methodological uncertainty has to be stated in words; the interval does not contain it.
The newspaper seller is delighted. The buyer now owns ten identical mistakes and the powerful feeling of having done research.