The core problem of Big Data

Run one comparison at a 5 per cent false-positive threshold and you may be fooled once in twenty. Run a thousand comparisons and analysts can produce a small festival of convincing rubbish. That is the awkward side of big data. Precision increases. So does the number of opportunities to find patterns after the fact. Returning visitors using a particular browser appear to buy 38 per cent more—until someone notices there were thirteen of them and one placed a very large order.

The analyst presents three views. First, every correlation the search produced: hundreds. Second, the handful supported by a plausible reason. Third, results from untouched data collected later. Most of the excitement disappears between views two and three. Another team could present only the surviving pattern and look like prophets. Analysts give bullshit a clean history when they hide failed questions, display the lucky answer and use hindsight to call the search a hypothesis. Honest analysis shows the discarded results too.

That disappearance is useful. A result discovered and confirmed on the same records is rather like asking a witness to choose the suspect after showing them one photograph repeatedly. The process has leaked the answer into the test.

Large datasets are brilliant for rare events, small differences and messy reality. They also let mediocre questions acquire six decimal places. Before celebrating a pattern, count how many patterns were examined, inspect the size of the effect and try it on fresh observations.

The browser result fails. A less exciting pattern survives: customers who experience a specific error rarely return. Engineering already knew about the error. It had fewer decimal places and far worse consequences.

Behavioural principles

Behavioural ideas at play in this post

Short, plain-English explanations of the principles behind this post, with links to related books and examples in the archive.