Big data gets bigger. So does the haystack

Search log:

4 audience types × 7 channels × 12 messages × 5 outcome windows.

That is 1,680 places for an exciting result to appear, before anyone starts checking age, device or whether the moon was in a persuasive mood.

One segment shows a 43 per cent lift. The chart is beautiful. There were nineteen people in it.

Big data increases the chance of learning something rare and important. It also gives wishful thinking industrial equipment. With enough slices, random variation eventually produces an impressive pattern. The analyst can then invent the mechanism afterwards and forget the 1,679 less convenient results.

A clean holdout is less glamorous. Freeze the proposed relationship and see whether it appears in observations untouched by the search. Record how many questions were asked. Show effect size and uncertainty beside the celebratory percentage. The impressive segment often evaporates; occasionally it survives and becomes genuinely interesting.

Starting with a plausible mechanism helps, though even a sensible story can be wrong. Exploration deserves a place. The dishonesty begins when exploration is described later as prediction and the lucky discovery receives a false childhood.

The nineteen-person segment goes into a second campaign. This time the lift is 2 per cent and comfortably compatible with nothing.

Someone asks whether the original chart should remain in the case study. The cursor hovers over “duplicate chart”.

Behavioural principles

Behavioural ideas at play in this post

Short, plain-English explanations of the principles behind this post, with links to related books and examples in the archive.