Skip to main content

We’re continuing our look at the “Deadly Dozen” data science mistakes – the costly errors that can undermine otherwise promising analytics projects. In Mistake 4, we looked at what happens when teams aim at the wrong target. But even when you’re asking the right question, the data itself can lead you astray if you treat it as the whole story. That brings us to Mistake 5: Listen Only to the Data.

Inducing models from data has the great virtue of looking at the data afresh unconstrained by old ideas. But, while you “let the data speak,” don’t ignore existing wisdom about the problem domain. Almost always, a good solution comes from close teaming between domain experts and analytics experts.

Often, nothing inside the data can protect us from making seemingly significant, but wrong, conclusions. We must look outside of the data and especially ask ourselves how the data that came to us might differ from the full picture.

When the Data Tells the Wrong Story

Let’s look at a real example in detail. The table below contains two variables about US high schools, by state: average cost and the rank of average SAT score. Let’s say our task is to model their relationship to advise a state of the costs of improving their educational standing, especially relative to nearby states. Figure 1 illustrates the relationship between the two variables and includes a linear regression line. The positive correlation is very significant statistically.[1] Astonishingly though, the relationship is the opposite of what is expected! That is, to improve our standing (lower our SAT ranking), the graph suggests we need to reduce school funding.

Spending and Rank of Average SAT Score by State in 1994

USA State SAT Rank $ Spent
AK 31 7877
AL 14 3648
AR 17 3334
AZ 25 4231
CA 34 4826
CO 23 4809
CT 35 7914
DC 49 8210
DE 37 6016
FL 40 5154
GA 50 4860
HI 44 5008
IA 1 4839
ID 22 3200
IL 10 5062
IN 47 5051
KS 6 5009
KY 18 4390
LA 16 4012
MA 33 6351
MD 32 6184
ME 41 5894
MI 20 5257
MN 3 5260
MO 13 4415
MS 12 3322
MT 19 5184
NB 8 4381
NC 48 4802
ND 2 3685
NH 28 5504
NJ 39 9159
NM 15 4446
NV 29 4564
NY 42 8500
OH 24 5639
OK 11 3742
OR 26 5291
PA 45 6534
RI 43 6989
SC 51 4327
SD 5 3730
TN 9 3707
TX 46 4238
UT 4 2993
VA 38 5360
VT 36 5740
WA 30 5045
WI 7 5946
WV 27 5046
WY 21 5255

Figure 1: Rank of a State (in average SAT score) vs. its spending per student (circa 1994) + the least-squares regression estimate line

 

 

When I show groups this problem participants will often suggest gathering further state data – for example, living costs or percent rural or urban population – to potentially help explain what is happening. But additional demographic variables, even if they have some small effect, miss the real problem. The breakthrough came for me by focusing on the best-ranked states, which all had something (outside the data) in common: their state universities (which are much more affordable than private ones) all require the competing ACT test. Students from ACT states who also take the SAT test are largely aiming at more prestigious institutions. Those samples are self-selected, causing their scores to be higher.

A state’s average SAT score (or the rank thereof) is not a good proxy for educational quality because the proportion of students taking the SAT varies widely by state. The states with a high proportion of students taking the SAT are highlighted in Figure 2, and their distribution is almost a perfect inverse of the states with high SAT reading scores, shown in Figure 3.[2] On reflection, this inverse relationship is obvious, and it reveals that the average SAT metric is useless. This further means that the strong negative correlation between spending and SAT scores is completely spurious. Remember, nothing in the original data could reveal this.[3]

Figure 2: Proportion of eligible students taking the SAT test by state

 

Figure 3: Average SAT reading scores by state

 

What the Data Leaves Out

The above example analysis application employed “opportunistic”, or found, data. Most data science problems rely on such data and there is always a risk that such data are in some way missing key examples (censored), self-selected, biased, or distributed differently from the problem’s true underlying data. But even data generated by a designed experiment needs external information. For example, a national defense project from the early days of Neural Networks attempted to distinguish aerial images of forests with and without tanks in them. Perfect performance was achieved on the training set, and then also on an out-of-sample set of data that had been gathered at the same time but not used for training. This was celebrated at first but, wisely, an additional confirming study was performed. New images were collected and the model performed very poorly on them. This drove investigation into the features driving the models and revealed them to be magnitude readings from specific locations of the images; i.e., background pixels. It turns out that the day the tanks had been photographed was sunny, and the day non-tank images were collected was cloudy![4] Even resampling the original data wouldn’t have protected against this error, as the flaw in the data was inherent in the generating experiment.

Another tank example from Dean Abbott[5] illustrates the same danger even more vividly. Dean had worked at a San Diego defense contractor, where they sought to use radar to distinguish tanks vs. trucks at different angles. Because radars and mechanized vehicles are bulky and expensive to move around, they fixed the radar installation and rotated a tank or a truck on separate large, rectangular platforms. Signals were beamed at different angles, and the returns were extensively processed – using polynomial network models of subsets of principal components of Fourier transforms of the signals – and great accuracy in classification was achieved.

However, in seeking greater transparency, Dean discovered, much to his chagrin, that the source of the key distinguishing features determining vehicle type turned out to be the bushes beside one platform![6] Further, it is suspected that the angle estimation accuracy came from the signal reflecting from the platform corners – not a feature one will encounter in the field. Again, no modeling technology alone could overcome flaws in the data; it took careful study of how the model worked to discover its weakness.

Looking Beyond the Data

The lesson is not to distrust the data, but to recognize what the data cannot tell you on its own. Whether you’re working with found data or a carefully designed experiment, ask how the data was generated, what might be missing or biased, and whether the patterns your model finds make sense in the real world. The strongest solutions come from combining what the data reveals with what domain experts know about the problem – and making sure the model is learning what you think it is.

[1] The t-statistic is over 4, suggesting that such a strong relationship occurs randomly only 1/10,000 times. This theoretical result is confirmed by Target Shuffling; it takes about 104 randomized trials (shuffle the rankings and re-fit) before a correlation this strong is measured. Target Shuffling is the advanced resampling technique, to be detailed in the 13th blog in this series, that can protect analysts from spurious correlations due to multiple comparisons. But even it does not solve the problem here.

[2] Figures 2 and 3 are from [economix.blogs.nytimes.com](http://economix.blogs.nytimes.com/)

[3] The remaining mystery is why ACT states spend so much less per student than SAT states. Meanwhile, if one instead employs NAEP (National Assessment of Educational Progress) scores — an educational metric with near-universal test participation — one finds no relationship between spending and outcome.

[4] PBS featured this project in a 1991 documentary series “The Machine that Changed the World”: Episode IV, “The Thinking Machine”.

[5] A close friend and former colleague who wrote a very useful book on data science.

[6] This excellent practice of trying to break one’s own work, is so hard to do even if one is convinced of its need, that managers should pit teams with opposite reward metrics against one another to proof-test solutions.