Percentiles and Distributions
/via https://flowingdata.com/2017/05/02/summary-stat/ If you don’t pay attention to your’ data’s distribution, you run the risk of focusing on the wrong thing entirely. The perfect examples here are the datasets generated by Justin Matejka & George Fitzmaurice (•), where, thought wildly different, each dataset has the same summary statistics (mean, standard deviation, and Pearson’s correlation) to 2 decimal places! To really belabor this point, absent knowledge of the actual distribution, you shouldn’t rely on baseline summary statistics. Take averages for example: 1. Outliers can — wildly! — skew your averages. You walk to work 200 days a year, and fly cross-country to the HQ in Los Angeles once a year. What’s the average distance to work? 2. On the other hand, you averages can — totally — hide your outliers. Your DMV processes 290 out of 300 applications in less than 15 minutes, but the other ten take greater than 2 hours....