1
00:00:05,660 --> 00:00:06,957
[MUSIC]. 
Okay, so where are we now? 

2
00:00:06,957 --> 00:00:11,115
So we motivated the discussion of 
statistical inference and estimation by 

3
00:00:11,115 --> 00:00:15,075
bringing up this so called decline 
effect, where the, effect size of 

4
00:00:15,075 --> 00:00:19,431
scientific results seems to be going down 
over time and reproducibility is 

5
00:00:19,431 --> 00:00:24,130
suffering. 
And so we gave some reasons for this; 

6
00:00:24,130 --> 00:00:28,560
publication bias, you know mistakes and 
fraud and this multiple hypothesis 

7
00:00:28,560 --> 00:00:31,300
problem. 
And so we used these, motivations, these 

8
00:00:31,300 --> 00:00:33,260
scenario, to bring up various topics and 
techniques. 

9
00:00:33,260 --> 00:00:36,880
So we talked a little bit about basic 
statistical inference Where I just give 

10
00:00:36,880 --> 00:00:40,890
you an overview and, and that's it. 
we talked about effect size. 

11
00:00:40,890 --> 00:00:45,480
We brought up the specific term, 
heteroskedasticity. 

12
00:00:45,480 --> 00:00:47,530
For fraud detection, we brought up 
Benford's Law. 

13
00:00:47,530 --> 00:00:50,120
And then we multiple hypothesis testing 
which is. 

14
00:00:50,120 --> 00:00:53,680
Perhaps the most important part of the 
discussion we talked about the familywise 

15
00:00:53,680 --> 00:00:57,760
error rate and the false discovery rate 
and gave correction procedures for both 

16
00:00:57,760 --> 00:01:00,790
of these. 
Okay, and so this hopefully was a tour of 

17
00:01:00,790 --> 00:01:06,590
not just some basic concepts but also 
some, if not advanced at least things 

18
00:01:06,590 --> 00:01:11,930
that don't necessarily come up in a a you 
know, Stats 101 course. 

19
00:01:11,930 --> 00:01:14,250
But I think it's pretty important for us 
data scientists to understand. 

20
00:01:14,250 --> 00:01:21,560
In fact, as a data scientist, there's a 
view amongst statisticians that these 

21
00:01:21,560 --> 00:01:25,150
topics are not very well understood. 
And in fact, they'll point to typical 

22
00:01:25,150 --> 00:01:30,000
machine learning classes where understand 
the population, understanding the various 

23
00:01:30,000 --> 00:01:33,830
biases, understanding how to correct for, 
for the problems that can arise is not 

24
00:01:33,830 --> 00:01:38,105
taught at all and it's more of a. 
Of a, you know, blind application of 

25
00:01:38,105 --> 00:01:40,520
algorithms. 
So I think it's pretty important to go 

26
00:01:40,520 --> 00:01:47,670
over this choice of topics now. 
So, What about big data? 

27
00:01:47,670 --> 00:01:52,990
What changes? 
Well, so, Brad Efron. 

28
00:01:52,990 --> 00:01:57,665
Who's a world renowned statistician you 
know, describes it this way. 

29
00:01:57,665 --> 00:02:02,351
Says classical statistics was fashioned 
for small problems, a few hundred data 

30
00:02:02,351 --> 00:02:05,990
points at most, in just a few parameters. 
And the bottom line is that we've entered 

31
00:02:05,990 --> 00:02:08,960
in an era of massive scientific data 
collection, with a demand for answers to 

32
00:02:08,960 --> 00:02:13,090
large scale inference problems that lie 
beyond the scope of classical statistics. 

33
00:02:13,090 --> 00:02:19,910
And so, Suggest that something is 
changing in the area of big data. 

34
00:02:19,910 --> 00:02:22,318
Now, what can go wrong here? 
Well, as we've talked about, you can find 

35
00:02:22,318 --> 00:02:29,870
spurious relationships in big data and so 
this is a picture that I got from a 

36
00:02:29,870 --> 00:02:32,060
colleague recently that was emailed to 
him. 

37
00:02:32,060 --> 00:02:35,490
Which is a plot that someone took the 
time to make, may or may not have been as 

38
00:02:35,490 --> 00:02:39,618
a joke, but as you can see here, it says 
"Internet Explorer versus the murder 

39
00:02:39,618 --> 00:02:45,970
rate." Rate, OK. 
And so this is the murders in the US in 

40
00:02:45,970 --> 00:02:49,772
blue, along with the market share of 
internet explorer in the green. 

41
00:02:49,772 --> 00:02:53,855
And the, you know, corresponding 
discussion that went along with this 

42
00:02:53,855 --> 00:02:59,300
plot, you know, was, was somewhat 
amusing. 

43
00:02:59,300 --> 00:03:03,830
Talking about various theories for why 
The murder rate might be going up as in 

44
00:03:03,830 --> 00:03:06,860
the next four market share. 
Our murder rate goes down as, as in the 

45
00:03:06,860 --> 00:03:11,720
next four market share also goes down. 
But, the point here is that without some 

46
00:03:11,720 --> 00:03:14,978
common sense or without the [UNKNOWN] the 
application of understanding the scenario 

47
00:03:14,978 --> 00:03:21,523
of the problem you can make, you know 
discoveries Of, of this form. 

48
00:03:21,523 --> 00:03:24,927
Okay. 
Alright, and so other examples that have 

49
00:03:24,927 --> 00:03:30,263
been talked about in the literature, 
again brought up, as as, you know, bad 

50
00:03:30,263 --> 00:03:35,420
examples, the number of police officers 
and the number of crimes. 

51
00:03:35,420 --> 00:03:38,130
So, why might these 2 things be 
correlated? 

52
00:03:38,130 --> 00:03:40,780
You know, maybe police officers cause 
crimes. 

53
00:03:40,780 --> 00:03:44,560
Well, no, probably because there's in 
pop, in densely populated areas, there 

54
00:03:44,560 --> 00:03:47,220
are both more police officers and there 
are more, and there are more crimes. 

55
00:03:47,220 --> 00:03:51,335
By the way, just to point out again, you 
know, these, these authors here are not 

56
00:03:51,335 --> 00:03:56,818
authors that claim, made these claims. 
These are authors that brought up the 

57
00:03:56,818 --> 00:04:01,900
mistake. 
Okay, amount of ice cream sold and deaths 

58
00:04:01,900 --> 00:04:04,820
by drowning. 
Why would these things be correlated? 

59
00:04:04,820 --> 00:04:06,966
Well, there's a seasonality. 
Right? 

60
00:04:06,966 --> 00:04:11,184
In the summertime you sell more ice cream 
and more people go swimming. 

61
00:04:11,184 --> 00:04:16,770
And then one is you know Stork sightings 
and population increase used as you know, 

62
00:04:16,770 --> 00:04:21,860
evidence that storks do indeed bring 
newborns to families. 

63
00:04:21,860 --> 00:04:25,984
Well again, in more densely populated 
areas there's more people actually 

64
00:04:25,984 --> 00:04:28,080
actually see the Storks and so you get an 
increase in sightings. 

65
00:04:28,080 --> 00:04:33,960
So, these kind of procedures to remove 
bias and these procedures to understand 

66
00:04:33,960 --> 00:04:40,515
the population you are sampling from and 
understand the possibilities as far as 

67
00:04:40,515 --> 00:04:44,695
correlations. 
These things are taught in statistics 

68
00:04:44,695 --> 00:04:49,240
programs, but are not typically taught in 
machine learning classes. 

69
00:04:49,240 --> 00:04:51,760
Okay. 
So what does that have to do with big 

70
00:04:51,760 --> 00:04:52,370
data? 
Well. 

71
00:04:52,370 --> 00:04:56,520
Might be a view that there's, you know, 
the curse of big data, as Vincent 

72
00:04:56,520 --> 00:04:59,500
Granville put it, is the fact that when 
you search for patterns in a very, very 

73
00:04:59,500 --> 00:05:01,968
large data sets with billions or 
trillions of data points and thousands of 

74
00:05:01,968 --> 00:05:06,130
metrics, you are bound to identify 
coincidences that have Predictive power 

75
00:05:06,130 --> 00:05:11,940
and so the example he gives is to 
consider stock prices for some large 

76
00:05:11,940 --> 00:05:13,830
number of companies over a one month 
period. 

77
00:05:13,830 --> 00:05:16,055
And then you check for correlations 
between all pairs. 

78
00:05:16,055 --> 00:05:20,670
And actually doesn't stop there, because 
that would be over the same exact one 

79
00:05:20,670 --> 00:05:23,530
month period. 
But you might want to account for lags. 

80
00:05:23,530 --> 00:05:27,900
Maybe the stock price of Google. 
a few days later effects the stock price 

81
00:05:27,900 --> 00:05:31,985
of smaller companies that depend on you. 
So now you're not just comparing every 

82
00:05:31,985 --> 00:05:39,000
500 squared checking the paralyzed 
correlation of these time series but you 

83
00:05:39,000 --> 00:05:43,640
are also checking the paralyzed slightly 
offset one okay and so these are the 

84
00:05:43,640 --> 00:05:52,487
cross correlation procedures. 
So very basic time series analysis this 

85
00:05:52,487 --> 00:05:55,290
is just to measure the correlation and I 
just wanted to throw the formulas up here 

86
00:05:55,290 --> 00:05:59,820
where the covariants of two data sets is 
measured this way. 

87
00:05:59,820 --> 00:06:07,020
Alright so you take the data point xi and 
subtract the mean of x. 

88
00:06:07,020 --> 00:06:09,630
And multiply that by y i minus the mean 
of y. 

89
00:06:09,630 --> 00:06:12,540
And all that up and that's the 
covariance. 

90
00:06:12,540 --> 00:06:16,290
And then you divide the covariance by the 
standard deviation of each data set 

91
00:06:16,290 --> 00:06:19,741
multiplied by each other. 
And so this gives you the correlation. 

92
00:06:19,741 --> 00:06:22,080
Okay. 
So, what does this experiment look like? 

93
00:06:22,080 --> 00:06:27,820
Well, I generated this plot by running 
random walks for stock prices that start 

94
00:06:27,820 --> 00:06:30,714
at $10. 
They all start at the same, the same 

95
00:06:30,714 --> 00:06:38,799
Point, and at each step, which is an hour 
of simulated time. 

96
00:06:38,799 --> 00:06:44,810
A, draw a sample for a normal 
distribution where the mean is the 

97
00:06:44,810 --> 00:06:48,540
current stock price. 
And the, a standard deviation is one 

98
00:06:48,540 --> 00:06:51,890
percent of that current stock price. 
Okay. 

99
00:06:51,890 --> 00:06:57,180
And this is, not especially defensible, 
but you can see just sort of visually 

100
00:06:57,180 --> 00:06:59,620
that it does generate stock price looking 
things. 

101
00:06:59,620 --> 00:07:05,306
And you do get some variance here. 
Alright. 

102
00:07:05,306 --> 00:07:14,060
So, clearly this is, this is random. 
This plot shows the number of corelations 

103
00:07:14,060 --> 00:07:18,530
at a level of 0.9. 
All right, that's a pretty strong 

104
00:07:18,530 --> 00:07:24,160
correlation as a function of the number 
of stock prices tracked. 

105
00:07:24,160 --> 00:07:29,040
So as I went up from 10 to 100, I didn't 
go all the way up to 500 which is what 

106
00:07:31,210 --> 00:07:35,610
Vincent Granville described in the 
thought experiment. 

107
00:07:35,610 --> 00:07:40,510
This is the number of spurious 
correlations I- You, you find, okay and 

108
00:07:40,510 --> 00:07:44,110
this is also not doing the lagged cross 
correlation, alright this is just 

109
00:07:44,110 --> 00:07:47,810
directly [INAUDIBLE] the correlation of 
these two [INAUDIBLE] of time series 

110
00:07:47,810 --> 00:07:50,410
across this month. 
And that's a pretty long period to, 

111
00:07:50,410 --> 00:07:54,170
across a month. 
So what's the point? 

112
00:07:54,170 --> 00:07:57,678
Well [INAUDIBLE] gives more opportunities 
For spurious findings. 

113
00:07:57,678 --> 00:08:00,228
Okay. 
Now, it's not all bad news. 

114
00:08:00,228 --> 00:08:05,736
So, how is big data different? 
Well, there's a notion of big p versus 

115
00:08:05,736 --> 00:08:09,000
big n. 
Where big p is sort of the number of 

116
00:08:09,000 --> 00:08:12,290
columns. 
And big n is the number of rows. 

117
00:08:12,290 --> 00:08:15,300
And in this experiment we just did with 
the time series. 

118
00:08:15,300 --> 00:08:18,607
This was sort of a big piece in here. 
We looked at more you know an increasing 

119
00:08:18,607 --> 00:08:21,840
number of companies and then we looked at 
all possible correlations between them so 

120
00:08:21,840 --> 00:08:26,800
this was growing sort of quadratically. 
Okay. 

121
00:08:26,800 --> 00:08:30,280
So the thing about big data though is 
marginal cost of increasing the number of 

122
00:08:30,280 --> 00:08:33,310
records is essentially zero. 
It's gotten cheaper and cheaper and 

123
00:08:33,310 --> 00:08:36,412
cheaper to collect data. 
Okay. 

124
00:08:36,412 --> 00:08:40,700
Great. 
Now that's very very powerful, right. 

125
00:08:40,700 --> 00:08:45,140
We want to, the increase in the number of 
records, adds statistical power and helps 

126
00:08:45,140 --> 00:08:51,270
us sort of, you know, get lower and lower 
p values but it also amplifies bias. 

127
00:08:51,270 --> 00:08:54,229
If you 're collecting the wrong data, if 
you're looking at the wrong population. 

128
00:08:55,330 --> 00:08:59,660
you're going to make, you know, so-called 
discoveries that are simply false. 

129
00:08:59,660 --> 00:09:04,100
And so, for example, log all the clicks 
to your website, you have a very, very 

130
00:09:04,100 --> 00:09:07,440
large data set and you can very precisely 
model user behavior. 

131
00:09:07,440 --> 00:09:10,580
But that would only model your current 
users. 

132
00:09:10,580 --> 00:09:13,580
When your hope, you know, perhaps the 
whole point of modeling. 

133
00:09:13,580 --> 00:09:16,140
user behaviors to try to attract new 
customers. 

134
00:09:16,140 --> 00:09:20,270
Well, for example, if you have early 
adopters, and your current user base is 

135
00:09:20,270 --> 00:09:22,190
early adopters, you're only going to be 
modeling their behavior. 

136
00:09:22,190 --> 00:09:23,982
You haven't actually sampled the 
population at large. 

137
00:09:23,982 --> 00:09:30,320
You know, another example is mobile data. 
And this comes up in polling, for say, 

138
00:09:30,320 --> 00:09:34,430
the presidential election. 
you know, you, you're only sampling 

139
00:09:34,430 --> 00:09:37,330
people that have cell phones. 
And this may or may not be the same 

140
00:09:37,330 --> 00:09:38,780
population, you want, you want to be 
sampling. 

141
00:09:38,780 --> 00:09:47,770
Okay, this may ignore lower income groups 
or different age groups, okay. 

142
00:09:47,770 --> 00:09:51,150
You need to be careful on multiple 
hypothesis tests as well, as we pointed 

143
00:09:51,150 --> 00:09:54,640
out. 
So there's a fantastic comment from XKCD 

144
00:09:54,640 --> 00:09:59,350
that makes this point very, very clear. 
where they sort of demonstrate that green 

145
00:09:59,350 --> 00:10:01,984
jelly beans cause acne. 
Right, and the story here is that there's 

146
00:10:01,984 --> 00:10:09,662
20 different [SOUND] colours of jelly 
beans, and for a P value of 0.05 [SOUND] 

147
00:10:09,662 --> 00:10:13,830
we do 20 experiments. 
And sure enough we find one of the colors 

148
00:10:13,830 --> 00:10:17,300
indeed causes acne. 
But that would be expected purely by 

149
00:10:17,300 --> 00:10:19,080
chance. 
And so I encourage you to look up that 

150
00:10:19,080 --> 00:10:22,520
comment. 
And the other comment I'll make that we 

151
00:10:22,520 --> 00:10:27,260
will probably come back to is Nassim 
Taleb's Black Swan events. 

152
00:10:27,260 --> 00:10:32,570
So this is- Things that are sort of 
inherently unpredictable or the 

153
00:10:32,570 --> 00:10:38,202
distribution of them does not follow a 
normal distribution, sort of a bell curve 

154
00:10:38,202 --> 00:10:44,370
distribution, where the tails of the bell 
curve mean that extreme values become 

155
00:10:44,370 --> 00:10:47,260
exponentially more rare. 
That's the sort of definition of the 

156
00:10:47,260 --> 00:10:50,055
normal distribution. 
But in some cases, extreme values are not 

157
00:10:50,055 --> 00:10:55,750
exponentially less common. 
They, they, they happen, okay? 

158
00:10:55,750 --> 00:11:01,230
And so the example that he uses in this 
case is that, you know, that if the, if a 

159
00:11:01,230 --> 00:11:07,970
turkey was to model your behavior, it 
would get increasingly more confident 

160
00:11:07,970 --> 00:11:13,440
that that you mean it, it no hard. 
And you mean it, you know, good will. 

161
00:11:13,440 --> 00:11:16,100
Every day you come and feed the turkey, 
and everyday you take care of it and you 

162
00:11:16,100 --> 00:11:19,670
look out for its well being. 
But then on the, you know day before 

163
00:11:19,670 --> 00:11:26,100
Thanksgiving it gets slaughtered. 
Perhaps and so that was Taleb's argument 

164
00:11:26,100 --> 00:11:29,450
for a Black Swam event. 
A black swan itself refers to the fact 

165
00:11:29,450 --> 00:11:32,050
that people didn't believe black swans 
existed and then. 

166
00:11:32,050 --> 00:11:34,410
Finds out that they did, so it was an 
unexpected event, okay. 

167
00:11:34,410 --> 00:11:37,060
All right we'll talk more about that in 
some detail. 

168
00:11:37,060 --> 00:11:37,060
[SOUND] 

