1
00:00:00,005 --> 00:00:06,272
[MUSIC]. 

2
00:00:06,272 --> 00:00:09,611
All right, so we've started a new week, 
starting a new section of the course, 

3
00:00:09,611 --> 00:00:12,725
where are we? 
We talked about what I maybe in this 

4
00:00:12,725 --> 00:00:16,483
slide [UNKNOWN] calling informatics. 
So the management, manipulation, 

5
00:00:16,483 --> 00:00:20,156
integration of data. 
And we have some emphasis on scale and we 

6
00:00:20,156 --> 00:00:25,159
have some emphasis on specific tools. 
Okay, so now we're moving into what I'll 

7
00:00:25,159 --> 00:00:27,496
call analytics. 
And so we're going to talk about 

8
00:00:27,496 --> 00:00:31,046
statistical estimation and prediction. 
And one of the points I want to make is 

9
00:00:31,046 --> 00:00:34,448
that this builds on the informatics, in 
that the things you're going to learn, 

10
00:00:34,448 --> 00:00:39,246
you can implement using the tools before. 
And we already saw a piece of this, where 

11
00:00:39,246 --> 00:00:41,294
we did sort of mult-, you know, matrix 
multiplication in various tools and so 

12
00:00:41,294 --> 00:00:43,786
on. 
Okay? 

13
00:00:43,786 --> 00:00:47,416
Okay. 
The other point, if you remember, we made 

14
00:00:47,416 --> 00:00:51,369
early on was that, you know, 80% of what 
people think of as analytics really boils 

15
00:00:51,369 --> 00:00:56,130
down to the ability to do sums and 
averages. 

16
00:00:56,130 --> 00:00:58,050
And so we'll see a little bit of that in 
here. 

17
00:00:58,050 --> 00:01:01,845
We're maybe, you know, understanding the 
problem and understanding the solution 

18
00:01:01,845 --> 00:01:06,262
is, you know, hard or easy, depending on 
your, maybe your background. 

19
00:01:06,262 --> 00:01:09,788
But as far as implementing it, it's, it's 
not too bad. 

20
00:01:09,788 --> 00:01:12,426
Okay? 
And then we'll move on to the visual, 

21
00:01:12,426 --> 00:01:16,716
visualization in a couple of weeks. 
[MUSIC]. 

22
00:01:16,716 --> 00:01:19,489
All right. 
So to get started on this [SOUND] I 

23
00:01:19,489 --> 00:01:23,953
want to call your attention to this 
article in 2010 from The New Yorker, with 

24
00:01:23,953 --> 00:01:29,290
the caveat that this is far from a 
research article. 

25
00:01:29,290 --> 00:01:32,938
and in fact, a lot of what the article 
has to say, I'm not sure I'd recommend 

26
00:01:32,938 --> 00:01:36,876
taking to, to heart. 
But the point the, you know, topic that 

27
00:01:36,876 --> 00:01:41,300
they bring up is that, you know, the 
title is here, is the truth wears off. 

28
00:01:41,300 --> 00:01:45,360
And what they're exploring is this notion 
that statistical results in the sciences 

29
00:01:45,360 --> 00:01:49,882
seem to have gotten weaker over time. 
And so, John Davis', you know, a 

30
00:01:49,882 --> 00:01:54,222
researcher at University of Illinois who 
does work on anti-depressants, is quoted 

31
00:01:54,222 --> 00:01:57,818
in this and, is discussed as talking 
about how a forthcoming analysis 

32
00:01:57,818 --> 00:02:01,724
demonstrating the efficacy of any 
anti-depressant has gone down as much as 

33
00:02:01,724 --> 00:02:08,600
three fold in recent decades. 
So this is that the effectiveness as 

34
00:02:08,600 --> 00:02:14,600
measured by clinical trials of these 
antidepressants has gotten a lot weaker. 

35
00:02:14,600 --> 00:02:20,304
The article also talked about Anders 
Moller who studied barn swallows and 

36
00:02:20,304 --> 00:02:26,192
discovered that the females were more 
likely to mate with males that had long 

37
00:02:26,192 --> 00:02:33,146
symmetrical feathers. 
And these, these findings sort of relied 

38
00:02:33,146 --> 00:02:39,368
on precise measurements of the symmetry. 
And so, this was a pretty significant 

39
00:02:39,368 --> 00:02:45,000
discovery, but over the course of the 
next five or six years, the effect size, 

40
00:02:45,000 --> 00:02:52,480
as discovered by himself and other 
researchers shrank by 80%. 

41
00:02:52,480 --> 00:02:54,666
Okay. 
A lot of the article talks about Jonathan 

42
00:02:54,666 --> 00:02:57,970
Schooler in 1990 who made a discovery of 
an effect that he called verbal 

43
00:02:57,970 --> 00:03:02,754
overshadowing. 
which was counterintuitive because it 

44
00:03:02,754 --> 00:03:08,084
showed that people that are asked to 
describe a face, you know, using English, 

45
00:03:08,084 --> 00:03:13,414
that they've, that they've seen were 
actually less likely to remember it than 

46
00:03:13,414 --> 00:03:20,575
those who had just seen the face. 
And so, the, you know, talking about the 

47
00:03:20,575 --> 00:03:24,475
face somehow overshadowed the effect of 
just seeing it alone. 

48
00:03:24,475 --> 00:03:28,060
Okay. 
And once again, this effect seemed to get 

49
00:03:28,060 --> 00:03:30,805
weaker over time and it became 
increasingly difficult to measure but you 

50
00:03:30,805 --> 00:03:35,054
including by Jonathan Schooler himself. 
And, in fact, he's quoted as saying this, 

51
00:03:35,054 --> 00:03:36,560
this frustrated him. 
Right? 

52
00:03:36,560 --> 00:03:37,732
He was having trouble replicating it. 
Okay. 

53
00:03:37,732 --> 00:03:42,774
And then they also bring up someone who's 
a little less respected. 

54
00:03:42,774 --> 00:03:47,394
Well a little off, fair amount unless 
respected in the scientific community as 

55
00:03:47,394 --> 00:03:51,396
an historical example. 
And its the person who actually coined 

56
00:03:51,396 --> 00:03:54,612
the term the decline effect, which is 
brought up over and over again in the 

57
00:03:54,612 --> 00:03:58,834
article. 
So, in the 1930s, Joseph Rhine tested 

58
00:03:58,834 --> 00:04:04,378
individuals with these card guessing 
experiments in an attempt to measure the 

59
00:04:04,378 --> 00:04:10,662
effect of extrasensory perception or ESP. 
And he's the one who actually coined that 

60
00:04:10,662 --> 00:04:12,996
term. 
So, he had a few students that achieved 

61
00:04:12,996 --> 00:04:17,008
multiple streaks of very low probability, 
you know, many, many cards in a row they 

62
00:04:17,008 --> 00:04:22,444
guessed right and so on. 
But there was a decline effect, in the 

63
00:04:22,444 --> 00:04:28,576
same candidates, the same participants, 
couldn't match the earlier performance, 

64
00:04:28,576 --> 00:04:32,726
okay. 
And so, the article touches on what is 

65
00:04:32,726 --> 00:04:37,150
essentially the correct explanation 
[LAUGH] to this effect. 

66
00:04:37,150 --> 00:04:41,294
And it also sort of brings up the 
possibility of a bunch of quasi, you 

67
00:04:41,294 --> 00:04:47,700
know, mystical, incorrect explanations of 
this, at least in my opinion. 

68
00:04:47,700 --> 00:04:50,580
So I want to tell you about, I wan-, I 
want to, you know, in the next couple of 

69
00:04:50,580 --> 00:04:53,460
segments, in the next few segments, I 
want to sort of explore this as a test 

70
00:04:53,460 --> 00:04:57,933
case for statistics, statistics and 
statistical estimation. 

71
00:04:57,933 --> 00:05:01,886
And I want to use it as a vehicle to 
introduce the fundamental concepts of 

72
00:05:01,886 --> 00:05:06,442
statistics and also some of the, somewhat 
more advanced concepts of statistics, 

73
00:05:06,442 --> 00:05:11,190
especially as, as they relate to big 
data. 

74
00:05:11,190 --> 00:05:13,540
Okay. 
So to get started, let's talk about the 

75
00:05:13,540 --> 00:05:17,740
background here, and then we'll come back 
to this specific article and, and explore 

76
00:05:17,740 --> 00:05:22,195
the reasons for why this, why the truth 
wears off. 

77
00:05:22,195 --> 00:05:24,675
Okay. 
So this is going to be, this is not 

78
00:05:24,675 --> 00:05:27,830
going to be a replacement for a 
introductory college statistics course. 

79
00:05:27,830 --> 00:05:30,530
This is going to be a quick overview of 
the terminology and the concepts you 

80
00:05:30,530 --> 00:05:32,456
should be familiar with. 
Okay. 

81
00:05:32,456 --> 00:05:34,260
So we're talking about statistical 
inference, here. 

82
00:05:34,260 --> 00:05:37,484
And so these are methods for drawing 
conclusions about a popula-, general 

83
00:05:37,484 --> 00:05:42,059
population from sample data. 
And there's two key methods that you can 

84
00:05:42,059 --> 00:05:46,780
use here: hypothesis tests and confidence 
intervals. 

85
00:05:46,780 --> 00:05:48,732
And we're going to bring up confidence 
intervals again later on, but I'm not 

86
00:05:48,732 --> 00:05:50,460
going to talk about them directly right 
now. 

87
00:05:50,460 --> 00:05:52,754
All right. 
So what is hypotheses testing? 

88
00:05:52,754 --> 00:05:57,885
Well, you're going to be comparing an 
experimental group, to a control group. 

89
00:05:57,885 --> 00:06:01,620
And, there's always going to be a null 
hypotheses. 

90
00:06:01,620 --> 00:06:05,480
And a null hypotheses is there's just no 
difference between these two groups. 

91
00:06:05,480 --> 00:06:07,516
Right? 
The one who received the treatment in 

92
00:06:07,516 --> 00:06:11,110
question, are no different than the ones 
who did not. 

93
00:06:11,110 --> 00:06:15,598
you know, the, the new website generates 
no more traffic than the old website, 

94
00:06:15,598 --> 00:06:20,190
than the control website, than the 
default, and so on. 

95
00:06:20,190 --> 00:06:24,120
Okay, so that's the null hypothesis. 
The alternative hypothesis is that there 

96
00:06:24,120 --> 00:06:25,968
is an effect. 
That there's a statistically significant 

97
00:06:25,968 --> 00:06:29,130
difference between the two. 
And so here difference is defined in 

98
00:06:29,130 --> 00:06:33,098
terms of some test statistic. 
And you can, most of these examples 

99
00:06:33,098 --> 00:06:36,348
you'll find in an introductory course, or 
really any course, there going to be 

100
00:06:36,348 --> 00:06:39,124
about comparing the means. 
Right? 

101
00:06:39,124 --> 00:06:42,601
So the average affect in the control 
group was different than the average 

102
00:06:42,601 --> 00:06:48,511
affect in the experimental group, okay. 
Now, a lot of what statistics is about is 

103
00:06:48,511 --> 00:06:54,080
actually designing the experiment to 
collect the data. 

104
00:06:54,080 --> 00:06:58,777
And in a data science regime, in a big 
data regime, we're actually less 

105
00:06:58,777 --> 00:07:03,628
frequently in the con-, in, in in a 
position to design these things in the 

106
00:07:03,628 --> 00:07:08,231
first place. 
A lot of times we are dealing with data 

107
00:07:08,231 --> 00:07:10,559
that we did not Is how they collect. 
Okay. 

108
00:07:10,559 --> 00:07:15,389
So that's maybe one difference between 
classical statistics and the, the way I 

109
00:07:15,389 --> 00:07:20,495
want to present this material for 
purposes of data science. 

110
00:07:20,495 --> 00:07:22,111
Okay. 
That being said, it's important to 

111
00:07:22,111 --> 00:07:25,261
understand that careful experimental 
design is really the most important 

112
00:07:25,261 --> 00:07:28,240
[LAUGH] thing there is in all this work, 
right. 

113
00:07:28,240 --> 00:07:33,980
The, the, the analysis techniques are 
second fiddle to the proper collection of 

114
00:07:33,980 --> 00:07:37,644
data, okay. 
So this includes things like randomized 

115
00:07:37,644 --> 00:07:41,780
trials, blinded and double blinded. 
And so, you know. 

116
00:07:41,780 --> 00:07:44,489
What is blinded, means that the, 
participants themselves do not know which 

117
00:07:44,489 --> 00:07:48,400
group they're in. 
And that's, pretty much non-negotiable, 

118
00:07:48,400 --> 00:07:51,067
right? 
You can't tell people that they're 

119
00:07:51,067 --> 00:07:55,225
getting, a placebo drug versus the, the 
actual drug or where they'll, they'll, 

120
00:07:55,225 --> 00:08:01,245
you know, the [LAUGH] they'll report 
their symptoms differently as an effect. 

121
00:08:01,245 --> 00:08:05,445
randomize is also, would be non 
negotiable except the fact that it's 

122
00:08:05,445 --> 00:08:09,821
difficult to achieve in practice, in some 
cases. 

123
00:08:09,821 --> 00:08:14,791
So randomized would mean we we draw a 
sample through some method and then we 

124
00:08:14,791 --> 00:08:19,691
assign them to the groups to the control 
group and the experimental group with no 

125
00:08:19,691 --> 00:08:23,975
process whatsoever. 
Right? 

126
00:08:23,975 --> 00:08:26,235
It's just purely random. 
Okay. 

127
00:08:26,235 --> 00:08:31,500
So this framework expressed in just these 
sort of few bullets at a high level is 

128
00:08:31,500 --> 00:08:34,961
unbelievably powerful. 
Right? 

129
00:08:34,961 --> 00:08:37,190
It's completely universal to data 
analysis. 

130
00:08:37,190 --> 00:08:40,063
It's really important to internalize 
these points. 

131
00:08:40,063 --> 00:08:43,391
And we'll go into some detail on the 
other aspects that aren't included in 

132
00:08:43,391 --> 00:08:46,970
these slides. 
So, some examples you can dream up, you 

133
00:08:46,970 --> 00:08:50,935
know, that the measuring the effect of a 
new ad placement on your website, 

134
00:08:50,935 --> 00:08:55,355
compared to the control group of the 
existing placement measuring the effect 

135
00:08:55,355 --> 00:09:02,200
of a treatment against a sugar pill, or 
the best existing treatment. 

136
00:09:02,200 --> 00:09:05,020
Okay. 
And everything else you might imagine. 

137
00:09:05,020 --> 00:09:08,896
So, to summarize hypothesis testing, you 
can organize the terminology into this 

138
00:09:08,896 --> 00:09:14,334
grid here where there's two possibilities 
[MUSIC] for this true state of the world. 

139
00:09:14,334 --> 00:09:18,102
One is that the null hypothesis is true. 
There is no difference between the 

140
00:09:18,102 --> 00:09:22,568
control group and the experimental group. 
And the other is that the null hypothesis 

141
00:09:22,568 --> 00:09:25,137
is false. 
That there is an effect that you're 

142
00:09:25,137 --> 00:09:26,444
measuring. 
Okay. 

143
00:09:26,444 --> 00:09:30,044
And so, then, there's also two 
possibilities for the outcome of your 

144
00:09:30,044 --> 00:09:33,904
statistical test. 
In one case, you do not reject the null 

145
00:09:33,904 --> 00:09:35,460
hypothesis. 
Right? 

146
00:09:35,460 --> 00:09:38,680
You find no evidence that there's any 
difference between the groups. 

147
00:09:38,680 --> 00:09:40,340
And the other is that you reject the null 
hypothesis. 

148
00:09:40,340 --> 00:09:43,146
You do find evidence that there's 
differences between the groups. 

149
00:09:43,146 --> 00:09:44,716
Okay? 
So if the null hypothesis is true, there 

150
00:09:44,716 --> 00:09:47,250
is no difference. 
But you detect a difference. 

151
00:09:47,250 --> 00:09:51,072
That's a Type 1 error. 
And the, rate at which that happens you 

152
00:09:51,072 --> 00:09:56,294
can, is, we will refer to as alpha. 
And that might come up at times as we 

153
00:09:56,294 --> 00:10:00,952
may, as we have this discussion. 
[INAUDIBLE] And if you make the correct 

154
00:10:00,952 --> 00:10:05,712
decision, then the probability of that is 
1 minus alpha, when the null hypothesis 

155
00:10:05,712 --> 00:10:10,274
is true. 
When the null hypothesis is false and you 

156
00:10:10,274 --> 00:10:14,960
fail to reject it, alright, there is an 
effect and you fail to measure it, that's 

157
00:10:14,960 --> 00:10:20,881
Type 2 error and that's beta. 
And when you do reject a null hypothesis 

158
00:10:20,881 --> 00:10:24,785
when it's false, right, you detect an 
effect when there is an effect to detect 

159
00:10:24,785 --> 00:10:31,580
that's one minus beta. 
And this is called this, the power of the 

160
00:10:31,580 --> 00:10:36,054
test, the statistical power. 

