1
23:59:59,500 --> 00:00:06,063
[MUSIC]. 

2
00:00:06,063 --> 00:00:08,599
Okay. 
So a third potential explanation of this 

3
00:00:08,599 --> 00:00:12,694
decline effect is exceedingly simple but 
exceedingly common. 

4
00:00:12,694 --> 00:00:17,309
And it's maybe the most important one to 
learn to apply in your own work when 

5
00:00:17,309 --> 00:00:21,918
you're, when you're doing statistical 
analysis. 

6
00:00:21,918 --> 00:00:24,838
Okay. 
And so this is the idea of multiple 

7
00:00:24,838 --> 00:00:28,461
hypothesis testing. 
So the problem here is that if you 

8
00:00:28,461 --> 00:00:31,588
perform experiments over and over and 
over again, you're bound to find 

9
00:00:31,588 --> 00:00:33,540
something. 
Right? 

10
00:00:33,540 --> 00:00:36,406
That's sort of the definition in fact. 
Right? 

11
00:00:36,406 --> 00:00:39,415
You, if you keep rolling dice long 
enough, you, you know, you can't just 

12
00:00:39,415 --> 00:00:42,380
keep rolling dice and yell Yahtzee, 
right? 

13
00:00:42,380 --> 00:00:45,634
You get one shot at it. 
And the same is true with these 

14
00:00:45,634 --> 00:00:48,742
experimental design. 
Okay, and so this is related to the 

15
00:00:48,742 --> 00:00:51,824
publication bias problem in that you're 
only showing your positive results but 

16
00:00:51,824 --> 00:00:55,618
it's a little bit different. 
Because here you're talking about the 

17
00:00:55,618 --> 00:00:59,499
same sample and you're testing different 
hypotheses over the same data. 

18
00:01:01,110 --> 00:01:05,400
And so in these situations, either you 
shouldn't do it at all or if you do have 

19
00:01:05,400 --> 00:01:11,630
to do it for various reasons, you need to 
adjust the significance level down. 

20
00:01:11,630 --> 00:01:15,750
That means you need to settle not for 
0.05 as the threshold. 

21
00:01:15,750 --> 00:01:17,035
You need to do something much, much 
lower. 

22
00:01:17,035 --> 00:01:22,819
Okay. 
So, to understand why, you know, consider 

23
00:01:22,819 --> 00:01:30,195
something pretty basic. 
Over completely random data and we'll, we 

24
00:01:30,195 --> 00:01:37,054
set the threshold at 0.05, as alpha 0.05. 
So the probability of detecting effect, 

25
00:01:37,054 --> 00:01:41,488
where there is none, is 0.05. 
Then the probability of [SOUND] detecting 

26
00:01:41,488 --> 00:01:44,250
an effect when it exists is 1 minus 
alpha. 

27
00:01:44,250 --> 00:01:48,342
Then the probability of detecting an 
effect when it exists on every experiment 

28
00:01:48,342 --> 00:01:52,434
you do, out of k experiments, is 1 minus 
alpha, times 1 minus alpha, times 1 minus 

29
00:01:52,434 --> 00:01:58,326
alpha, times 1 minus alpha, assuming that 
they're independent. 

30
00:01:58,326 --> 00:02:00,642
Right? 
We're [UNKNOWN] it's okay to multiply 

31
00:02:00,642 --> 00:02:04,810
probabilities together if those 
probabilities are independent. 

32
00:02:04,810 --> 00:02:07,003
Okay. 
Then the probability of, finally the 

33
00:02:07,003 --> 00:02:11,413
probability of detecting an effect, where 
there is none, on at least one experiment 

34
00:02:11,413 --> 00:02:16,956
is 1 minus that total, right? 
So, first we build up the probability of 

35
00:02:16,956 --> 00:02:21,500
being perfect, and then 1 minus that is 
the probability of not being perfect, 

36
00:02:21,500 --> 00:02:26,350
without making at least one mistake. 
Okay. 

37
00:02:26,350 --> 00:02:29,450
So if you plot these numbers, what you 
get is, you know, in the x, x here's the 

38
00:02:29,450 --> 00:02:32,500
number of texts, and the y axis is the 
probability of as least one spurious 

39
00:02:32,500 --> 00:02:34,950
finding. 
Right? 

40
00:02:34,950 --> 00:02:38,404
Making at least one mistake. 
Well, it goes up like this. 

41
00:02:38,404 --> 00:02:43,669
So, if, as you get sort of 50 hypothesis 
tests, you know your up at the 90% chance 

42
00:02:43,669 --> 00:02:48,242
of at least one spurious finding. 
Okay? 

43
00:02:48,242 --> 00:02:54,290
And so, controlling this is known as 
controlling the Familywise Error Rate. 

44
00:02:54,290 --> 00:02:56,910
This is the Familywise Error Rate of at 
least one mistake. 

45
00:02:56,910 --> 00:03:01,210
So this is a pretty stringent. 
So, what do we do about this multiple 

46
00:03:01,210 --> 00:03:04,326
testing problem? 
How do we control the familywise error 

47
00:03:04,326 --> 00:03:07,138
rate? 
Well, one solution is the Bonferroni 

48
00:03:07,138 --> 00:03:11,990
Correction which is you just divide by 
the number of hypotheses. 

49
00:03:11,990 --> 00:03:17,156
So if your significance level is Alpha, 
0.05 then you do 20 experiments, 20, 

50
00:03:17,156 --> 00:03:23,990
you're testing 20 hypothesis, you just 
divide 0.05 by 20. 

51
00:03:23,990 --> 00:03:27,125
So another correction is the Sidak 
Correction, which has this extra 

52
00:03:27,125 --> 00:03:31,340
condition where he, the tests are, it 
need to be independent. 

53
00:03:31,340 --> 00:03:34,442
So you know, we, we talked about it in 
the last slide that in order to make that 

54
00:03:34,442 --> 00:03:37,722
plot we were assuming that they were 
independent. 

55
00:03:37,722 --> 00:03:41,360
But the Bonferroni Correction in general 
does not need to assume that. 

56
00:03:41,360 --> 00:03:43,276
Okay. 
So if you're doing hypothesis tests that 

57
00:03:43,276 --> 00:03:47,200
are related to each other you could still 
do the Bonferroni correction. 

58
00:03:47,200 --> 00:03:49,632
However, to derive this Sidak Correction 
we're going to rely on the fact that 

59
00:03:49,632 --> 00:03:52,292
we're going to multiply the probabilities 
together, whenever you see probabilities 

60
00:03:52,292 --> 00:03:56,335
being multiplied together that means that 
you assuming they're independent. 

61
00:03:56,335 --> 00:04:00,340
Okay? 
So let's see if we can build this up. 

62
00:04:00,340 --> 00:04:03,512
So hear we are going to derive the 
individual task, the corrected 

63
00:04:03,512 --> 00:04:08,360
significance level from the overall 
significance level. 

64
00:04:08,360 --> 00:04:12,254
So we are going to set the overall 
significance level Alpha equal to the 

65
00:04:12,254 --> 00:04:17,026
probability that at least one of the test 
Is significant. 

66
00:04:17,026 --> 00:04:20,380
All right. 
So at least one is significant. 

67
00:04:20,380 --> 00:04:23,334
Well, what's that? 
That's 1 minus the probability that none 

68
00:04:23,334 --> 00:04:26,452
of them are significant. 
And the probability that none of them are 

69
00:04:26,452 --> 00:04:29,764
significant, assuming independence, is 
the probability that the first one is not 

70
00:04:29,764 --> 00:04:32,740
significant times the probability the 
second one is insignificant is not 

71
00:04:32,740 --> 00:04:38,212
significant and so one. 
So that's 1 minus alpha c raised to the k 

72
00:04:38,212 --> 00:04:41,913
experiment. 
All right, 1 minus alpha c, times 1 minus 

73
00:04:41,913 --> 00:04:44,820
alpha c, times 1 minus alpha c, and so 
on. 

74
00:04:44,820 --> 00:04:47,831
And that's what this expression says. 
So [MUSIC] fine, so now we just solved 

75
00:04:47,831 --> 00:04:51,796
for alpha c and we get this expression, 1 
minus, 1 minus alpha, raised to the one 

76
00:04:51,796 --> 00:04:56,483
over k, raised to, raised to, raised to 
the k root. 

77
00:04:56,483 --> 00:05:06,380
The kth root of 1 minus alpha. 
Okay. 

78
00:05:06,380 --> 00:05:08,180
Okay. 
So showing the same plot from before. 

79
00:05:08,180 --> 00:05:13,292
But now zooming the scale in down around 
0.05, where the original s-, significance 

80
00:05:13,292 --> 00:05:16,750
level was. 
You can see the difference between these 

81
00:05:16,750 --> 00:05:19,630
2 corrections. 
So the Sidak Correction is more 

82
00:05:19,630 --> 00:05:24,695
conservative than [UNKNOWN] than the 
Bonferroni one correction. 

83
00:05:24,695 --> 00:05:28,880
So Bonferroni evens it out across. 
So, instead of, instead of increasing the 

84
00:05:28,880 --> 00:05:32,312
likelihood of making a mistake quickly, 
which is what the previous plot shows, 

85
00:05:32,312 --> 00:05:35,536
this is zooming in at 0.05 and showing 
that the Bonferroni makes a constant 

86
00:05:35,536 --> 00:05:40,460
across, regardless how many tests. 
Which makes sense, you're just dividing 

87
00:05:40,460 --> 00:05:43,260
them by the number of tests you've done. 
Okay? 

88
00:05:43,260 --> 00:05:46,050
The Sidak Correction is even more 
conservative. 

89
00:05:46,050 --> 00:05:49,830
All right. 
That's what to remember. 

90
00:05:49,830 --> 00:05:56,990
Both of these are considered to be more 
conservative than is perhaps necessary. 

91
00:05:56,990 --> 00:06:01,140
You lose too, you give up too much 
statistical power when you use these. 

92
00:06:01,140 --> 00:06:04,374
And in fact, any correction for the, i-, 
it goes back to the actual definition of 

93
00:06:04,374 --> 00:06:08,160
family wise error rate is considered to 
be too conservative. 

94
00:06:08,160 --> 00:06:12,613
So another way of controlling for 
multiple hypothesis tests that is less 

95
00:06:12,613 --> 00:06:17,670
conservative, is by considering the false 
discovery rate. 

96
00:06:17,670 --> 00:06:20,001
Okay. 
And so, the false recovery state you can 

97
00:06:20,001 --> 00:06:24,133
understand by going back to our grid and 
labeling it a slightly different way. 

98
00:06:24,133 --> 00:06:28,226
Okay? 
So here, the, excuse me, the mnemonic 

99
00:06:28,226 --> 00:06:34,814
here is that, the total, lets see, T and 
F stand for true and false and D and N 

100
00:06:34,814 --> 00:06:44,350
stand for discovery and nondiscovery. 
So, false discovery is FD. 

101
00:06:44,350 --> 00:06:49,296
True discovery is TD. 
True nondiscovery is TN and false 

102
00:06:49,296 --> 00:06:53,082
nondiscovery is FN. 
Okay. 

103
00:06:54,900 --> 00:06:59,660
With this notation, the false discovery 
rate, FDR, which is sometimes called Q is 

104
00:06:59,660 --> 00:07:05,600
the number of false discoveries over the 
total number of discoveries. 

105
00:07:05,600 --> 00:07:12,180
And so here in this notation, by the way, 
D, you know, is equal to FD plus TD. 

106
00:07:12,180 --> 00:07:17,280
So these are counts, these are the number 
of, of, of, you know, true relationships 

107
00:07:17,280 --> 00:07:21,436
and false relationships and so on. 
Okay? 

108
00:07:21,436 --> 00:07:25,220
So this is the rate you're trying to 
control for. 

109
00:07:25,220 --> 00:07:30,039
So the Bonferroni Correction and other, 
and other Familywise error rate 

110
00:07:30,039 --> 00:07:36,490
corrections tend to wipe out evidence of 
the most interesting effects. 

111
00:07:36,490 --> 00:07:39,397
We say they suffer from low power. 
So the false discovery rate controls 

112
00:07:39,397 --> 00:07:42,405
offer a way to increase power while 
maintaining still some principled bound 

113
00:07:42,405 --> 00:07:44,580
on error. 
Okay. 

114
00:07:44,580 --> 00:07:48,444
And so it's more, intuitively is based on 
the assessment that, you know, four false 

115
00:07:48,444 --> 00:07:52,252
discoveries out of ten, you know, if you 
reject the null hypothesis ten times, you 

116
00:07:52,252 --> 00:07:57,435
make, you make five, this quote, quote, 
you know, discoveries. 

117
00:07:57,435 --> 00:08:03,400
Four, having four of those be false is 
really bad. 

118
00:08:03,400 --> 00:08:12,705
But if, you know, much worse than making 
20 false discoveries out of 100. 

119
00:08:12,705 --> 00:08:16,060
[INAUDIBLE] is that, you know, finding 
true effects is a good thing. 

120
00:08:16,060 --> 00:08:18,812
And so even though you're going to make 
some mistakes, if you can, the more you 

121
00:08:18,812 --> 00:08:22,900
find the more further value you've added. 
So how can you control the false 

122
00:08:22,900 --> 00:08:26,060
discovery rate? 
Well the Benjamini-Hochberg procedure 

123
00:08:26,060 --> 00:08:29,360
gives you a way to do this. 
And so, here's how it works. 

124
00:08:29,360 --> 00:08:32,609
You compute the p-value of your 
hypothesis, and then you sort them in 

125
00:08:32,609 --> 00:08:35,910
descending order. 
Such that the ones with the lowest 

126
00:08:35,910 --> 00:08:38,850
p-value, which are the most likely 
hypothesis. 

127
00:08:38,850 --> 00:08:39,958
Right? 
The ones that are best supported by the 

128
00:08:39,958 --> 00:08:42,539
evidence. 
Come first, okay. 

129
00:08:42,539 --> 00:08:46,193
And then you apply this condition where 
the P value is subject to a more 

130
00:08:46,193 --> 00:08:52,770
stringent condition then just alpha. 
Remember alpha is your 0.05, your cutoff. 

131
00:08:52,770 --> 00:08:56,410
And what we're trying to do is correct 
from multiple hypotheses testing. 

132
00:08:56,410 --> 00:09:02,081
So we want a much more stringent alpha. 
And so that more stringent alpha is this 

133
00:09:02,081 --> 00:09:06,512
ratio i over m. 
And so i [SOUND] is just the rank order, 

134
00:09:06,512 --> 00:09:13,937
of the hypotheses you're testing. 
And m is your total number of hypo, 

135
00:09:13,937 --> 00:09:23,242
hypotheses that you're testing. 
Number of hypotheses. 

136
00:09:23,242 --> 00:09:28,108
[BLANK_AUDIO] All right. 
And the procedure says, well find that 

137
00:09:28,108 --> 00:09:32,262
highest i for which this condition holds 
and then reject the null hypotheses for 

138
00:09:32,262 --> 00:09:37,442
all i lower than that except everything 
up until that point. 

139
00:09:37,442 --> 00:09:43,333
Right, okay. 
And so, here's what it might look like 

140
00:09:43,333 --> 00:09:48,060
with 50. 
The first and, and 0.05, I suppose I 

141
00:09:48,060 --> 00:09:53,931
should've put that. 
So the first, your first hypotheses has 

142
00:09:53,931 --> 00:10:00,195
to, the first hypothesis is compared with 
a pretty stringent conditions zero, you 

143
00:10:00,195 --> 00:10:04,655
know, 1 in 1,000. 
And the second one is double that, and 

144
00:10:04,655 --> 00:10:08,593
third one is triple that, and so on. 
All the way up to the 50th one, which 

145
00:10:08,593 --> 00:10:12,640
would be 50 over 50, which is just your 
original alpha 0.05. 

146
00:10:12,640 --> 00:10:14,510
Okay? 
So this is a much tighter condition. 

147
00:10:14,510 --> 00:10:18,158
And so what they were able to prove is 
that the false, under these conditions, 

148
00:10:18,158 --> 00:10:22,034
you know, following this procedure the 
false discovery rate is less than [SOUND] 

149
00:10:22,034 --> 00:10:27,154
T over m, times alpha. 
Where, if you remember T was, the total 

150
00:10:27,154 --> 00:10:31,770
number of, cases where the null 
hypothesis is true. 

151
00:10:31,770 --> 00:10:39,800
Okay. 
So here's what it looks like graphically. 

152
00:10:39,800 --> 00:10:42,640
The little x's are mean, are above this 
line. 

153
00:10:42,640 --> 00:10:48,110
And the dots are below this line. 
And the x axis is rank order. 

154
00:10:48,110 --> 00:10:53,890
And these are all your 50 hypotheses 
sorted, in increasing P value. 

155
00:10:53,890 --> 00:10:57,920
And the line re-, represents that 
threshold condition, and slopes up with 

156
00:10:57,920 --> 00:11:02,780
rank order, as we'd expect. 
And so we'd say we find the highest i for 

157
00:11:02,780 --> 00:11:08,743
which this is, this condition holds and 
accept everything lower than this. 

158
00:11:08,743 --> 00:11:11,985
And here we had you know, a pretty good 
run. 

159
00:11:11,985 --> 00:11:14,770
We accepted sort of 30 out of 50 
hypothesis. 

160
00:11:14,770 --> 00:11:19,786
And notice that these are actually are 
above the line, but we would still accept 

161
00:11:19,786 --> 00:11:21,924
them. 
[SOUND]. 

162
00:11:21,924 --> 00:11:22,908
Okay. 

