[MUSIC]. Okay. So a third potential explanation of this decline effect is exceedingly simple but exceedingly common. And it's maybe the most important one to learn to apply in your own work when you're, when you're doing statistical analysis. Okay. And so this is the idea of multiple hypothesis testing. So the problem here is that if you perform experiments over and over and over again, you're bound to find something. Right? That's sort of the definition in fact. Right? You, if you keep rolling dice long enough, you, you know, you can't just keep rolling dice and yell Yahtzee, right? You get one shot at it. And the same is true with these experimental design. Okay, and so this is related to the publication bias problem in that you're only showing your positive results but it's a little bit different. Because here you're talking about the same sample and you're testing different hypotheses over the same data. And so in these situations, either you shouldn't do it at all or if you do have to do it for various reasons, you need to adjust the significance level down. That means you need to settle not for 0.05 as the threshold. You need to do something much, much lower. Okay. So, to understand why, you know, consider something pretty basic. Over completely random data and we'll, we set the threshold at 0.05, as alpha 0.05. So the probability of detecting effect, where there is none, is 0.05. Then the probability of [SOUND] detecting an effect when it exists is 1 minus alpha. Then the probability of detecting an effect when it exists on every experiment you do, out of k experiments, is 1 minus alpha, times 1 minus alpha, times 1 minus alpha, times 1 minus alpha, assuming that they're independent. Right? We're [UNKNOWN] it's okay to multiply probabilities together if those probabilities are independent. Okay. Then the probability of, finally the probability of detecting an effect, where there is none, on at least one experiment is 1 minus that total, right? So, first we build up the probability of being perfect, and then 1 minus that is the probability of not being perfect, without making at least one mistake. Okay. So if you plot these numbers, what you get is, you know, in the x, x here's the number of texts, and the y axis is the probability of as least one spurious finding. Right? Making at least one mistake. Well, it goes up like this. So, if, as you get sort of 50 hypothesis tests, you know your up at the 90% chance of at least one spurious finding. Okay? And so, controlling this is known as controlling the Familywise Error Rate. This is the Familywise Error Rate of at least one mistake. So this is a pretty stringent. So, what do we do about this multiple testing problem? How do we control the familywise error rate? Well, one solution is the Bonferroni Correction which is you just divide by the number of hypotheses. So if your significance level is Alpha, 0.05 then you do 20 experiments, 20, you're testing 20 hypothesis, you just divide 0.05 by 20. So another correction is the Sidak Correction, which has this extra condition where he, the tests are, it need to be independent. So you know, we, we talked about it in the last slide that in order to make that plot we were assuming that they were independent. But the Bonferroni Correction in general does not need to assume that. Okay. So if you're doing hypothesis tests that are related to each other you could still do the Bonferroni correction. However, to derive this Sidak Correction we're going to rely on the fact that we're going to multiply the probabilities together, whenever you see probabilities being multiplied together that means that you assuming they're independent. Okay? So let's see if we can build this up. So hear we are going to derive the individual task, the corrected significance level from the overall significance level. So we are going to set the overall significance level Alpha equal to the probability that at least one of the test Is significant. All right. So at least one is significant. Well, what's that? That's 1 minus the probability that none of them are significant. And the probability that none of them are significant, assuming independence, is the probability that the first one is not significant times the probability the second one is insignificant is not significant and so one. So that's 1 minus alpha c raised to the k experiment. All right, 1 minus alpha c, times 1 minus alpha c, times 1 minus alpha c, and so on. And that's what this expression says. So [MUSIC] fine, so now we just solved for alpha c and we get this expression, 1 minus, 1 minus alpha, raised to the one over k, raised to, raised to, raised to the k root. The kth root of 1 minus alpha. Okay. Okay. So showing the same plot from before. But now zooming the scale in down around 0.05, where the original s-, significance level was. You can see the difference between these 2 corrections. So the Sidak Correction is more conservative than [UNKNOWN] than the Bonferroni one correction. So Bonferroni evens it out across. So, instead of, instead of increasing the likelihood of making a mistake quickly, which is what the previous plot shows, this is zooming in at 0.05 and showing that the Bonferroni makes a constant across, regardless how many tests. Which makes sense, you're just dividing them by the number of tests you've done. Okay? The Sidak Correction is even more conservative. All right. That's what to remember. Both of these are considered to be more conservative than is perhaps necessary. You lose too, you give up too much statistical power when you use these. And in fact, any correction for the, i-, it goes back to the actual definition of family wise error rate is considered to be too conservative. So another way of controlling for multiple hypothesis tests that is less conservative, is by considering the false discovery rate. Okay. And so, the false recovery state you can understand by going back to our grid and labeling it a slightly different way. Okay? So here, the, excuse me, the mnemonic here is that, the total, lets see, T and F stand for true and false and D and N stand for discovery and nondiscovery. So, false discovery is FD. True discovery is TD. True nondiscovery is TN and false nondiscovery is FN. Okay. With this notation, the false discovery rate, FDR, which is sometimes called Q is the number of false discoveries over the total number of discoveries. And so here in this notation, by the way, D, you know, is equal to FD plus TD. So these are counts, these are the number of, of, of, you know, true relationships and false relationships and so on. Okay? So this is the rate you're trying to control for. So the Bonferroni Correction and other, and other Familywise error rate corrections tend to wipe out evidence of the most interesting effects. We say they suffer from low power. So the false discovery rate controls offer a way to increase power while maintaining still some principled bound on error. Okay. And so it's more, intuitively is based on the assessment that, you know, four false discoveries out of ten, you know, if you reject the null hypothesis ten times, you make, you make five, this quote, quote, you know, discoveries. Four, having four of those be false is really bad. But if, you know, much worse than making 20 false discoveries out of 100. [INAUDIBLE] is that, you know, finding true effects is a good thing. And so even though you're going to make some mistakes, if you can, the more you find the more further value you've added. So how can you control the false discovery rate? Well the Benjamini-Hochberg procedure gives you a way to do this. And so, here's how it works. You compute the p-value of your hypothesis, and then you sort them in descending order. Such that the ones with the lowest p-value, which are the most likely hypothesis. Right? The ones that are best supported by the evidence. Come first, okay. And then you apply this condition where the P value is subject to a more stringent condition then just alpha. Remember alpha is your 0.05, your cutoff. And what we're trying to do is correct from multiple hypotheses testing. So we want a much more stringent alpha. And so that more stringent alpha is this ratio i over m. And so i [SOUND] is just the rank order, of the hypotheses you're testing. And m is your total number of hypo, hypotheses that you're testing. Number of hypotheses. [BLANK_AUDIO] All right. And the procedure says, well find that highest i for which this condition holds and then reject the null hypotheses for all i lower than that except everything up until that point. Right, okay. And so, here's what it might look like with 50. The first and, and 0.05, I suppose I should've put that. So the first, your first hypotheses has to, the first hypothesis is compared with a pretty stringent conditions zero, you know, 1 in 1,000. And the second one is double that, and third one is triple that, and so on. All the way up to the 50th one, which would be 50 over 50, which is just your original alpha 0.05. Okay? So this is a much tighter condition. And so what they were able to prove is that the false, under these conditions, you know, following this procedure the false discovery rate is less than [SOUND] T over m, times alpha. Where, if you remember T was, the total number of, cases where the null hypothesis is true. Okay. So here's what it looks like graphically. The little x's are mean, are above this line. And the dots are below this line. And the x axis is rank order. And these are all your 50 hypotheses sorted, in increasing P value. And the line re-, represents that threshold condition, and slopes up with rank order, as we'd expect. And so we'd say we find the highest i for which this is, this condition holds and accept everything lower than this. And here we had you know, a pretty good run. We accepted sort of 30 out of 50 hypothesis. And notice that these are actually are above the line, but we would still accept them. [SOUND]. Okay.