[MUSIC]. Okay, so, so far we've been discussing statistical inference from a particular perspective, which is the frequentest perspective. And so frequentists are concerned with the Probability of seeing a particular data sample, given the null hypothesis. Okay, and that's what the p-value gives you. But there's another perspective to statistics called the Bayesian approach, which is concerned with the probability of Having a certain outcome given the data we've already seen. Okay, and so I want to talk a little bit about Bayesian approaches and give an example of their power. Okay. What are the differences here? Well, you can think about the differences In terms of what is fixed. So the frequency's perspective is that data are a repeatable random sample, right? You're always allowed to go do another experiment the same way, and in fact you need to at least virtually, to reason about, the probabilities that get produced by these methods. Okay. So there is some sort of frequency you can reason about, frequency of of achieving a certain outcome. Okay. And the underlying parameters of the population remain constant during this repeatable process. Right, your, your, the population stays fixed and you run experiments to determine, you know likelihoods. Okay, meanwhile, with the Bayesian approach, the data observed from a realized sample, and the parameters A, of the population or unknown, but can be described probabilistically, right. So there not sort of, fixed values. However, the data are fixed, right, you don't think about going back and sampling more data, you just have the observations you have, okay. And so, the Bayesian approach is 100% concerned with the application of Bayesian rule which is this. Okay, so this, what Thomas Bayesian did was relate these conditional probabilities with the prior beliefs. Okay, so it allows you to take your- Belief about the probabilities of certain events happening and update them when more data is collected, okay. And so here's what it says, it says the probability of event A happening, given that event B has already happened. Is equal to the probability of b happening given that a has already happened, multiplied by the probability of a happening across the board, divided by the probability of be happening across the board. And if that's not clear, that's okay. We're, we're going to go into more detail here. So right now though, recognize that the key benefit here is the ability to incorporate prior knowledge, which is not the case with the frequency approach. Now, a key weakness here is that you need to incorporate prior knowledge So if you, so there's a couple of problems with this. One is if you don't know anything about the population you're modeling then there is not much you can do with the prior, there's not a good way to model this, the distributions. Now, there are some techniques to sort of derive so called uninformative priors that try not to influence things to much but- Give you a plugin to be able to rule. But still, that is an issue. fine. Perhaps more insideously and the reason why Bayes' the Bayesian approach, was popular and then it fell out of favor in the earlier part of the twentieth century. Was that, you can use this rule to kind of, do anything you want, to confirm or deny the affect of any sort of evidence, right? Just by plugging in your own prior belief. Okay. So, given a fixed set of data, two different people can come up with two different conclusion about the data because they had prior beliefs. They modeled the, the, the prior distributions differently, okay. And this was seen as a major flaw, and gave rise to the frequency approach, which. Looks more or less objectively at the data itself. Okay. Right. And so that's what is spelled out here in a nice essay by Matthews 1998, is that different people could use Bayesian Theorem and get different results. And so, faced with some, some experimental evidence per se ESP true believers could uses Bayes's Theorem to sh show that the new results confirmed it, while skeptics could use it to show that it the ESP didn't exist. Okay, and both views are possible because Bayesian Theorem only shows how to alter one's prior level of belief. And different people can start out with different opinions. Okay, so that's the issue. However the frequency oriented, the frequentist approach So, the frequency approach, laid out by Fisher and others, were able to achieve what was thought to be impossible. Which is a way of judging the significance of experimental data independent of any prior beliefs. Okay. So he'd found a way that anyone could use to show that a result was too impressive, too statistically significant to be dismissed as a fluke. And, you know, all you had to do was convert your raw data into this thing called a P-value. Alright. Now, you have to have some sort of threshold to measure the significance of this P-value, and that was written up as 0.05 in by Fisher. You know, in a, in a, in a original paper. So what were the insights that led to this particular value of 0.05? Well, Fisher admitted that there weren't any. He simply decided 0.05 because it was mathematically convenient. Okay? And you know, in the many people have pointed out, say from the 60s on, there's been periods of time where there's been a significant amount of work showing that this approach, with these particular thresholds are able to produce a lot of incorrect conclusions. So James Berger Purdue wrote his entire series of papers warning about the quote astonishing tendency of Fisher's p-values to exaggerate significance. Findings that met the 120 standard can actually arise when the data provide little, very little of no evidence in favor of an effect. Okay, so that's the problem, so now perhaps the, we can say that the pendulum is in some sense swinging back towards the Bayesian approach. One more problem with the Bayesian approach that I'll bring up that I don't necessarily have a slide on, is that, which we'll see in a little bit of detail in. Perhaps the next segment, is that these conditional probabilities and these prior probabilities end up forcing you to, to model a situation mathematically, and that situation may be very difficult to model mathematically. Okay. So what you end up with is these, these chains of complicated conditional probabilities that need to be integrated in order to apply Bayes' rule, and so this was computationally intractable, and there were a norm, a huge number of tricks and simplications and, and ideas used to make this more tractable. Okay. But, thanks to the development of methods that allow you sample these complicated distributions computationally these, this approach is becoming increasingly popular. Okay. And there's a lot of thinkers in this space that believe that, you know, for the 21st century, a combination of frequent and Bayesian approaches is going to be dominant. Okay, so for these reasons we're going to make sure we talk about the Bayesian approach at a very introductory level and leading up to machine learning algorithm called naive bayes.