[MUSIC]. As we talked about Bayesian versus Frequentist approaches to two statistics, and this cartoon is my attempt to a quick summary. So the Frequentist approach is concerned with using the data and only the data to make a decision. While the Bayesian approach incorporates prior belief as well as the data, okay. And that's both a strength and a weakness. Strength in that, if you have good information, you can incorporate it. If the weakness in that, you have to actually incorporate information and that it's somewhat subjective, depending on who you're talking to, fine. So Bayes' Theorem, breaking it down in a little more detail, you've got this terminology you can apply to it. So the prior, as we said, here in blue, is the probability of the hypothesis being true before you've collected any data, right? So what's the, of the, in the global population, how many people are above a certain height? Okay, then the marginal probability is what is the probability of collecting this data, this particular data observations under all possible hypotheses, okay. And then the likelihood is the probability of collecting this data given that our hypothesis is actually true. And then finally the posterior is what we're usually trying to compute, which is the probability of our hypothesis being true, given the data collected, okay. So, let's think about it this way, arrange this into a grid. This is like, you know, basically, it's actually pretty trivial to derive this theorem, if you think carefully about what's going on. So, a nice way to explain it is so build a two by two grid like this. Where you have A, and B, and not A and not B. And, the box here is the probability of A and B happening, okay. Well, A and B is the probability of A given that B has already occurred, multiplied by the probability that B occurs, right? And you should take a moment and convince yourself of that. Equivalently, is the probability that B of B occurring, given that A has already occurred, multiplied by the probability that A actually occurs, okay. So these things are all equivalent. Set these two things equal to each other, and just, you know, divide. So set these two things equal to one another, and divide the both sides by the probability of B occurring and you have, Bayses' rule. All right so let's try this out. So, there's a question. Say that you know that 1% of women at age 40 who participate in routine screening, have Breast Cancer. And let's say you know 80% of women who actually do have Breast Cancer will get a positive, result from the test. Further, say you know that 9.6% of women who do not have Breast Cancer will also get positive results. So these, this is the false positive rate, okay. Now given that you know a woman in this age group had a positive test. All right, the test came back positive in a routine screening. What is the probability that she actually has breast cancer? And you should take a minute to sort of work this out. So what's sort of remarkable is that, this is a fairly straightforward application of Bayes' rule. But intuitively it's easy to make sort of mistakes in, in the reasoning. And in fact when you ask, there's been some famous studies that are now somewhat out of date where they asked doctors this question. And the doctors came up with wildly wrong answers. Only 15% of the, doctors they asked were able to answer the question correctly. Okay, so, here's how to break it down, again with our two by two grid. We've got 1% of cases that have cancer, and we've got 99, which means we have 99% of cases that do not have cancer. This is in the global population, right? And let's say we just know this. All right, we know that 1% of people, across the board, have cancer, or at least in this group that we're studying. I shouldn't say across the board, okay. Now, we also know that if you have cancer, there's an 80% chance that when you take the test it'll come back positive, okay. So the true positive rate is 1% have cancer times 80%. Now, we're also given that if you do not have cancer and you take the test, you'll get a positive result of 9, at a rate of 9.6%. And so the false positive rate is the rate of not having cancer, 99% times 9.6%. Now, suddenly for the false negative and the true negative tests, 1% times 20 times 1 minus 80 and 99% times 1 minus 9.6 is 90.4. Okay, so we have all the information we need. So going back to actually answering the questions we can, write down, Bayes' Rule. So, what is the probability of having cancer given that we have a positive test result? Well, that's the probability of getting a positive test result, given we have cancer. Multiplied by the probability of having cancer overall, divided by the probability of a positive test overall, okay. So, the only one of these terms, so we actually know this terms straight out, we're given this one, all right, which is 80%. And we're given this one, the probability of cancer overall is, is 1%. We don't have this denominator given, so we need to figure that out. Well, that's the chance of a positive test in all other occurrences, the chance of a positive test. Given that we have cancer and the chance of a positive test given that they don't have cancer. In each case multiplied by the probability of that happening. Okay, so you can, you can decompose this. So that's 0.8 times the 1% probability of actually having cancer plus 9.6% times 99% which gives you this number of 10.3%. So that's the overall probability of getting a positive test result, all right? And so now, you plug this in, you end up with a number of 7.8% for our, of our answer. If we have a positive test result, the chance of actually having cancer is 7.8%. So this is lower than you might come up with if you don't think about it carefully, right? You might think that boy, I got a positive test result. There's something, there's this 80% floating around. You know, boy, there's probably a 70 to 80% change that that, I have cancer. But because of the very low percentage of having caner in the, in the prior probabilities. The actual number is still pretty low, okay? So it's easy to make mistakes with this stuff. Now, this was a remarkably simple case. First of all, we're given all this information. Second of all, there's only two possibilities, these sort of binary variables. So let's think about something a little more complicated. Okay, so let's consider a classic application of Bayes' rule to a big data problem, which is spam filtering. Okay, so here our task is to determine whether an email message is spam. The probability that an email message is spam, given the words in the email message. Okay, and with Bayes' rule, you can express that probability as the probability that the email message is spam overall. Multiplied by the probability of seeing these particular words in the message, given that we already know it's spam. And all that divided by the probability of seeing these words in the message. Now, the interesting one here is this numerator. And the reason is that, the probability of words appearing doesn't involve the unknown label of whether it's spam or not. And so all we're trying to do is get a relative frequency of spam or not spam, okay? And so, just dividing by a constant factor of the probability of seeing these words doesn't change our decision at all. So, we don't care about the actual number, we just care about the decision of spam or not spam. Okay, so fine, so, re-expressing is before we get rid of the denominator. Re-expressing this, what do we mean by words? Well, this will, you can write this as probability that the email message is spam, given that the word viagra appears in the message. Given that the word rich, appears in the message. Given that the word, something more innocuous, perhaps like friend appears in the message. So all the words, you know, in the English language or, or, all the words of, of interest to us in this test, okay. So that's, re-expressed that way. Now, this numerator can be rewritten in the following way. Given that it's a conditional probability, we can apply, a chain rule, repeated, a repeated application of the definition of conditional probability to obtain this. The probability that it's spam multiplied by, so let's see, so this expression rather, can be expressed as the probability seeing the word viagra. Given that it's spam, multiplied by the probability of seeing all these other words. given that it's spam and given that it's viagra. Or given that the email message contains viagra. And you can keep going. This probability times the probability of seeing the word rich, given that it's spam, and given that is viagra. Multiplied by the probability of all the other words, given that it is spam, given that it contains the word viagra. Given that it contains the word rich, and so on, okay. So this is a long, complicated, conditional probability. And this is where the Naive Bayes assumption comes in. So, under the Naive Bayes assumption we say that the probabilities of these different words appearing in this email message are completely independent. That it's no more likely for you to see the word rich, when you see the word wealth than it is, you know, without the word wealth there. Okay, this isn't true, right? Obviously words go, go together. There are co-occurrence rates, right? But you just ignore that, and just treat everything as completely independent, which allows you to simplify this expression as just a sequence of probabilities. What is the probability of seeing the word viagra, given that it's spam? What is the probability of seeing the word rich, given that it's spam, and so on? Now, how do you get these probabilities? Well, you have data, right? You had a set of documents that had been pre-labeled as spam and you can look at the number of them that contain the word rich. And divide by the total number, okay. And so now you can calculate the probability of, of the two classes, spam and not spam. And apply a decision procedure to call it. And in fact, a simple one is just whichever one is more likely, whichever one has the higher probability. It's called the MAP decision rule which stands for the Maximum A Posteriori. Okay.