Hi there, I'm Amy Braverman from JPL and this is, a module on a review of basic probability. In preparation, for a discussion of inference and uncertainty in big data analytics. So here's the outline of this short section,. Some of this maybe beneath many of you. But maybe not. So just proceed and see what happens. I'm going to talk about the following topics. What is probability? Sample spaces and events. The axioms of probability and some of those corollaries of those axioms. And joint and conditional probabilities. I think we all have an idea in our minds about what probability is. Probability it will rain, probability if I throw to dice, I'll get two sixes, probability that the free way will be jammed this morning when I get on it to go to work. Or the probability that global warming is real. Historically, and classically probability has been defined as long run relative frequency which conforms, I think, to those examples that I just gave. The canonical example is flipping a coin. If I flip a coin one time, what's the probability I get a head? Well, it's a half. And one might say that the reason we think that is because, I believe that if I flipped a coin many, many, many times, I should get about heads half the time. And I believe that because I believe that the coin is physically built to come up either head or tail with equal. Long run relative frequency. probability, fortunately, is also a mathematical object. We can say that it's a set function in that it assigns a number, in this case, a number between zero and one to a set. So. Let's, we're going to need a couple of definitions before we, before we proceed with that. Let's define a trial, to be a measurement or observation of some outcome, of some phenomenon. The set of all possible outcomes of a trial is called the sample space and it's traditionally denoted by a capital S. Examples are the possible sexes of a newborn baby. That would be a boy or a girl or the minutes, number of minutes I have to wait to get on the freeway to go to work this morning and that would be a number between zero and infinity tending toward infinity if you live here in Southern California. An event is a subset of the sample space. So event e might be the event that the newborn baby was a boy or the event that I was able to get on the freeway in five minutes in less. So sample spaces and events are sets. And people are very fond of depicting sets and set operations with these things called Venn diagrams, which probably everyone has seen. We represent the whole sample space, S, by the square black rectangle with the white background. And we find subsets of that sample space that represent certain events. So here I have the even A, and the event B. Event A is the red one, the event B is the blue one. And I'm going to define certain operations on those two set. For example the union of A and B show in the upper right. are, all the elements of the sample space that belong to both A and B. I'm sorry, A or B. I make the mistake that I was about to tell you not to make. Or means that the, that the event that you're interested in belongs to. A or B. The intersection is shown in the lower left. I'm sure you're all familiar with that. That's the end operator. That says the element of the sample space is a member of both A and B simultaneously. And over the right, I tried to depict the give you an idea of what the complement is. The complement means not. And in my little example here, what I've shown you here, is B and not A. So hopefully that is clear, and hopefully that is not the first time you've ever seen anything like that. so, set theory offers us many Logical consequences of the mathematical formulation of what I just explained. And there are some laws of set operations, which are fairly basic, which you've probably seen before as well. The Commutative Law, the Associative Law and the Distributive Law, the Distributive Law, excuse me. And these things at the bottom, called DeMorgan's Laws. Which essentially says that the complement of the union of a bunch of disjoint sets is equal to the intersection of the complements. And similarly, the other way around. These are one of those things that we may or may not end up using later but it has to be said or I would probably be run out of the business. Okay. So there are three axioms of probability. Only three things that we have to assume. Probabilities are numbers between zero and one inclusive. The probability of the entire sample space is one meaning something must happen on your trial. And Axiom three says that, for any sequence of disjoint events, in other words, events which have an intersection that's empty, they do not overlap, in terms of those Venn diagrams. The probability of the union of those events, is equal to the sum of their individual probabilities. All other rule of probability can be derived from these three. So we could stop there and if there's anything else you want to know about probability you could pull out your pencil and paper and figure it out but you probably be advised to go to a book, someones already done it. Here's a few of the highlights of the things that you could prove from those three axioms, but the probability of [INAUDIBLE] is simply one minus the probability of A. If A is a subset of B. Meaning in terms of those Venn diagrams, if the event A was the red circle, and it lay entirely within the bounds of the red circle that defined B. Then the probability of A would be less than or equal to the probability of B. And, in general, if the probabil, the probability of A union B, that's A or B. Is the probability of a plus the probability of b minus the intersection, minus the probability of the intersection. And that is a little bit of a generalization of the rule about disjoint events. Here what we're saying is A, if A and B are not disjoint, I have to remember to remove one copy. Of the intersection, otherwise I would have double-counted the intersection. In the fourth point there, you see the generalization of the, of the A union B rule when we have many events, not just two. And I will leave you to go and look that up in a book. If you try to work out that little example with A just A1 and A2, you should get the result on the third line. And we will move on now to defining joint and conditional probabilities. The joint probabilities of events A and B we already talked about that, that's the intersection. That's elements of the sample space that belong to both events at the same time. The conditional probability of A given that B has occurred. Is simply the probability of A intersection B divided by the probability of B. And the way to visualize that, I think, with this diagram is to say well that's, you might think of that as the area of the intersection, but not now normalized by the area of the entire box labelled S, which would be one, but simply normalized by the area of the blue circle. The unconditional probability of A, is the intersection of the A, with S, divided by the probability of S. Since the probability of S is one, we, the unconditional probability of A is just A intersectional with a sample space, itself. And. The usual definition of statistical independence is that the probability of A intersection B is equal to the product of the probabilities of A and B. Now, I would take this moment to, to make a comment. I often encounter with, particularly with my physics friends up at the lab, different uses of the word independent. Sometimes the word independence is taken in an English language sense/ and whenever we use the word independence in this set of lectures, and in generally, when you're talking about probability, you mean this very precise definition of independence. So that's something important to keep in mind, if you are a Physics friend or you're talking to one or to anyone else for that matter. finally, we can see that another equivalent definition of statistical independence is that the conditional probability of A given B is simply equal to the probability of A. And that has a sort of pleasing interpretation, which is that you learn nothing from knowing that B occurred. So if the probability if A given B is exactly the same as the probability of A without. Anything to do with B then we would say that A and B are independent. Okay, now we are getting into the good stuff. There is something out there called the law of total probability and all this says is that if I have the partitioning of the Sample Space Say my, say I'm going to going to take away, what I was calling event A before, and I'm going to simply divide the sample space into three different events A1, A2 and A3. Then, I can reconstruct the probability of B there, by simply taking the individual probabilities of the intersection of each of those As. Oh, I'm sorry, the conditional probability. Of, B given A for each of the As. And then weighting it by the probability of each of those individual As. [SOUND] And this should make some kind of intuitive sense. And in fact, that is nothing more than the definition of conditional probability that we saw in the earlier slide. now. You've probably all heard of this, Bayes' Rule or Bayes' Theorem. Bayes' Theorem is nothing more than another definition of conditional probability. It simply says that the probability of A given B, can be expressed, well, we know it's definition is the probability of A intersection B, which is on the. An equivalent expression is in the numerator of the central term there. Divided by the probability of B. So, I can leave the numerator alone. And then I can notice that I can use the law of total probability to re express the probability of B. In the form on the, in the denominator of the right side of the equation. And that's Bayes' rule and what's good about Bayes' rule or what's useful about is it allows to express the probability of a A given B in terms of the probability of B given A and sometimes we know more about the probability of B given A than we do about the probability of A given B and here's probably the one example that I'll take. From things I do every day which would be suppose A is the event that the true CO2 concentration in the atmosphere over a column of atmosphere, is greater than 400 parts per million. And let B be the event that the Orbiting Carbon Observatory instrument, which just launched in July. Observes CO2 concentration of 398 parts per million, but what we would really like to know is what the conditional distribution of A given B is in that case. What we would like to know, how likely is it that the true concentration is 400 parts per million if what CO2 observed was 398 parts per million. But we don't have a good way of knowing that, but because we built OCO-2 ourselves and we know how the instrument works and we know how the instrument behaves under different physical conditions, we know much more about the probability that OCO-2 would observe 398 parts per million. If the true concentration in the atmospheric column was 400 parts per million. So that's a case where we can use that, and in fact is used to retrieve OCO2, CO2 concentration measurements that base therom is how they do their so called retrieval, their measurement of CO2 in the atmosphere, so that may be the little. Cocktail party tidbit. Okay, so at the end of each of these sections. I'm going to put down some references books typically that I'll, I'll be embarrassed to say in most cases the books I learned these things from. Or, or, later editions of them, so that if you wanted to go and look these things. There are many places you can go, of course, to read about probability. A First Cour in probability by Sheldon Ross is a good, compact, discussion of these items and many, many more. And the classic. Is William Feller's Volume I and II Introduction to Probability Theory and its Applications which is quite old by now. But fortunately probability doesn't change very much, over time. The Volume I is devoted exclusively to discreet probabilities, which we'll get to in a minute. Volume two is, devoted to, continuous probabilities, and they're both excellent and they have great problems in the back if you like doing that sort of thing. So in the module now, we'll discuss how these rules of probability will translate for settings in which we model numerical phenomenon. In other words data.