Hi, welcome tot he second module on Basic Probability for the JPL of Caltech Virtual Summer School on Big Data Analytics. In this section, we are going to, introduce a mathematical formalism for coding and describing the outcome of uncertain phenomena. We will talk about random variables. Distributions, densities, and mass functions and expectation, or expected values. This is all building on the module, which is the prior module which was part one of a review of basic probability. Okay, so a random variable is a numerical coding of the out come of a trial, or set of trials. Simple example, I toss a coin and I let X, I give X the value one if the coin comes up heads, and I give it zero if it comes up tails. So now instead of identifying the outcome of my trial or my observations with h's and t's as I might have done previously I'm now going to use ones and zeros. Random variables can be discreet, taking on at most a countable number of values, or continuous, taking on a continuous, taking on values in a continuous range. An example of a discreet random variable might be the number of times I say hello today. An example of a continuous random variable. Might be the height of the next person that I meet. Now, notation is, going to be really important for the rest of these lectures, so I'm going to dwell on it for an entire slide here. I have a friend who's fond of saying, that if you have good notation you can actually learn things from the notation and I personally have experienced that and it's quite. Quite amazing when it happens. If you stick to your notation, and you see something that doesn't look right, and you've actually done the notation correctly, you might have discovered something. So, we're going let random variables or scalars, random variables be denoted by capital letters like big X there. And ordinary variables, Which are the kind of variables we learned about it Algebra class in high school take on fixed but possibly arbitrary values will let those be denoted by lower case letters. So for example, you probably remember the formula for the equation of a line y equals mx plus b. We talked about x and y being the variables and the point there was that x and y could be. Any numbers that obey that certain rule they are arbitrary but fixed. Random variables, on the other hand, are variables which we think of as behaving according to a probability distribution. In other words, they're not fixed. so, when we talk about the event capital X equals little x. We say that X is a realization, of capital X, one thing, one of, a number of possible outcomes actually occurred when it was realized. We could also talk about random vectors as opposed to random variables which I indicated would be scalars, random vectors would be a collection of random variables represent, representing a point that you could think of as a high dimensional space. And we'll denote those by a bold bold, bold text. So we can have a capital bold X, and we can also have a lowercase bold x. An example of a discreet random variable, the number of times I say hello. An example of a continuous random variable, the height of the next person I meet. And I think that probably duplicates what I had on the previous slide. Anyway. Okay, so, now instead of talking about probabilities and looking at Venn diagrams, it's convenient to have mathematical functions to describe these things. The behavior of a random variable, instead of my writing down, expressions for P, a behavior of a random variable might be well described by something called its cumulative distribution function. Which is nothing, other than the probability that random variable capital X takes on a value less than or equal to little x. So that's just a function, and I call that capital F with a subscript, big X, and the argument to that function is little x. The function PX equals x is called the probability mass function if x is a discreet random variable. And if it's a, continuous random variable, we call it a probability density function. So people, tend to use PDF to represent either of these things, but I'm a bit of a stickler so I want to be sure that we call a PDF a PDF and a PMF a PMF. But in both cases, they are the probability that. Random variable capl, capital X in this instance takes on a value less than or equal to A as shown, one's an integral one's a sum. Let us also point out that in the low in the in the bottom equation there I've defined a new term little f of. Subscript capital X of X, and that is just the derivative of, capital F in the case of the continuous random variable. Right. So here are two examples, what I've shown here on the left is the probability mass function of a discreet random variable. And on the right is its cumulative distribution function. And you can see that the cumulative distribution function goes up in steps because we sort of lurch from one value to the next. The transition is not nice and smooth. Here are the PDF and CDF of a continuous random variable. I've shown a normal random variable here, a normal PDF. And the normal CDF on the right and you can see that one is smooth. Okay, so these, definitions all generalized and very straight forward ways to higher dimensions and I'm going to show you a few examples that involving two dimensions because I can draw those, I cannot visualize or make nice, pretty pictures. Of PEFs or CEFs or much else for that matter in higher dimensions then that but all the math carries through which is of course the benefit of having math. So here on the left I'm showing you the PEF of a bivariate random vector, bold capital X which has components X1 an X2. And you're probably thinking that looks like a bivariate Gaussian, and it is, because that's how I generate it. But there it is, the height of that function gives you, the value of the PDF at the combination of X1 and X2 in the domain on the floor of the plot. And the graph on the right is the cumulative distribution function. That goes with that PDF which shows how you pick up mass as you roam around in that space down on the floor on the plot. okay, so let's define something new let's go back to that what was on the left in the earlier, in the previous graphic I have my Gaussian. Bump there and that is a joint distribution. It's telling me something about the joint behavior of X1 and X2 and I'm going to define the marginal distributions of X1 and X2 now the marginal distribution of X1 would be obtained by. Let's say, standing over on the right where I have written f sub X1 comma sub X2, open paren, X1 comma X2, close paren. If I stood over there and I looked straight at my Gaussian, bump and I imagined taking a bulldozer and pushing all that mass, all the way over. On to the axis that I've labeled X1. Then I would have, the marginal density of random variable X1. And if I did the same thing in the other direction, then I would have the marginal density of random variable X2. Essentially what this means is that, we're simply integrating over, the unwanted variable. To obtain the distribution of the wanted variable, and this, simply, if you think about it is just another application of the law of total probability. Now, conditional densities we already started to talk about how useful conditional probabilities are. Conditional density does something a little bit different. It says. I'm going to slice, that joint PDF fx1, X2. At specific values of random variable X2. And, if I take those slices, what those are, are joint distributions at fixed values of random variable X2, which I've denoted here, let's say we're going to look at that one in the middle, which is identified by C2 on the X2 axis. So the conditional density were almost there, it would be that slice, but the area under that slice is not one, because the area under the whole bump, is one so the area under a single slice cannot be one so we simply have to renormalize, the values in that function. By the, total area under that curve and then that becomes a conditional, density function. And that's how we define the conditional distribution of X1 given X2, for some particular fixed value of X2, which I've called C2. Okay, now, now we're, going to have fun. Supposing I have a random variable X, and I know it's distribution function, I want to know what the distribution function is of a random variable Y, that is simply a transformation of random variable X. And that can be any transformation. And, here's how I would do it. This is one of the fun, little things you'll find in the Ross book. Which is simply a, logical chain from left to right across the line. The equations in the, line on the first bullet point, which tells us how we could find the. CDF of random variable y, knowing what the transformation g is, and what the CDF of the random variable X is. And of course, g has to be invertible here. For this to work. But all it amounts to doing, is finding all the values in the X domain that correspond to a given value in the Y domain. And, collecting them up together and then assigning probability equal to the sum of those probabilities. If X and Y are continuous there's one little, extra little thing you have to do there and that's multiply by the determinate of the derivative of that inverse transformation. and, that is how we would get the. The joint distribution, I'm sorry that's how we would get the distribution of Y as a function X. The joint distribution of X and Y, I've written that purposely down below there because I'm going to need that on a, on a couple of slides from now to show you something. That distribution there is simply the, same thing as the definition of conditional of probability that we solved with the Ps earlier. Okay. So here's my cartoon of what I said about collecting up all the values of X that corresponds with the same value of Y. If the transformation was g. g, y equals g of X, and g of X is X squared. Then this is all we'd be doing. We'd be saying x has a PDF that looks like the thing on the back wall there, on the left. And I'm going to, look, I'm going to just sort of walk across that axis. And, for each value on that axis, I will find its value x squared. And then I'll take the mass and I'll shove it over there, to the wall on the right and I will collect those up. And because Y equals X squared has the same Y value for both the, positive and negative value of X that have the same absolute value, I'm going to get double the mass over there on the right. And all the values. Of the random variable y have to be positive. So that's how that works. Now let's talk about random vectors. We just looked at the PDF of a function of a random variable, everything is still true, if we're talking about random vectors, and here's an illustration of what happens if your random vector is a bivariate random vector. Things are just a little bit more complicated because now we might have two transformations, g1 and g2. And in order for, as before, in order for this to work, things have to be invertible, and certain conditions have to be met, but the formula I've written down at the very bottom of the page looks like the formula that we saw earlier. Only that thing on the right that term J is sometimes called the Jacobian. But it's completely analogous to the one dimensional case. Okay. Bear with me now. Now we're, at the point where we want to talk about an important, function of a, of, of a random variable called its expected value. The expected value of a random variable sometimes we also, call it the mean although we really shouldn't we should call it the expected value when we're talking about probability distributions we can call it the mean when we're talking about things we compute from samples. >> Which we call statistics the expected value of a random variable is a typical value that you might expect it to assume. It's the weighted average of all the potential realizations that random variable can take, where, where the weights are provided by, the probabilities given by the probability density or mass function. So, you can see under the first bullet point on the line that begins with E there, we have a definition of expected value for discrete random variable, and a definition of expected value for a continuous random variable. And in fact, it's often useful to just think of E as an operator and, that depending on whether we are talking about discrete or continuous random variables, you would substitute in what are the sum or the interval. and, all I have done in the next line down is to substitute in, f where we had Ps up above. The expected value of a random vector is simply the vector of expected values of it's components. The expected value of a function of a random variable, this is actually kind of cute. It is simply the weighted average, or the expected the weighted average value of the, derived random variable. But now because, the values of the original X variable correspond to specific values for the derived y variable, which I've called g of X here the formulas generalize in this way. Meaning I can just compute all the different possible values of g of X, and weight them by the appropriate values of X. The expected deviation of X from it's own expected value, which we sometimes call the mean, is called the bias of the random variable X relative to it's mean or relative to it's own expected value. We like to use the Greek letter mu. To denote that and we subscript it by X to make it completely clear that we are talking about random variable X and random variable X's distribution. The variants of a random variable is the expected squared deviation from it's own mean. And there are formulas for it right there. You've probably seen all of that before. Here things will get just slightly tricky. The covariance of two random variables, you've probably seen that before too. It's like a generalization of a variance. The variance of a random vector is a matrix, which is called the variance covariance matrix. And the variance, covariance matrix must be square and symmetric. And it's simply has the variances on the diagonal, and the covariances on the off diagonal, corresponding to the different elements of X. We often use sigma squared as our shorthand for, the variance and sigma Xi, Xj as a shorthand for a covariance. This is the tricky part. The cross covariance between two random vectors that's a different matrix. That's not the same variance, covariance matrix the cross covariance matrix need not be symmetric or square because you can think of the rows of that matrix as corresponding to the components of X and the columns corresponding to the components of Y. But other than that the form looks entirely familiar. Based on what we've already seen. So, there's a bit of tedious equations on this slide, and I don't want to, bore you too much. The important thing here is to say that point number one, and X is a linear operator. So if I take the expectation of the sum of two random variables. It's the, sum of their expectations. And, if I multiply either or both of those random variable by a constant, that constant simply comes out front. And you can, prove this yourself if you want to by using that formula for the expected value of a transformation of a random variable. Now here's an important point. If, X1 and X2 are independent, then the expected value of the product of two random variables based on, those two independent random variables are simply the product of their expectations. And that's very nice and very handy when you're trying to figure out the expected value of something. That represents a product. however, it's really, really, really important to recognize that, that is not true the other way around. It is not true that just because covariance is zero, that, that implies that these random variables are independent. Perhaps I should, back track just slightly, and say that. The expect the product of expectations equaling, the expectation of the product is is essentially defines covariance, there will be the subtraction of the mean terms involved but if you assume everything has zero mean to start with then you don't worry about that, and what we're saying here is that the If the random variables are independent, then, the covariants will be zero. the, only condition, the only case in which the converse is true, namely that, the covariants being zero implies independence, is if you're two, random variables, X1 and X2 are bivariate Gaussian. Now finally, this may be, this is one of the most cool things that there is to say about this topic. And that is to look at the notion of conditional independence, which is I'm sorry the notion of conditional expectation. The conditional expected value of random variable X1 given X2. Is the ordinary definition of expectation, but applied with the conditional distribution of X1 given X2. And if we think of this object as a function of random variable X2, this defines the regression of exponent X1 on X2. And many of you may be thinking, but that doesn't sound like, simple linear regression, or multiple regression that I learned in my textbook at school. Well in fact it is. It's just that what you learn in school, pertains to conditions where X1 and X2 are bivariate Gaussian. And if X1 and X2 are bivariate Gaussian, then if you look at the value of this function, namely the expected value of X1 given X2 is the function of X2, that will lie on a line. And that's, where simple linear regression comes from. But this concept is much more general and could apply in lots of places for example if you applied a clustering algorithm to a data set, and you thought about X2 as being a cluster identifier a number that identifies a cluster and the value of X1 as being a random draw from all the Xs that belong to that cluster. Then you could say that the regression of X1 on X2 is, are the mean functions of each of those clusters, in other words if you made a function that had cluster number on the x axis and the expected value of X1 for each of those clusters separately that would be a regression. So it's a very general term. Now perhaps one of the most useful things that ever existed in probability and statistics. That you may we may not see again in this lecture but it is definitely worth knowing about, it's called the law of iterated conditional expectation. So we defined the expected value of X1 given X2 on the earlier slide. And for clarity here, let's say we have a bivariate distribution, X1 and X2 in the lower left, that's a very, exaggerated bump there. And then I, slice that distribution along. For values defined along X2, C1, C2, C3, and C4. And of course I'm only showing you certain slices out of that distribution. But, I could slice at any location. And if I found the conditional expected value of each of those conditional distributions, and then I averaged them over the possible realizations of X2. I could reconstruct the expected value of X1, and that again, ends up being very very very useful in sort of the same way that Bayes' theorem ends up being useful in that, sometimes it's easier for me to know things conditionally than unconditionally. And this gives me a way to get the unconditional expected value of X1 if I know something, about the conditional behavior of X1 given X2. [SOUND] And just for completeness, we have to talk about the variance as well. And I'm going to, I'm, I'm stating all these results basically for continuous random variables. But there are analogs for the discrete and vector cases as well. The variance of a linear function is not the sum of the variances. it, it, you have to account for the co-variance term, and that comes about because of those Venn diagrams when we were looking at things like probability of a intersection b. And we had to worry about, not double counting the intersection. That intersection ends up being sort of, related to the covariances here and that we have to worry about. So if you see on the first line of the first bullet it says the variance of, two random variables, just look at the thing on the left side of the plus sign and the thing on the right side of the plus sign, is not merely the sum of their two variances, but there is this covariance term, on the right. And by the way, the variance of a constant times a random variable is the, square of the constant times the variance of the random variable, because the variance is a squared thing. Variance of a non linear function, we could, in principle go back ad try to work it out the way we did with the expectations. But there's actually an easier way, because a lot of times that's far too difficult to do in the case of the variance. And we appeal to, a Taylor series expansion. On the function G, if we're talking about, random variable Y, which is a function of random variable X, and you'll see on the right side of the second bullet there. I've simply written out a, first-order tailor expansion about the expected value of X. For the function g of X. And then if I apply the variance formula to the thing on the right side of the approximately equal sign, I can get this approximate relationship that does turn out to be pretty handy. And yeah, here's my point about the covariance and independence. I guess I got to it a little sooner than I meant to. Okay, so analogous to the conditional, conditional expectations, there is a conditional variance formula, which is going to reinforce something that I said a few minutes ago, which is, you might think that what you could do is take these additional distributions and average up their variances. To get the variance of the random that you really care about, which is X1 in this case, but you can't quite do that. If that were true, we wouldn't have the first term on the right side of the equal sign. because if you did that, you would be missing the notion of variability between different values of X2. So, this also ends up being an extremely formula. You can read more about in the, the Ross Book, or in a number of other places. But in, in if you care about the variance of a random variable and you don't know what it is, but you do know something about what it is conditionally, this a nice way to figure it out. Okay, so, I wouldn't blame you if by now you were thinking, why should I care about all of this, it's seems a bit. Often mathland well the reason is because we're going to build models of unknown or uncertain populations, with probability distributions and we want to call these things, let's call these process distributions. And we make inferences about process distributions by computing statistics from samples. The statistics themselves are random variables because they were computed from a sample that was chosen randomly, and they have distributions of their own. We'll call these sampling distributions. The discipline of statistics is largely concerned with understanding the relationship between a process distribution parameter and the sampling distribution of a statistic that is designed to estimate it. So that's where we're headed next. Here's I think it's the same two references that I gave you earlier so I won't repeat that. But now we will move on to two modules on basic concepts of inference.