So we talked about that the concept of dependence, right? So X and Y are dependent upon one another and if I tell you about X that gives you some information, additional information about what Y is. Now, there are many, many, many different kinds of dependencies that can exist in data. Right? And the concept of covariance and correlation measures only one special kind of dependency, and that's called linear dependency. All right so covariance and correlation measure the extent to which X and Y move together in a straight line fashion. Okay. If there is a positive relationship then when you plot when X versus Y you see an upward sloping linear relationship. If there's negative covariance correlation and if you plot X versus Y then there's a downward sloping linear relationship between them. It is very important to understand that covariance and correlation only measures linear association. You know, but linear association is just one kind. You could have a non-linear association, that is Y didn't have to be related in a quadratic form, or you know, a logarithmic form, or an exponential form, or, or whatever. Turns out that it's very difficult to define measures of non-linear dependence, okay? There's just, there, you know, not because there are so many ways of defining non-linear dependence there isn't a general way to do it, and so when, we look at dependents with random variables, we focus on linear dependents only. Now, this may seem like it's, it's highly restrictive, but you know, one of the facts of the world is, a lot of relationships are approximately linear over, you know, generally, you know, viewed ranges of values. So even something that's nonlinear is approximately linear over, you know, say, small intervals, or something like that. So. Linear dependence could be a good approximation for even non-linear dependence in, in many situations and when we start looking at data and we, when we plot, you know, different values of X and Y together, we'll often see that the dependence that comes from the data looks pretty linear and so we don't actually have to go to something that's more complicated. Alright, so what is covariance? So, covariance is a measure of direction but not strength of the linear relationship in the data, okay? So it's, so covariance, so we use the Greek symbol sigma with the subscript x, y to represent the covariance between x and y. The definition of covariance is the expectation of x minus its mean multiplied by y minus its mean. And so if we have a discreet random variable, we would take all values of X and Y in the joint sample space and then we would take the product of X minus its mean, Y minus its mean, weight by the probabilities, and add it all up. If we have a continuous distribution, we would take from the summation into an integral. We would take X minus the mean, Y minus the mean, weighted by the probability curve and, and then compute that value. Alright? Now, when you look at the formula it's not necessarily apparent that this is a measure of linear association, right? And so what I want to do now is just show you a graph and, and show graphically why this computation gives you a, a measure of the direction of linear dependance. So. Let's see. Here we go. So you're going to do a graphical description of covariance. So I'm going to do a graph. I'm going to take Y minus the mean of Y, I'm going to plot on this axis, and I'm going to plot X minus the mean on this axis. And what I'm going to do is, I'm going to plot what I call a probability scatter plot. So I'm going to look at you know, say like a discrete distribution between X and Y, where every single point in that distribution is equally likely. Okay? And so I'll give a, a distribution that's going to look like this. Okay. So looking at the probability scatter plot, you know, I mean, what does your intuition tell you about the direction of linear dependence? Is it positive or negative? It's positive. And why is it positive? Yeah, so it's upward sloping. So when we look at values of X above its mean, so when X minus mu is positive we're up over here. Right? So in this quadrant we have values of X above its mean and then notice that when X is above its mean, in, this direction, Y also is above its mean. So as X increases, Y increases. And then if we look down here, also as well when X is below its mean then Y tends to be below its mean as well so that's just showing you that X and Y move together in a positive way. Now the so how does this give, rise to a positive covariance and remember covariance is the expectation of X minus the mean of X times Y minus the mean of Y? And, in this case, this would be a probability weighted average when we multiplied by the probabilities between X and Y. So let's just break up this probability scatter plot into four quadrants, so we'll have quadrant one, quadrant two, quadrant three, and quadrant four. Now in quadrant one, X is, tends to be below its mean, but Y tends to be above its mean. Right? So, in this quadrant we have X minus mu of X is less than zero, but Y minus mu of Y is greater than zero. So we take the product of X minus mu of X, and Y minus mu of Y, we get a negative number times a positive number. And so we get a negative number. So if we see points in this quadrant, and we want to know what is contribution to covariance of points in this quadrant, it's equal to X minus the mean times Y minus the mean, weighted by the probability, probabilities are always positive. So the contributions to covariance in this quadrant are all negative values, okay? Now notice in this graph, there are no contributions to covariance in this quadrant. Now let's look at quadrant number two. In quadrant number two, X is above it's mean. And Y is above it's mean. So you get, X minus the mean of X is positive, Y minus the mean of Y is positive, so when you take the product of X minus its mean and Y minus its mean. You get a positive times a positive, so that's a positive number. So these three blue dots here are equal to X minus the mean, Y minus the mean. They're positive numbers, they're weighted by a probability that's positive. So all these three dots here give a positive contribution to covariance.Okay? And then you could do the last for the quadrant number three. So X minus the mean is less than zero. Y minus the mean is less than zero. So when you take the product of X minus the mean, and Y minus the mean, you get a negative times a negative number, which is positive. So in this quadrant down here, these three blue dots give a positive contribution to covariance. Because if you have multiplied two negative numbers. Then finally in quadrant four, we have X minus the mean of X is a positive, but Y minus the mean of Y is a negative. So we take the product, we get a negative number. So that's a positive times a negative, so that's less than zero. So, any dots in this range, give a negative contribution to covariance. So what is covariance? It's the expected value of the product, so its the sum of the product weighted by the probabilities, so I take these three points, they're all positive, multiply by positive numbers, its all positive. There's no contribution here, no contribution here, positive contribution here, so we add everything up, we get a positive co-variance. So this is a case where sigma X Y is greater than zero. Okay? Now covariance only gives you the direction of linear dependence. Right? So if a covariance is positive we know X and Y move together in a positive relationship. Covariance doesn't tell you the strength of the relationship. Okay. And that's where correlation comes into play. Correlation tells you the direction and strength. So correlation is defined to be the covariance divided by the product of the standard deviations. And it's just the scale value of the covariance. So we now know what covariance is. Covariance measures the direction of linear association between two random variables. So, how do we compute it? So if we have a discrete distribution we know that the covariance between X and Y, you just take the X values minus the mean, Y values minus the mean, weight it by the probabilities, add them all up. So you got a discrete distribution, things are, it's a brute-force calculation. So with discreet distributions I said, computation is brute-force and, and, and not fun. So here I have a, an example of computing covariance. So, if you have a discreet distribution, there's no real easy way to end up doing the computation. So in a spreadsheet, you can take the values of X and the values of Y. And the joint probabilities. And then you have to compute X minus the mean and Y minus the mean, which you can do in a table. And then you can take the product of X minus the mean and Y minus the mean, and then weight that by the probability. And then add them all up at the end of the day. So this is just doing the brute-force calculation. The sum of X minus the mean, Y minus the mean, weighted by the probability. So there, there's no easy way to do the calculation, it, it's just brute-force. Now in the, for my discrete distribution here, my example distribution one way of viewing relationships between the random variables X and Y, is to do the probability scatter plot. So this probability scatter plot, I have the values of Y here, and the values of X here, and these bubbles here represent the probabilities associated with a joint occurrence between X and Y. And the size of the bubble is related to the magnitude of the probability. So we see that there's a bigger joint probability for this point and this point than for these other points. And you can sort of see that there's a upward sloping relationship in this probability scatter plot. And that's indicating that it looks like there's a positive relationship between X and Y. So if there's any justice in the world, if we actually compute the covariance, we should get a positive number. And so, this is what happens. We compute the covariance between X and Y and it turns out to be 0.25. Okay? So the covariance between X and Y is a positive number and that means that as X goes up Y tends to go up as well. And so that's just this number here which means that it computes the covariance and do the group work calculation. It's a quarter 0.25. Alright. Now, there are some important properties of covariance that are very useful and we're going to be using them a lot in our modeling. So that the first property is covariance is a symmetric relationship, so if you look at the covariance between X and Y, that's the same as the covariance between Y and X. Right? So, covariance is a symmetric relationship and you can get that just from the, the definition. Now, another result if we take X we multiply it by some scalar a. We take Y and multiply it by some scalar b so we have say 2X and 5Y.< /i>< /i> The covariance of this scaled value of X and Y is equal to the scalars ab< /i> times the covariance between X and Y. Okay, now this result comes from the, just the definition of covariance so if you look at the covariance between aX and bY, that's the expectation of aX minus a times the mean of X times bY - b times the mean of Y, right? And so that's just the definition of covariance between aX and bY. And so then this is equal to the expectation of aX minus its mean and bY minus its mean. And a and b are just numbers, so they come out of the expectation. So then this is just A times B times the covariance between X and Y. So, if you take a linear function, so if you just multiply, if you change the scale of X and Y, then the covariance changes by these values. Now this is very important because that partially tells you that I can make the value of the covariance between X and Y anything that I want just by multiplying X and Y by numbers a and b. Or, in other words, this is, this relationship tells us that the magnitude of the covariance depends upon the units of measurement between X and Y. Right? So say X is return and Y is return right? So we get equal variance between X and Y, what is the units of covariance? Well it's the product of returns, it's return squared because I'm taking X minus its mean and multiplying it by Y minus its mean. Now if I multiply returns by a 100, so I have returns in percentage. So I, if a is a hundred and b is a 100, then the covariance between X and Y becomes 100100 times the< /i> covariance between X and Y. So I, I'll inflate the magnitude of the covariance a lot, but I don't change the relationship between X and Y when I change the scale, right? If I multiple X and Y by ten, both by ten, that doesn't change the relationship that just changes the scale that we measure things on. So this relationship tell us that, very often we, we don't care about the magnitude of covariance by itself, we just care weather or not it's a positive or negative number. The coverage between X and itself, will have its X pull varied with itself but that is just the variants of X, okay? If X and Y are independent there is no relationship between X and Y at all. So, there is no linear relationship so the covariance is equal to zero. Now the converse is not true. If the covariance between X and Y is zero, we just know that there is no linear relationship between X and Y. But X and Y can have a non-linear relationship, alright? So knowledge that the covariance is zero does not mean that X and Y are independent. Now, the final result here is just a convenient way of computing covariance. So it turns out that the covariance between X and Y could be also, could be computed as the expectation of X times Y minus the expectation of X times the expectation of Y. And, and this result just follows from you know, brute force calculations. So if you want to calculate the covariance between X and Y, the expectation of X minus its mean, Y minus its mean and then you just multiply everything out. We get XY - X times the mean of Y - Y times the mean of X, plus the mean of X times the mean of Y. So this becomes expectation of XY minus mu of Y times the expectation of X minus the expectation of Y times the mean of X + the mean of X times the mean of Y. So this is the expectation of XY - UY times mu X Minus mu Y times mu X, plus mu X times mu Y. So this is the expectation of XY minus the mean of X times the mean of Y. So, by the definition of covariance, and the linearity properties of expectation, we can get that covariance can also be computed as the expectation of the product of X times Y, minus the mean of X times the mean of Y. Now, this is often a more convenient way to do the computation, you know, in