So we talked about that the concept of
dependence, right?
So X and Y are dependent upon one another
and if I tell you about X that gives you
some information, additional information
about what Y is.
Now, there are many, many, many different
kinds of dependencies that can exist in
data.
Right?
And the concept of covariance and
correlation measures only one special kind
of dependency, and that's called linear
dependency.
All right so covariance and correlation
measure the extent to which X and Y move
together in a straight line fashion.
Okay.
If there is a positive relationship then
when you plot when X versus Y you see an
upward sloping linear relationship.
If there's negative covariance correlation
and if you plot X versus Y then there's a
downward sloping linear relationship
between them.
It is very important to understand that
covariance and correlation only measures
linear association.
You know, but linear association is just
one kind.
You could have a non-linear association,
that is Y didn't have to be related in a
quadratic form, or you know, a logarithmic
form, or an exponential form, or, or
whatever.
Turns out that it's very difficult to
define measures of non-linear dependence,
okay?
There's just, there, you know, not because
there are so many ways of defining
non-linear dependence there isn't a
general way to do it, and so when, we look
at dependents with random variables, we
focus on linear dependents only.
Now, this may seem like it's, it's highly
restrictive, but you know, one of the
facts of the world is, a lot of
relationships are approximately linear
over, you know, generally, you know,
viewed ranges of values.
So even something that's nonlinear is
approximately linear over, you know, say,
small intervals, or something like that.
So.
Linear dependence could be a good
approximation for even non-linear
dependence in, in many situations and when
we start looking at data and we, when we
plot, you know, different values of X and
Y together, we'll often see that the
dependence that comes from the data looks
pretty linear and so we don't actually
have to go to something that's more
complicated.
Alright, so what is covariance?
So, covariance is a measure of direction
but not strength of the linear
relationship in the data, okay?
So it's, so covariance, so we use the
Greek symbol sigma with the subscript x, y
to represent the covariance between x and
y.
The definition of covariance is the
expectation of x minus its mean multiplied
by y minus its mean.
And so if we have a discreet random
variable, we would take all values of X
and Y in the joint sample space and then
we would take the product of X minus its
mean, Y minus its mean, weight by the
probabilities, and add it all up. If we
have a continuous distribution, we would
take from the summation into an integral.
We would take X minus the mean, Y minus
the mean, weighted by the probability
curve and, and then compute that value.
Alright?
Now, when you look at the formula it's not
necessarily apparent that this is a
measure of linear association, right?
And so what I want to do now is just show
you a graph and, and show graphically why
this computation gives you a, a measure of
the direction of linear dependance.
So.
Let's see.
Here we go.
So you're going to do a graphical
description of covariance.
So I'm going to do a graph.
I'm going to take Y minus the mean of Y,
I'm going to plot on this axis, and I'm
going to plot X minus the mean on this
axis.
And what I'm going to do is, I'm going to
plot what I call a probability scatter
plot.
So I'm going to look at you know, say like
a discrete distribution between X and Y,
where every single point in that
distribution is equally likely.
Okay?
And so I'll give a, a distribution that's
going to look like this.
Okay.
So looking at the probability scatter
plot, you know, I mean, what does your
intuition tell you about the direction of
linear dependence?
Is it positive or negative?
It's positive.
And why is it positive?
Yeah, so it's upward sloping.
So when we look at values of X above its
mean, so when X minus mu is positive we're
up over here.
Right?
So in this quadrant we have values of X
above its mean and then notice that when X
is above its mean, in, this direction, Y
also is above its mean.
So as X increases, Y increases.
And then if we look down here, also as
well when X is below its mean then Y tends
to be below its mean as well so that's
just showing you that X and Y move
together in a positive way.
Now the so how does this give, rise to a
positive covariance and remember
covariance is the expectation of X minus
the mean of X times Y minus the mean of Y?
And, in this case, this would be a
probability weighted average when we
multiplied by the probabilities between X
and Y.
So let's just break up this probability
scatter plot into four quadrants, so we'll
have quadrant one, quadrant two, quadrant
three, and quadrant four.
Now in quadrant one, X is, tends to be
below its mean, but Y tends to be above
its mean.
Right?
So, in this quadrant we have X minus mu of
X is less than zero, but Y minus mu of Y
is greater than zero.
So we take the product of X minus mu of X,
and Y minus mu of Y, we get a negative
number times a positive number.
And so we get a negative number.
So if we see points in this quadrant, and
we want to know what is contribution to
covariance of points in this quadrant,
it's equal to X minus the mean times Y
minus the mean, weighted by the
probability, probabilities are always
positive.
So the contributions to covariance in this
quadrant are all negative values, okay?
Now notice in this graph, there are no
contributions to covariance in this
quadrant.
Now let's look at quadrant number two.
In quadrant number two, X is above it's
mean.
And Y is above it's mean.
So you get, X minus the mean of X is
positive, Y minus the mean of Y is
positive, so when you take the product of
X minus its mean and Y minus its mean.
You get a positive times a positive, so
that's a positive number.
So these three blue dots here are equal to
X minus the mean, Y minus the mean.
They're positive numbers, they're weighted
by a probability that's positive.
So all these three dots here give a
positive contribution to covariance.Okay?
And then you could do the last for the
quadrant number three.
So X minus the mean is less than zero.
Y minus the mean is less than zero.
So when you take the product of X minus
the mean, and Y minus the mean, you get a
negative times a negative number, which is
positive.
So in this quadrant down here, these three
blue dots give a positive contribution to
covariance.
Because if you have multiplied two
negative numbers.
Then finally in quadrant four, we have X
minus the mean of X is a positive, but Y
minus the mean of Y is a negative.
So we take the product, we get a negative
number.
So that's a positive times a negative, so
that's less than zero.
So, any dots in this range, give a
negative contribution to covariance.
So what is covariance?
It's the expected value of the product, so
its the sum of the product weighted by the
probabilities, so I take these three
points, they're all positive, multiply by
positive numbers, its all positive.
There's no contribution here, no
contribution here, positive contribution
here, so we add everything up, we get a
positive co-variance.
So this is a case where sigma X Y is
greater than zero.
Okay?
Now covariance only gives you the
direction of linear dependence.
Right?
So if a covariance is positive we know X
and Y move together in a positive
relationship.
Covariance doesn't tell you the strength
of the relationship.
Okay.
And that's where correlation comes into
play.
Correlation tells you the direction and
strength.
So correlation is defined to be the
covariance divided by the product of the
standard deviations.
And it's just the scale value of the
covariance.
So we now know what covariance is.
Covariance measures the direction of
linear association between two random
variables.
So, how do we compute it?
So if we have a discrete distribution we
know that the covariance between X and Y,
you just take the X values minus the mean,
Y values minus the mean, weight it by the
probabilities, add them all up.
So you got a discrete distribution, things
are, it's a brute-force calculation.
So with discreet distributions I said,
computation is brute-force and, and, and
not fun.
So here I have a, an example of computing
covariance.
So, if you have a discreet distribution,
there's no real easy way to end up doing
the computation.
So in a spreadsheet, you can take the
values of X and the values of Y.
And the joint probabilities.
And then you have to compute X minus the
mean and Y minus the mean, which you can
do in a table.
And then you can take the product of X
minus the mean and Y minus the mean, and
then weight that by the probability.
And then add them all up at the end of the
day.
So this is just doing the brute-force
calculation.
The sum of X minus the mean, Y minus the
mean, weighted by the probability.
So there, there's no easy way to do the
calculation, it, it's just brute-force.
Now in the, for my discrete distribution
here, my example distribution one way of
viewing relationships between the random
variables X and Y, is to do the
probability scatter plot.
So this probability scatter plot, I have
the values of Y here, and the values of X
here, and these bubbles here represent the
probabilities associated with a joint
occurrence between X and Y.
And the size of the bubble is related to
the magnitude of the probability.
So we see that there's a bigger joint
probability for this point and this point
than for these other points.
And you can sort of see that there's a
upward sloping relationship in this
probability scatter plot.
And that's indicating that it looks like
there's a positive relationship between X
and Y.
So if there's any justice in the world, if
we actually compute the covariance, we
should get a positive number.
And so, this is what happens.
We compute the covariance between X and Y
and it turns out to be 0.25.
Okay?
So the covariance between X and Y is a
positive number and that means that as X
goes up Y tends to go up as well.
And so that's just this number here which
means that it computes the covariance and
do the group work calculation.
It's a quarter 0.25.
Alright.
Now, there are some important properties
of covariance that are very useful and
we're going to be using them a lot in our
modeling.
So that the first property is covariance
is a symmetric relationship, so if you
look at the covariance between X and Y,
that's the same as the covariance between
Y and X.
Right?
So, covariance is a symmetric relationship
and you can get that just from the, the
definition.
Now, another result if we take X we
multiply it by some scalar a.
We take Y and multiply it by some scalar b
so we have say 2X and 5Y.< /i>< /i>
The covariance of this scaled value of X
and Y is equal to the scalars ab< /i>
times the covariance between X and Y.
Okay, now this result comes from the, just
the definition of covariance so if you
look at the covariance between aX and bY,
that's the expectation of aX minus a times
the mean of X times bY - b times the mean
of Y, right?
And so that's just the definition of
covariance between aX and bY.
And so then this is equal to the
expectation of aX minus its mean and bY
minus its mean.
And a and b are just numbers, so they come
out of the expectation.
So then this is just A times B times the
covariance between X and Y.
So, if you take a linear function, so if
you just multiply, if you change the scale
of X and Y, then the covariance changes by
these values.
Now this is very important because that
partially tells you that I can make the
value of the covariance between X and Y
anything that I want just by multiplying X
and Y by numbers a and b.
Or, in other words, this is, this
relationship tells us that the magnitude
of the covariance depends upon the units
of measurement between X and Y.
Right?
So say X is return and Y is return right?
So we get equal variance between X and Y,
what is the units of covariance?
Well it's the product of returns, it's
return squared because I'm taking X minus
its mean and multiplying it by Y minus its
mean.
Now if I multiply returns by a 100, so I
have returns in percentage.
So I, if a is a hundred and b is a 100,
then the covariance between X and Y
becomes 100100 times the< /i>
covariance between X and Y.
So I, I'll inflate the magnitude of the
covariance a lot, but I don't change the
relationship between X and Y when I change
the scale, right?
If I multiple X and Y by ten, both by ten,
that doesn't change the relationship that
just changes the scale that we measure
things on.
So this relationship tell us that, very
often we, we don't care about the
magnitude of covariance by itself, we just
care weather or not it's a positive or
negative number.
The coverage between X and itself, will
have its X pull varied with itself but
that is just the variants of X, okay?
If X and Y are independent there is no
relationship between X and Y at all.
So, there is no linear relationship so the
covariance is equal to zero.
Now the converse is not true.
If the covariance between X and Y is zero,
we just know that there is no linear
relationship between X and Y.
But X and Y can have a non-linear
relationship, alright?
So knowledge that the covariance is zero
does not mean that X and Y are
independent.
Now, the final result here is just a
convenient way of computing covariance.
So it turns out that the covariance
between X and Y could be also, could be
computed as the expectation of X times Y
minus the expectation of X times the
expectation of Y.
And, and this result just follows from you
know, brute force calculations.
So if you want to calculate the covariance
between X and Y, the expectation of X
minus its mean, Y minus its mean and then
you just multiply everything out.
We get XY - X times the mean of Y - Y
times the mean of X, plus the mean of X
times the mean of Y.
So this becomes expectation of XY minus mu
of Y times the expectation of X minus the
expectation of Y times the mean of X + the
mean of X times the mean of Y.
So this is the expectation of XY - UY
times mu X Minus mu Y times mu X, plus mu
X times mu Y.
So this is the expectation of XY minus the
mean of X times the mean of Y.
So, by the definition of covariance, and
the linearity properties of expectation,
we can get that covariance can also be
computed as the expectation of the product
of X times Y, minus the mean of X times
the mean of Y.
Now, this is often a more convenient way
to do the computation, you know, in