1
00:00:00,840 --> 00:00:04,480
Hi, welcome tot he second
module on Basic Probability for

2
00:00:04,480 --> 00:00:08,080
the JPL of Caltech Virtual Summer School
on Big Data Analytics.

3
00:00:10,734 --> 00:00:13,963
In this section, we are going to,
introduce a mathematical formalism for

4
00:00:13,963 --> 00:00:16,892
coding and describing the outcome
of uncertain phenomena.

5
00:00:16,892 --> 00:00:19,900
We will talk about random variables.

6
00:00:19,900 --> 00:00:21,240
Distributions, densities,
and mass functions and

7
00:00:21,240 --> 00:00:23,250
expectation, or expected values.

8
00:00:23,250 --> 00:00:25,650
This is all building on the module,

9
00:00:25,650 --> 00:00:28,588
which is the prior module which was part
one of a review of basic probability.

10
00:00:28,588 --> 00:00:32,690
Okay, so

11
00:00:32,690 --> 00:00:38,480
a random variable is a numerical coding of
the out come of a trial, or set of trials.

12
00:00:38,480 --> 00:00:42,900
Simple example, I toss a coin and
I let X, I give X the value one if

13
00:00:42,900 --> 00:00:46,580
the coin comes up heads, and
I give it zero if it comes up tails.

14
00:00:46,580 --> 00:00:50,560
So now instead of identifying
the outcome of my trial or

15
00:00:50,560 --> 00:00:53,810
my observations with h's and

16
00:00:53,810 --> 00:00:57,820
t's as I might have done previously
I'm now going to use ones and zeros.

17
00:00:57,820 --> 00:01:02,350
Random variables can be discreet, taking
on at most a countable number of values,

18
00:01:02,350 --> 00:01:06,830
or continuous, taking on a continuous,
taking on values in a continuous range.

19
00:01:07,860 --> 00:01:10,890
An example of a discreet random
variable might be the number of times I

20
00:01:10,890 --> 00:01:11,740
say hello today.

21
00:01:13,060 --> 00:01:15,250
An example of a continuous
random variable.

22
00:01:16,580 --> 00:01:18,630
Might be the height of
the next person that I meet.

23
00:01:21,430 --> 00:01:26,870
Now, notation is, going to be really
important for the rest of these lectures,

24
00:01:26,870 --> 00:01:29,980
so I'm going to dwell on it for
an entire slide here.

25
00:01:29,980 --> 00:01:33,900
I have a friend who's fond of saying,
that if you have good notation you can

26
00:01:33,900 --> 00:01:37,320
actually learn things from the notation
and I personally have experienced that and

27
00:01:37,320 --> 00:01:38,610
it's quite.

28
00:01:38,610 --> 00:01:39,980
Quite amazing when it happens.

29
00:01:39,980 --> 00:01:43,570
If you stick to your notation, and
you see something that doesn't look right,

30
00:01:43,570 --> 00:01:45,800
and you've actually done
the notation correctly,

31
00:01:45,800 --> 00:01:47,760
you might have discovered something.

32
00:01:47,760 --> 00:01:49,860
So, we're going let random variables or

33
00:01:49,860 --> 00:01:56,130
scalars, random variables be denoted
by capital letters like big X there.

34
00:01:56,130 --> 00:01:59,940
And ordinary variables,
Which are the kind of

35
00:01:59,940 --> 00:02:04,970
variables we learned about it Algebra
class in high school take on fixed but

36
00:02:04,970 --> 00:02:09,060
possibly arbitrary values will let
those be denoted by lower case letters.

37
00:02:09,060 --> 00:02:11,690
So for example,
you probably remember the formula for

38
00:02:11,690 --> 00:02:14,820
the equation of a line y equals mx plus b.

39
00:02:14,820 --> 00:02:16,050
We talked about x and

40
00:02:16,050 --> 00:02:20,450
y being the variables and
the point there was that x and y could be.

41
00:02:20,450 --> 00:02:26,050
Any numbers that obey that certain
rule they are arbitrary but fixed.

42
00:02:27,210 --> 00:02:31,060
Random variables, on the other hand,
are variables which we think of

43
00:02:31,060 --> 00:02:34,010
as behaving according to
a probability distribution.

44
00:02:34,010 --> 00:02:35,120
In other words, they're not fixed.

45
00:02:36,410 --> 00:02:39,700
so, when we talk about the event
capital X equals little x.

46
00:02:39,700 --> 00:02:42,868
We say that X is a realization,
of capital X, one thing,

47
00:02:42,868 --> 00:02:47,825
one of, a number of possible outcomes
actually occurred when it was realized.

48
00:02:47,825 --> 00:02:52,101
We could also talk about random vectors
as opposed to random variables which I

49
00:02:52,101 --> 00:02:53,746
indicated would be scalars,

50
00:02:53,746 --> 00:02:57,629
random vectors would be a collection
of random variables represent,

51
00:02:57,629 --> 00:03:02,500
representing a point that you could
think of as a high dimensional space.

52
00:03:02,500 --> 00:03:07,250
And we'll denote those by a bold bold,
bold text.

53
00:03:07,250 --> 00:03:11,450
So we can have a capital bold X, and
we can also have a lowercase bold x.

54
00:03:12,580 --> 00:03:15,530
An example of a discreet random variable,
the number of times I say hello.

55
00:03:15,530 --> 00:03:17,220
An example of a continuous
random variable,

56
00:03:17,220 --> 00:03:19,410
the height of the next person I meet.

57
00:03:19,410 --> 00:03:22,700
And I think that probably duplicates
what I had on the previous slide.

58
00:03:22,700 --> 00:03:23,775
Anyway.

59
00:03:23,775 --> 00:03:28,030
Okay, so, now instead of
talking about probabilities and

60
00:03:28,030 --> 00:03:29,420
looking at Venn diagrams,

61
00:03:29,420 --> 00:03:33,620
it's convenient to have mathematical
functions to describe these things.

62
00:03:33,620 --> 00:03:36,570
The behavior of a random variable,
instead of my writing down,

63
00:03:36,570 --> 00:03:41,000
expressions for P,
a behavior of a random variable might be

64
00:03:41,000 --> 00:03:44,310
well described by something called
its cumulative distribution function.

65
00:03:44,310 --> 00:03:48,820
Which is nothing, other than the
probability that random variable capital X

66
00:03:48,820 --> 00:03:51,430
takes on a value less than or
equal to little x.

67
00:03:51,430 --> 00:03:56,990
So that's just a function, and I call that
capital F with a subscript, big X, and

68
00:03:56,990 --> 00:03:58,660
the argument to that function is little x.

69
00:03:59,878 --> 00:04:04,490
The function PX equals x is called
the probability mass function

70
00:04:04,490 --> 00:04:06,890
if x is a discreet random variable.

71
00:04:06,890 --> 00:04:11,460
And if it's a, continuous random variable,
we call it a probability density function.

72
00:04:11,460 --> 00:04:15,600
So people, tend to use PDF to represent
either of these things, but I'm a bit of

73
00:04:15,600 --> 00:04:21,300
a stickler so I want to be sure that
we call a PDF a PDF and a PMF a PMF.

74
00:04:22,570 --> 00:04:27,640
But in both cases,
they are the probability that.

75
00:04:27,640 --> 00:04:30,970
Random variable capl, capital X in this
instance takes on a value less than or

76
00:04:30,970 --> 00:04:36,340
equal to A as shown,
one's an integral one's a sum.

77
00:04:36,340 --> 00:04:39,830
Let us also point out that in the low in

78
00:04:39,830 --> 00:04:45,140
the in the bottom equation there
I've defined a new term little f of.

79
00:04:45,140 --> 00:04:49,930
Subscript capital X of X, and
that is just the derivative of,

80
00:04:49,930 --> 00:04:53,065
capital F in the case of
the continuous random variable.

81
00:04:53,065 --> 00:04:55,400
Right.

82
00:04:55,400 --> 00:04:58,090
So here are two examples,
what I've shown here on

83
00:04:58,090 --> 00:05:03,290
the left is the probability mass
function of a discreet random variable.

84
00:05:03,290 --> 00:05:06,150
And on the right is its
cumulative distribution function.

85
00:05:06,150 --> 00:05:09,270
And you can see that the cumulative
distribution function goes up in

86
00:05:09,270 --> 00:05:12,880
steps because we sort of lurch
from one value to the next.

87
00:05:12,880 --> 00:05:14,310
The transition is not nice and smooth.

88
00:05:15,480 --> 00:05:19,200
Here are the PDF and
CDF of a continuous random variable.

89
00:05:19,200 --> 00:05:22,380
I've shown a normal random variable here,
a normal PDF.

90
00:05:22,380 --> 00:05:25,456
And the normal CDF on the right and
you can see that one is smooth.

91
00:05:25,456 --> 00:05:31,420
Okay, so these, definitions

92
00:05:31,420 --> 00:05:34,820
all generalized and very straight
forward ways to higher dimensions and

93
00:05:34,820 --> 00:05:38,730
I'm going to show you a few examples
that involving two dimensions because I

94
00:05:38,730 --> 00:05:43,806
can draw those, I cannot visualize or
make nice, pretty pictures.

95
00:05:43,806 --> 00:05:49,980
Of PEFs or CEFs or much else for that
matter in higher dimensions then that but

96
00:05:49,980 --> 00:05:53,470
all the math carries through which is
of course the benefit of having math.

97
00:05:54,830 --> 00:06:00,310
So here on the left I'm showing you
the PEF of a bivariate random vector,

98
00:06:00,310 --> 00:06:03,190
bold capital X which has
components X1 an X2.

99
00:06:04,290 --> 00:06:07,680
And you're probably thinking that
looks like a bivariate Gaussian, and

100
00:06:07,680 --> 00:06:10,370
it is, because that's how I generate it.

101
00:06:10,370 --> 00:06:14,630
But there it is, the height of that
function gives you, the value of

102
00:06:14,630 --> 00:06:19,800
the PDF at the combination of X1 and
X2 in the domain on the floor of the plot.

103
00:06:19,800 --> 00:06:24,821
And the graph on the right is
the cumulative distribution function.

104
00:06:24,821 --> 00:06:28,460
That goes with that PDF
which shows how you pick up

105
00:06:28,460 --> 00:06:32,100
mass as you roam around in that
space down on the floor on the plot.

106
00:06:35,930 --> 00:06:39,320
okay, so let's define something new let's
go back to that what was on the left in

107
00:06:39,320 --> 00:06:43,830
the earlier, in the previous
graphic I have my Gaussian.

108
00:06:43,830 --> 00:06:47,260
Bump there and
that is a joint distribution.

109
00:06:47,260 --> 00:06:50,520
It's telling me something about
the joint behavior of X1 and X2

110
00:06:52,440 --> 00:06:57,080
and I'm going to define the marginal
distributions of X1 and

111
00:06:57,080 --> 00:07:02,668
X2 now the marginal distribution
of X1 would be obtained by.

112
00:07:02,668 --> 00:07:07,964
Let's say, standing over on
the right where I have written f

113
00:07:07,964 --> 00:07:13,680
sub X1 comma sub X2, open paren,
X1 comma X2, close paren.

114
00:07:13,680 --> 00:07:18,270
If I stood over there and
I looked straight at my Gaussian, bump and

115
00:07:18,270 --> 00:07:24,090
I imagined taking a bulldozer and
pushing all that mass, all the way over.

116
00:07:24,090 --> 00:07:25,940
On to the axis that I've labeled X1.

117
00:07:25,940 --> 00:07:32,480
Then I would have, the marginal
density of random variable X1.

118
00:07:32,480 --> 00:07:34,540
And if I did the same thing
in the other direction,

119
00:07:34,540 --> 00:07:37,400
then I would have the marginal
density of random variable X2.

120
00:07:39,040 --> 00:07:40,550
Essentially what this means is that,

121
00:07:40,550 --> 00:07:44,370
we're simply integrating over,
the unwanted variable.

122
00:07:44,370 --> 00:07:50,490
To obtain the distribution of the wanted
variable, and this, simply, if you

123
00:07:50,490 --> 00:07:53,420
think about it is just another application
of the law of total probability.

124
00:07:55,370 --> 00:07:58,090
Now, conditional densities
we already started to

125
00:07:58,090 --> 00:08:00,910
talk about how useful
conditional probabilities are.

126
00:08:00,910 --> 00:08:03,120
Conditional density does
something a little bit different.

127
00:08:03,120 --> 00:08:04,340
It says.

128
00:08:04,340 --> 00:08:10,700
I'm going to slice,
that joint PDF fx1, X2.

129
00:08:10,700 --> 00:08:14,750
At specific values of random variable X2.

130
00:08:14,750 --> 00:08:20,240
And, if I take those slices,
what those are, are joint distributions

131
00:08:20,240 --> 00:08:24,645
at fixed values of random variable X2,
which I've denoted here, let's say we're

132
00:08:24,645 --> 00:08:29,940
going to look at that one in the middle,
which is identified by C2 on the X2 axis.

133
00:08:29,940 --> 00:08:34,290
So the conditional density were almost
there, it would be that slice, but

134
00:08:34,290 --> 00:08:39,770
the area under that slice is not one,
because the area under the whole bump,

135
00:08:39,770 --> 00:08:43,230
is one so the area under
a single slice cannot be one so

136
00:08:43,230 --> 00:08:46,930
we simply have to renormalize,
the values in that function.

137
00:08:48,200 --> 00:08:51,090
By the, total area under that curve and

138
00:08:51,090 --> 00:08:53,790
then that becomes a conditional,
density function.

139
00:08:54,980 --> 00:08:58,170
And that's how we define the conditional
distribution of X1 given X2,

140
00:08:59,180 --> 00:09:02,690
for some particular fixed value of X2,
which I've called C2.

141
00:09:05,360 --> 00:09:09,100
Okay, now, now we're, going to have fun.

142
00:09:09,100 --> 00:09:14,520
Supposing I have a random variable X,
and I know it's distribution function,

143
00:09:14,520 --> 00:09:18,630
I want to know what the distribution
function is of a random variable Y,

144
00:09:18,630 --> 00:09:21,850
that is simply a transformation
of random variable X.

145
00:09:21,850 --> 00:09:24,130
And that can be any transformation.

146
00:09:24,130 --> 00:09:25,400
And, here's how I would do it.

147
00:09:25,400 --> 00:09:30,510
This is one of the fun, little
things you'll find in the Ross book.

148
00:09:30,510 --> 00:09:36,310
Which is simply a, logical chain
from left to right across the line.

149
00:09:36,310 --> 00:09:37,490
The equations in the,

150
00:09:37,490 --> 00:09:41,060
line on the first bullet point,
which tells us how we could find the.

151
00:09:43,400 --> 00:09:47,830
CDF of random variable y,
knowing what the transformation g is, and

152
00:09:47,830 --> 00:09:51,020
what the CDF of the random variable X is.

153
00:09:51,020 --> 00:09:53,090
And of course,
g has to be invertible here.

154
00:09:55,270 --> 00:09:56,300
For this to work.

155
00:09:56,300 --> 00:10:00,230
But all it amounts to doing,
is finding all the values in

156
00:10:00,230 --> 00:10:03,210
the X domain that correspond to
a given value in the Y domain.

157
00:10:04,470 --> 00:10:06,080
And, collecting them up together and

158
00:10:06,080 --> 00:10:09,390
then assigning probability equal
to the sum of those probabilities.

159
00:10:10,470 --> 00:10:11,300
If X and

160
00:10:11,300 --> 00:10:15,490
Y are continuous there's one little, extra
little thing you have to do there and

161
00:10:15,490 --> 00:10:20,950
that's multiply by the determinate of the
derivative of that inverse transformation.

162
00:10:20,950 --> 00:10:23,250
and, that is how we would get the.

163
00:10:25,340 --> 00:10:26,790
The joint distribution,

164
00:10:26,790 --> 00:10:31,920
I'm sorry that's how we would get
the distribution of Y as a function X.

165
00:10:31,920 --> 00:10:33,770
The joint distribution of X and

166
00:10:33,770 --> 00:10:37,278
Y, I've written that purposely down below
there because I'm going to need that on a,

167
00:10:37,278 --> 00:10:39,420
on a couple of slides from
now to show you something.

168
00:10:40,790 --> 00:10:46,120
That distribution there is simply the,
same thing as the definition of

169
00:10:46,120 --> 00:10:48,400
conditional of probability that
we solved with the Ps earlier.

170
00:10:50,170 --> 00:10:52,146
Okay.
So here's my cartoon of what I said

171
00:10:52,146 --> 00:10:56,720
about collecting up all the values of X
that corresponds with the same value of Y.

172
00:10:56,720 --> 00:10:58,420
If the transformation was g.

173
00:10:59,890 --> 00:11:04,040
g, y equals g of X, and
g of X is X squared.

174
00:11:04,040 --> 00:11:05,010
Then this is all we'd be doing.

175
00:11:05,010 --> 00:11:10,270
We'd be saying x has a PDF that looks
like the thing on the back wall there,

176
00:11:10,270 --> 00:11:11,840
on the left.

177
00:11:11,840 --> 00:11:16,750
And I'm going to, look, I'm going to
just sort of walk across that axis.

178
00:11:16,750 --> 00:11:21,770
And, for each value on that axis,
I will find its value x squared.

179
00:11:21,770 --> 00:11:23,260
And then I'll take the mass and

180
00:11:23,260 --> 00:11:27,130
I'll shove it over there, to the wall on
the right and I will collect those up.

181
00:11:27,130 --> 00:11:30,380
And because Y equals X squared
has the same Y value for

182
00:11:30,380 --> 00:11:36,420
both the, positive and negative value
of X that have the same absolute value,

183
00:11:36,420 --> 00:11:38,610
I'm going to get double the mass
over there on the right.

184
00:11:38,610 --> 00:11:40,000
And all the values.

185
00:11:40,000 --> 00:11:42,600
Of the random variable
y have to be positive.

186
00:11:42,600 --> 00:11:43,800
So that's how that works.

187
00:11:45,570 --> 00:11:47,610
Now let's talk about random vectors.

188
00:11:47,610 --> 00:11:51,740
We just looked at the PDF of a function
of a random variable, everything is

189
00:11:51,740 --> 00:11:55,620
still true, if we're talking about random
vectors, and here's an illustration of

190
00:11:55,620 --> 00:11:59,430
what happens if your random vector
is a bivariate random vector.

191
00:11:59,430 --> 00:12:02,190
Things are just a little bit more
complicated because now we might have

192
00:12:02,190 --> 00:12:04,490
two transformations, g1 and g2.

193
00:12:04,490 --> 00:12:08,360
And in order for, as before, in order for
this to work, things have to be

194
00:12:08,360 --> 00:12:13,110
invertible, and certain conditions have to
be met, but the formula I've written down

195
00:12:13,110 --> 00:12:15,980
at the very bottom of the page looks
like the formula that we saw earlier.

196
00:12:17,370 --> 00:12:21,020
Only that thing on the right that term
J is sometimes called the Jacobian.

197
00:12:22,050 --> 00:12:24,060
But it's completely analogous
to the one dimensional case.

198
00:12:25,780 --> 00:12:26,770
Okay.

199
00:12:26,770 --> 00:12:27,430
Bear with me now.

200
00:12:27,430 --> 00:12:31,800
Now we're, at the point where we
want to talk about an important,

201
00:12:33,450 --> 00:12:37,560
function of a, of, of a random
variable called its expected value.

202
00:12:37,560 --> 00:12:40,500
The expected value of a random
variable sometimes we also,

203
00:12:40,500 --> 00:12:43,210
call it the mean although we
really shouldn't we should call

204
00:12:43,210 --> 00:12:47,030
it the expected value when we're talking
about probability distributions we can

205
00:12:47,030 --> 00:12:50,248
call it the mean when we're talking
about things we compute from samples.

206
00:12:50,248 --> 00:12:54,750
>> Which we call statistics the expected
value of a random variable is

207
00:12:54,750 --> 00:12:57,670
a typical value that you
might expect it to assume.

208
00:12:57,670 --> 00:13:00,920
It's the weighted average of all
the potential realizations that

209
00:13:00,920 --> 00:13:04,160
random variable can take, where,
where the weights are provided by,

210
00:13:04,160 --> 00:13:09,070
the probabilities given by
the probability density or mass function.

211
00:13:09,070 --> 00:13:12,700
So, you can see under the first bullet
point on the line that begins with

212
00:13:12,700 --> 00:13:15,860
E there, we have a definition of expected
value for discrete random variable, and

213
00:13:15,860 --> 00:13:19,710
a definition of expected value for
a continuous random variable.

214
00:13:19,710 --> 00:13:25,970
And in fact, it's often useful to
just think of E as an operator and,

215
00:13:25,970 --> 00:13:29,110
that depending on whether we are talking
about discrete or continuous random

216
00:13:29,110 --> 00:13:32,750
variables, you would substitute in
what are the sum or the interval.

217
00:13:34,510 --> 00:13:37,650
and, all I have done in the next
line down is to substitute in,

218
00:13:37,650 --> 00:13:41,590
f where we had Ps up above.

219
00:13:41,590 --> 00:13:45,330
The expected value of a random vector is
simply the vector of expected values of

220
00:13:45,330 --> 00:13:45,860
it's components.

221
00:13:47,810 --> 00:13:50,080
The expected value of a function
of a random variable,

222
00:13:50,080 --> 00:13:51,200
this is actually kind of cute.

223
00:13:52,760 --> 00:13:55,340
It is simply the weighted average, or

224
00:13:55,340 --> 00:14:01,710
the expected the weighted average
value of the, derived random variable.

225
00:14:01,710 --> 00:14:04,980
But now because,
the values of the original X

226
00:14:06,000 --> 00:14:10,090
variable correspond to specific values for
the derived y variable,

227
00:14:10,090 --> 00:14:14,800
which I've called g of X here
the formulas generalize in this way.

228
00:14:14,800 --> 00:14:18,590
Meaning I can just compute all the
different possible values of g of X, and

229
00:14:18,590 --> 00:14:20,649
weight them by
the appropriate values of X.

230
00:14:22,450 --> 00:14:27,010
The expected deviation of X
from it's own expected value,

231
00:14:27,010 --> 00:14:32,570
which we sometimes call the mean,
is called the bias of the random

232
00:14:32,570 --> 00:14:36,600
variable X relative to it's mean or
relative to it's own expected value.

233
00:14:36,600 --> 00:14:38,620
We like to use the Greek letter mu.

234
00:14:39,832 --> 00:14:43,025
To denote that and we subscript it by
X to make it completely clear that we

235
00:14:43,025 --> 00:14:47,430
are talking about random variable X and
random variable X's distribution.

236
00:14:47,430 --> 00:14:51,060
The variants of a random variable is
the expected squared deviation from

237
00:14:51,060 --> 00:14:52,976
it's own mean.

238
00:14:52,976 --> 00:14:54,380
And there are formulas for it right there.

239
00:14:54,380 --> 00:14:55,880
You've probably seen all of that before.

240
00:14:58,150 --> 00:15:00,160
Here things will get just slightly tricky.

241
00:15:00,160 --> 00:15:04,180
The covariance of two random variables,
you've probably seen that before too.

242
00:15:04,180 --> 00:15:06,030
It's like a generalization of a variance.

243
00:15:07,140 --> 00:15:09,520
The variance of a random
vector is a matrix,

244
00:15:09,520 --> 00:15:11,828
which is called the variance
covariance matrix.

245
00:15:11,828 --> 00:15:15,760
And the variance, covariance matrix
must be square and symmetric.

246
00:15:15,760 --> 00:15:19,410
And it's simply has the variances on
the diagonal, and the covariances on

247
00:15:19,410 --> 00:15:24,440
the off diagonal, corresponding
to the different elements of X.

248
00:15:24,440 --> 00:15:29,080
We often use sigma squared as our
shorthand for, the variance and

249
00:15:29,080 --> 00:15:31,510
sigma Xi, Xj as a shorthand for
a covariance.

250
00:15:33,890 --> 00:15:34,790
This is the tricky part.

251
00:15:34,790 --> 00:15:38,540
The cross covariance between two random
vectors that's a different matrix.

252
00:15:38,540 --> 00:15:39,775
That's not the same variance,

253
00:15:39,775 --> 00:15:45,290
covariance matrix the cross covariance
matrix need not be symmetric or

254
00:15:45,290 --> 00:15:49,490
square because you can think of the rows
of that matrix as corresponding to

255
00:15:49,490 --> 00:15:53,310
the components of X and the columns
corresponding to the components of Y.

256
00:15:53,310 --> 00:15:55,500
But other than that the form
looks entirely familiar.

257
00:15:57,660 --> 00:15:58,820
Based on what we've already seen.

258
00:16:00,420 --> 00:16:05,050
So, there's a bit of tedious
equations on this slide, and

259
00:16:05,050 --> 00:16:08,120
I don't want to, bore you too much.

260
00:16:08,120 --> 00:16:12,070
The important thing here is
to say that point number one,

261
00:16:12,070 --> 00:16:13,620
and X is a linear operator.

262
00:16:13,620 --> 00:16:16,949
So if I take the expectation of
the sum of two random variables.

263
00:16:18,120 --> 00:16:20,340
It's the, sum of their expectations.

264
00:16:21,500 --> 00:16:25,080
And, if I multiply either or both of
those random variable by a constant,

265
00:16:25,080 --> 00:16:29,440
that constant simply comes out front.

266
00:16:29,440 --> 00:16:33,540
And you can, prove this yourself if
you want to by using that formula for

267
00:16:33,540 --> 00:16:35,730
the expected value of a transformation
of a random variable.

268
00:16:37,270 --> 00:16:39,610
Now here's an important point.

269
00:16:39,610 --> 00:16:41,330
If, X1 and X2 are independent,

270
00:16:41,330 --> 00:16:45,850
then the expected value of the product
of two random variables based on,

271
00:16:45,850 --> 00:16:48,560
those two independent random variables are
simply the product of their expectations.

272
00:16:49,860 --> 00:16:51,080
And that's very nice and

273
00:16:51,080 --> 00:16:54,950
very handy when you're trying to figure
out the expected value of something.

274
00:16:54,950 --> 00:16:56,040
That represents a product.

275
00:16:57,100 --> 00:17:00,710
however, it's really, really,
really important to recognize that,

276
00:17:00,710 --> 00:17:02,650
that is not true the other way around.

277
00:17:02,650 --> 00:17:08,060
It is not true that just because
covariance is zero, that,

278
00:17:08,060 --> 00:17:11,955
that implies that these random
variables are independent.

279
00:17:11,955 --> 00:17:15,630
Perhaps I should,
back track just slightly, and say that.

280
00:17:15,630 --> 00:17:19,050
The expect the product of
expectations equaling,

281
00:17:19,050 --> 00:17:24,310
the expectation of the product is
is essentially defines covariance,

282
00:17:24,310 --> 00:17:27,890
there will be the subtraction of
the mean terms involved but if you

283
00:17:27,890 --> 00:17:31,000
assume everything has zero mean to start
with then you don't worry about that, and

284
00:17:31,000 --> 00:17:36,800
what we're saying here is that the If
the random variables are independent,

285
00:17:36,800 --> 00:17:39,900
then, the covariants will be zero.

286
00:17:41,980 --> 00:17:46,370
the, only condition, the only case
in which the converse is true,

287
00:17:46,370 --> 00:17:49,730
namely that, the covariants
being zero implies independence,

288
00:17:49,730 --> 00:17:55,040
is if you're two, random variables,
X1 and X2 are bivariate Gaussian.

289
00:17:56,450 --> 00:17:58,070
Now finally, this may be,

290
00:17:58,070 --> 00:18:01,200
this is one of the most cool things
that there is to say about this topic.

291
00:18:02,340 --> 00:18:06,880
And that is to look at the notion
of conditional independence,

292
00:18:06,880 --> 00:18:10,020
which is I'm sorry the notion
of conditional expectation.

293
00:18:11,118 --> 00:18:15,020
The conditional expected value
of random variable X1 given X2.

294
00:18:15,020 --> 00:18:17,310
Is the ordinary definition of expectation,
but

295
00:18:17,310 --> 00:18:20,810
applied with the conditional
distribution of X1 given X2.

296
00:18:20,810 --> 00:18:25,340
And if we think of this object as
a function of random variable X2,

297
00:18:25,340 --> 00:18:28,860
this defines the regression
of exponent X1 on X2.

298
00:18:28,860 --> 00:18:32,950
And many of you may be thinking, but
that doesn't sound like, simple linear

299
00:18:32,950 --> 00:18:37,600
regression, or multiple regression that
I learned in my textbook at school.

300
00:18:37,600 --> 00:18:38,820
Well in fact it is.

301
00:18:38,820 --> 00:18:43,040
It's just that what you learn in school,
pertains to conditions where X1 and

302
00:18:43,040 --> 00:18:44,930
X2 are bivariate Gaussian.

303
00:18:44,930 --> 00:18:49,320
And if X1 and X2 are bivariate Gaussian,
then if you look at the value of

304
00:18:49,320 --> 00:18:53,710
this function, namely the expected value
of X1 given X2 is the function of X2,

305
00:18:53,710 --> 00:18:55,290
that will lie on a line.

306
00:18:55,290 --> 00:18:58,180
And that's,
where simple linear regression comes from.

307
00:18:58,180 --> 00:19:02,480
But this concept is much more general and
could apply in lots of places for example

308
00:19:02,480 --> 00:19:07,260
if you applied a clustering algorithm to
a data set, and you thought about X2 as

309
00:19:07,260 --> 00:19:12,750
being a cluster identifier a number
that identifies a cluster and the value

310
00:19:12,750 --> 00:19:18,890
of X1 as being a random draw from all
the Xs that belong to that cluster.

311
00:19:18,890 --> 00:19:24,120
Then you could say that
the regression of X1 on X2 is,

312
00:19:24,120 --> 00:19:28,100
are the mean functions of each of those
clusters, in other words if you made

313
00:19:28,100 --> 00:19:33,630
a function that had cluster number on the
x axis and the expected value of X1 for

314
00:19:33,630 --> 00:19:35,859
each of those clusters separately
that would be a regression.

315
00:19:37,270 --> 00:19:38,180
So it's a very general term.

316
00:19:39,832 --> 00:19:43,545
Now perhaps one of the most
useful things that

317
00:19:43,545 --> 00:19:47,690
ever existed in probability and
statistics.

318
00:19:47,690 --> 00:19:51,990
That you may we may not see again in
this lecture but it is definitely worth

319
00:19:51,990 --> 00:19:55,350
knowing about, it's called the law
of iterated conditional expectation.

320
00:19:55,350 --> 00:20:01,980
So we defined the expected value of
X1 given X2 on the earlier slide.

321
00:20:01,980 --> 00:20:07,430
And for clarity here, let's say we
have a bivariate distribution, X1 and

322
00:20:07,430 --> 00:20:11,190
X2 in the lower left, that's a very,
exaggerated bump there.

323
00:20:12,220 --> 00:20:16,500
And then I, slice that distribution along.

324
00:20:16,500 --> 00:20:22,240
For values defined along X2,
C1, C2, C3, and C4.

325
00:20:22,240 --> 00:20:25,340
And of course I'm only showing you
certain slices out of that distribution.

326
00:20:25,340 --> 00:20:27,080
But, I could slice at any location.

327
00:20:28,260 --> 00:20:31,640
And if I found the conditional expected
value of each of those conditional

328
00:20:31,640 --> 00:20:37,170
distributions, and then I averaged them
over the possible realizations of X2.

329
00:20:37,170 --> 00:20:41,330
I could reconstruct the expected
value of X1, and that again,

330
00:20:41,330 --> 00:20:44,365
ends up being very very very useful
in sort of the same way that

331
00:20:44,365 --> 00:20:47,490
Bayes' theorem ends up being useful
in that, sometimes it's easier for

332
00:20:47,490 --> 00:20:50,280
me to know things conditionally
than unconditionally.

333
00:20:50,280 --> 00:20:54,000
And this gives me a way to get
the unconditional expected value of X1

334
00:20:54,000 --> 00:20:57,895
if I know something, about
the conditional behavior of X1 given X2.

335
00:20:57,895 --> 00:21:02,778
[SOUND] And just for completeness, we
have to talk about the variance as well.

336
00:21:02,778 --> 00:21:06,559
And I'm going to, I'm, I'm stating
all these results basically for

337
00:21:06,559 --> 00:21:08,330
continuous random variables.

338
00:21:08,330 --> 00:21:11,750
But there are analogs for
the discrete and vector cases as well.

339
00:21:12,900 --> 00:21:16,220
The variance of a linear function
is not the sum of the variances.

340
00:21:17,530 --> 00:21:19,130
it, it, you have to account for

341
00:21:19,130 --> 00:21:23,610
the co-variance term, and that comes about
because of those Venn diagrams when we

342
00:21:23,610 --> 00:21:26,440
were looking at things like
probability of a intersection b.

343
00:21:26,440 --> 00:21:30,180
And we had to worry about,
not double counting the intersection.

344
00:21:30,180 --> 00:21:34,740
That intersection ends up being sort of,
related to the covariances here and

345
00:21:34,740 --> 00:21:35,830
that we have to worry about.

346
00:21:37,060 --> 00:21:40,970
So if you see on the first line of the
first bullet it says the variance of, two

347
00:21:40,970 --> 00:21:43,360
random variables, just look at the thing
on the left side of the plus sign and

348
00:21:43,360 --> 00:21:46,120
the thing on the right side of the plus
sign, is not merely the sum of their two

349
00:21:46,120 --> 00:21:49,980
variances, but there is this
covariance term, on the right.

350
00:21:51,110 --> 00:21:56,820
And by the way, the variance of a constant
times a random variable is the,

351
00:21:56,820 --> 00:21:59,060
square of the constant times
the variance of the random variable,

352
00:21:59,060 --> 00:22:00,990
because the variance is a squared thing.

353
00:22:02,580 --> 00:22:05,760
Variance of a non linear function,
we could,

354
00:22:05,760 --> 00:22:10,700
in principle go back ad try to work it
out the way we did with the expectations.

355
00:22:10,700 --> 00:22:15,850
But there's actually an easier way,
because a lot of times that's far too

356
00:22:15,850 --> 00:22:17,480
difficult to do in
the case of the variance.

357
00:22:17,480 --> 00:22:20,280
And we appeal to,
a Taylor series expansion.

358
00:22:21,554 --> 00:22:25,691
On the function G, if we're talking about,
random variable Y, which is a function of

359
00:22:25,691 --> 00:22:29,680
random variable X, and you'll see on
the right side of the second bullet there.

360
00:22:29,680 --> 00:22:31,491
I've simply written out a,

361
00:22:31,491 --> 00:22:35,500
first-order tailor expansion
about the expected value of X.

362
00:22:37,130 --> 00:22:37,910
For the function g of X.

363
00:22:37,910 --> 00:22:40,750
And then if I apply the variance
formula to the thing on

364
00:22:40,750 --> 00:22:42,930
the right side of
the approximately equal sign,

365
00:22:42,930 --> 00:22:47,460
I can get this approximate relationship
that does turn out to be pretty handy.

366
00:22:47,460 --> 00:22:51,700
And yeah, here's my point about
the covariance and independence.

367
00:22:51,700 --> 00:22:55,480
I guess I got to it a little
sooner than I meant to.

368
00:22:56,890 --> 00:23:02,340
Okay, so analogous to the conditional,
conditional expectations,

369
00:23:02,340 --> 00:23:06,670
there is a conditional variance formula,
which is going to reinforce something that

370
00:23:06,670 --> 00:23:09,810
I said a few minutes ago, which is,
you might think that what you could do

371
00:23:09,810 --> 00:23:14,270
is take these additional distributions and
average up their variances.

372
00:23:14,270 --> 00:23:17,520
To get the variance of the random
that you really care about,

373
00:23:17,520 --> 00:23:20,360
which is X1 in this case,
but you can't quite do that.

374
00:23:20,360 --> 00:23:21,100
If that were true,

375
00:23:21,100 --> 00:23:23,641
we wouldn't have the first term on
the right side of the equal sign.

376
00:23:23,641 --> 00:23:26,940
because if you did that,
you would be missing the notion of

377
00:23:26,940 --> 00:23:29,920
variability between
different values of X2.

378
00:23:29,920 --> 00:23:33,002
So, this also ends up being
an extremely formula.

379
00:23:33,002 --> 00:23:37,680
You can read more about in the, the Ross
Book, or in a number of other places.

380
00:23:37,680 --> 00:23:42,060
But in, in if you care about
the variance of a random variable and

381
00:23:42,060 --> 00:23:46,030
you don't know what it is, but you do know
something about what it is conditionally,

382
00:23:46,030 --> 00:23:47,370
this a nice way to figure it out.

383
00:23:48,760 --> 00:23:51,530
Okay, so, I wouldn't blame you
if by now you were thinking,

384
00:23:51,530 --> 00:23:54,510
why should I care about all of this,
it's seems a bit.

385
00:23:56,460 --> 00:24:00,910
Often mathland well the reason is because
we're going to build models of unknown or

386
00:24:00,910 --> 00:24:04,270
uncertain populations,
with probability distributions and

387
00:24:04,270 --> 00:24:07,900
we want to call these things,
let's call these process distributions.

388
00:24:07,900 --> 00:24:10,600
And we make inferences about
process distributions by

389
00:24:10,600 --> 00:24:13,290
computing statistics from samples.

390
00:24:13,290 --> 00:24:16,260
The statistics themselves are random
variables because they were computed from

391
00:24:16,260 --> 00:24:19,570
a sample that was chosen randomly, and
they have distributions of their own.

392
00:24:19,570 --> 00:24:21,530
We'll call these sampling distributions.

393
00:24:21,530 --> 00:24:25,690
The discipline of statistics is
largely concerned with understanding

394
00:24:25,690 --> 00:24:29,850
the relationship between a process
distribution parameter and

395
00:24:29,850 --> 00:24:33,150
the sampling distribution of a statistic
that is designed to estimate it.

396
00:24:33,150 --> 00:24:34,509
So that's where we're headed next.

397
00:24:36,178 --> 00:24:41,201
Here's I think it's the same two
references that I gave you earlier so

398
00:24:41,201 --> 00:24:42,800
I won't repeat that.

399
00:24:42,800 --> 00:24:46,100
But now we will move on to two modules
on basic concepts of inference.

