1
00:00:00,790 --> 00:00:07,920
Okay, this is the, start of the second
of our sub-topics in the lectures

2
00:00:07,920 --> 00:00:11,840
on Uncertainty and Inference, with a Big
Data Analytics Virtual Summer School.

3
00:00:13,050 --> 00:00:16,790
Now we're going to talk a little about
inference and hopefully things will

4
00:00:16,790 --> 00:00:21,820
now be migrating in a direction
that seems a bit more practical or

5
00:00:21,820 --> 00:00:24,000
at least getting to more practical.

6
00:00:24,000 --> 00:00:29,500
So, here we're going to talk really about
a, a very small subset of what you would

7
00:00:29,500 --> 00:00:33,390
really think about as the topic
of inference and statistics.

8
00:00:33,390 --> 00:00:37,960
We're going to talk about
some basic concepts and

9
00:00:37,960 --> 00:00:40,220
then the principle of maximum likelihood.

10
00:00:42,080 --> 00:00:44,540
And we'll use the principle
of maximum likelihood,

11
00:00:44,540 --> 00:00:48,250
to illustrate the concept of
the uncertainty of the estimate and

12
00:00:48,250 --> 00:00:51,510
talk about desirable properties of
estimates, which as it turns out,

13
00:00:51,510 --> 00:00:56,100
maximum likelihood estimates tend to have,
which is a handy thing.

14
00:00:56,100 --> 00:00:58,340
So, basic concepts of inference.

15
00:00:58,340 --> 00:01:02,460
Remember we talked about
a process distribution which I

16
00:01:02,460 --> 00:01:08,140
might represent with a random variable x,
having a, pdf or pmf.

17
00:01:08,140 --> 00:01:10,080
Let's talk about pdf's from now on,

18
00:01:10,080 --> 00:01:15,510
and you'll just understand that what I
mean, is either pdf or pmf as appropriate.

19
00:01:15,510 --> 00:01:17,690
So the process distribution, is f of x.

20
00:01:19,820 --> 00:01:25,390
And let's say that what we are interested
in, is in knowing something about what

21
00:01:25,390 --> 00:01:29,220
the expected value of that distribution
is, what is the most typical outcome I

22
00:01:29,220 --> 00:01:33,330
could expect, if I took a random
draw from that distribution.

23
00:01:34,670 --> 00:01:38,957
Let and
we'll call that expected value mu sub x.

24
00:01:38,957 --> 00:01:41,620
So what would I do about this?

25
00:01:41,620 --> 00:01:45,951
I'm, I'm sure you have seen this before in
statistics classes,you take a sample, and

26
00:01:45,951 --> 00:01:46,586
let's say,

27
00:01:46,586 --> 00:01:50,300
we're going to draw a sample of size
capital N from the process distribution.

28
00:01:50,300 --> 00:01:55,507
And there's my cartoon picture of sampling
there, which is supposed to look like

29
00:01:55,507 --> 00:02:01,670
a filter, and so, I'm going to let some
of those realizations pass through.

30
00:02:01,670 --> 00:02:05,872
And I will obtain a sample size in
which I have written now as Y1 to YN,

31
00:02:05,872 --> 00:02:08,220
over there on the right.

32
00:02:08,220 --> 00:02:11,920
So, note that the distribution
I'm interested in, is a scalar,

33
00:02:11,920 --> 00:02:15,300
it's a distribution of
a scalar random variable.

34
00:02:15,300 --> 00:02:16,810
But the sample that I obtain,

35
00:02:16,810 --> 00:02:20,740
is an n dimensional object,
living in n dimensional space.

36
00:02:22,260 --> 00:02:27,060
The form of the probability
distribution of

37
00:02:27,060 --> 00:02:32,650
the sample random vector is determined
by the form of f, not surprisingly,

38
00:02:32,650 --> 00:02:37,030
that's where we got the observations from,
or where we got the values of y from.

39
00:02:37,030 --> 00:02:39,470
But also by the sampling
procedure that was used,

40
00:02:39,470 --> 00:02:41,410
by the nature of that screening process.

41
00:02:42,600 --> 00:02:45,700
I will just make a comment that
that's actually very important for

42
00:02:45,700 --> 00:02:49,790
those of us that work in
areas like remote sensing.

43
00:02:49,790 --> 00:02:53,490
Because going back to what I said earlier
in the first module about observational

44
00:02:53,490 --> 00:02:59,350
studies and experiments we kind of have
to live with the data that we get.

45
00:02:59,350 --> 00:03:03,336
And in remote sensing, the samples
that we get are, what you might call

46
00:03:03,336 --> 00:03:08,280
the systematic samples, in the sense that,
satellite orbits tend to repeat.

47
00:03:08,280 --> 00:03:10,930
So we tend to see things always
at the same time of day, for

48
00:03:10,930 --> 00:03:13,360
example, if our satellite
is in polar orbit.

49
00:03:13,360 --> 00:03:15,530
So in that case,
you would not be able to say that,

50
00:03:15,530 --> 00:03:19,740
that sampling, that filter,
that screen there was sort of a, kind of

51
00:03:19,740 --> 00:03:24,180
an unbiased representation of all possible
outcomes that we could have gotten.

52
00:03:24,180 --> 00:03:27,190
So we do need to keep that in mind,
but for the moment, let's ignore that.

53
00:03:29,240 --> 00:03:33,290
So our sample is identified as
the random vector Y1 to Yn there.

54
00:03:35,430 --> 00:03:38,016
And it has a PDF of it's own.

55
00:03:38,016 --> 00:03:43,590
And if each draw of y from f sub x,

56
00:03:43,590 --> 00:03:47,890
was independent and from the same
distribution which we call independent and

57
00:03:47,890 --> 00:03:49,840
identically distributed, IID for short.

58
00:03:50,850 --> 00:03:55,940
Then the joint pdf of all n of those
random variables Y1 through YN,

59
00:03:55,940 --> 00:03:58,109
is simply the product of their
individual probabilities.

60
00:03:59,270 --> 00:04:01,740
And that is what I have written down
at the bottom of the page there.

61
00:04:01,740 --> 00:04:07,470
It might look a little funny but the joint
probability of the vector little Y,

62
00:04:07,470 --> 00:04:11,780
which I'm, I'm now, being very
careful now to have the argument to

63
00:04:11,780 --> 00:04:16,100
my function there be a lower case letter,
because it's a realized value.

64
00:04:16,100 --> 00:04:20,765
Is whatever the pdf of capital Y1 is
evaluated at the observed value of little

65
00:04:20,765 --> 00:04:24,998
y1 and so on, through that multiplication
and I can write that sort of

66
00:04:24,998 --> 00:04:29,736
more compactly with the product operator
at the lower [INAUDIBLE] on the left and

67
00:04:29,736 --> 00:04:33,323
all I've done is to absorb that
each of the y's actually has

68
00:04:33,323 --> 00:04:37,986
the same distribution as x because that's
where I drew the y's from [SOUND].

69
00:04:37,986 --> 00:04:42,380
Okay, so let's look at a really,
really cartoonishly simple example here.

70
00:04:42,380 --> 00:04:48,120
Let's pretend, I take a sample of size
three, from my process distribution.

71
00:04:48,120 --> 00:04:49,980
So I've Y1, Y2, and Y3.

72
00:04:49,980 --> 00:04:53,890
And that's a random vector, it's a single
point, in a three dimensional space.

73
00:04:56,370 --> 00:05:01,130
And it is shown on the right, there's a
cube that's a three dimensional space and

74
00:05:01,130 --> 00:05:04,607
I've identified the sample I
actually got with the green dot.

75
00:05:08,725 --> 00:05:12,570
So and the question is,
what can learn about mu of x?

76
00:05:12,570 --> 00:05:15,010
Front the sample that we have,
Y equals little y.

77
00:05:16,600 --> 00:05:22,600
So the PDF of the sample, actually
is related to whatever mu sub x is.

78
00:05:22,600 --> 00:05:28,110
It's going to that, there, where the
center of that distribution is going to

79
00:05:28,110 --> 00:05:31,979
have an effect on where in that three
dimensional space, that green dot is.

80
00:05:33,140 --> 00:05:37,790
So it would seem,
that a logical way of trying to

81
00:05:37,790 --> 00:05:43,110
get at what mu of x is, is to ask
ourselves, what the probability is,

82
00:05:43,110 --> 00:05:47,430
of obtaining the y we actually obtained,
for different values of mu of x.

83
00:05:47,430 --> 00:05:49,529
And that is the principle
of maximum likelihood.

84
00:05:51,270 --> 00:05:54,600
Here what I'm trying to show,
is I have three copies of

85
00:05:54,600 --> 00:05:58,620
the three-dimensional space,
and I have the same sample,

86
00:05:58,620 --> 00:06:01,340
which is the location of the colored dots,
in each of the three.

87
00:06:01,340 --> 00:06:02,920
because I got, I got the answer I got.

88
00:06:02,920 --> 00:06:06,000
But what I've done is I've color-coded it,

89
00:06:06,000 --> 00:06:11,060
according to the probability
that I would get that answer,

90
00:06:11,060 --> 00:06:15,710
that I would get that sample under
three different choices of mu of x.

91
00:06:15,710 --> 00:06:19,290
And then I've made the whole
picture a lot more compact,

92
00:06:19,290 --> 00:06:22,977
by simply drawing those probabilities
as a function of mu of x and

93
00:06:22,977 --> 00:06:27,620
the little tiny postage stamp sized
graphic on the right whose labels you

94
00:06:27,620 --> 00:06:32,610
can't read, probably,
the X axis in that plot is labeled mu x.

95
00:06:32,610 --> 00:06:36,930
The Y axis I call likelihood and
the title is likelihood.

96
00:06:38,000 --> 00:06:43,140
so, if we think of the PDF of
the sample viewed as a function of

97
00:06:43,140 --> 00:06:47,580
the parameter we care about mu of x that's
called the likelihood function of mu of x.

98
00:06:47,580 --> 00:06:52,600
And instead of writing F sub Y of Y and
mu of x which we could do.

99
00:06:53,660 --> 00:06:56,298
We're going to write it
as L of mu of x of y

100
00:06:56,298 --> 00:07:00,921
to sort of highlight the importance
of the choice of mu here.

101
00:07:00,921 --> 00:07:05,153
We're now, basically going to think of y
as fixed at the value we got and we're

102
00:07:05,153 --> 00:07:09,738
going to twitter around with mu sub x, in
order to try to maximise that probability.

103
00:07:09,738 --> 00:07:16,393
The maximum likelihood estimate maximizes
L for the real life sample All right.

104
00:07:16,393 --> 00:07:19,200
So, you might have seen this example too.

105
00:07:19,200 --> 00:07:23,340
Another freakishly simple thing but
it gets the point across.

106
00:07:23,340 --> 00:07:26,980
If the original distribution of X,
that we care about,

107
00:07:26,980 --> 00:07:31,880
is Gaussian, with an unknown expected
value, mu sub x and a variance equal to 1.

108
00:07:31,880 --> 00:07:33,730
Let's pretend we know that.

109
00:07:33,730 --> 00:07:36,980
Then we know what the form of f sub x, is.

110
00:07:36,980 --> 00:07:39,250
And we know what the form of f sub y is.

111
00:07:39,250 --> 00:07:42,180
Because, those Ys were
drawn independently.

112
00:07:42,180 --> 00:07:46,770
So the joint distribution of the three Ys,
is just the product of three copies,

113
00:07:46,770 --> 00:07:50,134
of the Gaussian distribution,
with a mean at mu x and a variance of 1.

114
00:07:52,190 --> 00:07:58,630
I'm going to write that, as I did at the
bottom to include mu X as a variable in,

115
00:07:58,630 --> 00:08:01,580
in the notation there,
just to emphasize the dependence on mu.

116
00:08:03,580 --> 00:08:05,860
In cases where we know
the form of f sub y,

117
00:08:05,860 --> 00:08:10,060
it's possible to solve for the maximum
likelihood estimate analytically.

118
00:08:10,060 --> 00:08:12,370
By finding the derivatives of
the likelihood function and

119
00:08:12,370 --> 00:08:14,000
setting them equal to zero.

120
00:08:14,000 --> 00:08:15,230
Okay?

121
00:08:15,230 --> 00:08:16,660
It's often easier to solve for

122
00:08:16,660 --> 00:08:19,900
the log of the likelihood, which is
a fair thing to do, because the log of

123
00:08:19,900 --> 00:08:23,250
a likelihood will have a maximum same
place as the raw likelihood does.

124
00:08:24,980 --> 00:08:26,100
And this is how you would do it.

125
00:08:26,100 --> 00:08:27,760
I probably don't have to go through this.

126
00:08:27,760 --> 00:08:32,120
This is why we end up with
the sample mean, y bar,

127
00:08:32,120 --> 00:08:35,310
as the maximum likelihood estimate
of the unknown parameter mu x.

128
00:08:35,310 --> 00:08:38,259
So, it's not just because,
it seems like the right thing to do,

129
00:08:38,259 --> 00:08:40,310
it's because it maximizes the likelihood.

130
00:08:42,830 --> 00:08:44,730
Okay.
So now let's talk about uncertainty.

131
00:08:44,730 --> 00:08:46,600
That was all, I think,
fairly straightforward and

132
00:08:46,600 --> 00:08:48,940
fairly intuitively reasonable.

133
00:08:50,340 --> 00:08:54,100
The part that may be a less little
intuitively reasonable is that

134
00:08:54,100 --> 00:08:57,630
the one sample we got,
the green dot in the cube over there,

135
00:08:57,630 --> 00:09:00,300
was only one of many different
things we could have got.

136
00:09:00,300 --> 00:09:01,880
And I've tried to indicate that,

137
00:09:01,880 --> 00:09:04,980
with all the other light green dots
that I've put in that box too.

138
00:09:06,470 --> 00:09:10,630
Each of those, represents, a different
sample that we could have gotten.

139
00:09:10,630 --> 00:09:13,060
And there is a mu hat of X, or

140
00:09:13,060 --> 00:09:16,830
a maximum likelihood estimate,
that goes with each and every one of them.

141
00:09:16,830 --> 00:09:21,210
And it is the distribution over all those
possible samples and their values, their

142
00:09:21,210 --> 00:09:25,980
corresponding values of mu hat sub X, that
defines the uncertainty of our estimate.

143
00:09:27,060 --> 00:09:30,120
And that's all we're trying
to get at in inference and

144
00:09:30,120 --> 00:09:32,180
with a quantification of uncertainty.

145
00:09:32,180 --> 00:09:37,320
So, let's redraw the picture,
as I've done here.

146
00:09:37,320 --> 00:09:40,360
It's the same three-dimensional box,
with the dark green dot and

147
00:09:40,360 --> 00:09:42,360
the light green dots on the left.

148
00:09:42,360 --> 00:09:47,590
And I've simply listed there in
a color that I hope everyone can see.

149
00:09:47,590 --> 00:09:50,180
The different mu hat x's that we would get

150
00:09:51,270 --> 00:09:54,770
from from each of those
light green samples.

151
00:09:54,770 --> 00:09:57,520
And if I made a histogram of those or
better yet,

152
00:09:57,520 --> 00:10:01,600
I was able to draw the probability
density function of those mu hat of x's,

153
00:10:01,600 --> 00:10:04,580
I might get something
that we see on the right.

154
00:10:04,580 --> 00:10:10,180
And there would be desirable properties of
the way we would want that PDF to look.

155
00:10:10,180 --> 00:10:13,810
In other words, the way we want our
estimate theta hats of x to behave.

156
00:10:13,810 --> 00:10:17,510
We would want it to be unbiased, meaning
that we would want the peak of that

157
00:10:17,510 --> 00:10:22,891
distribution, its expected value,
to be, at the true value of mu sub x.

158
00:10:24,680 --> 00:10:27,560
And that's the unbasedness condition
there, its saying the same thing.

159
00:10:27,560 --> 00:10:31,280
Which is to say that the expected value of
the difference, between the estimate and

160
00:10:31,280 --> 00:10:34,210
the target, mu sub x would be 0.

161
00:10:34,210 --> 00:10:38,680
And we would also like that distribution
to be very narrow around that true value.

162
00:10:39,780 --> 00:10:43,010
Or around whatever the expected value is,
which is to simply say,

163
00:10:43,010 --> 00:10:47,780
that we want the variance of
the statistic mu hut sub x to have,

164
00:10:47,780 --> 00:10:52,090
to be smaller than, the variance for
any other estimate.

165
00:10:52,090 --> 00:10:56,868
And any other estimate which
I've called Mu tilde here.

166
00:10:56,868 --> 00:10:58,519
And estimate that has
both those properties is

167
00:10:58,519 --> 00:10:59,846
called minimum variance unbiased.

168
00:10:59,846 --> 00:11:00,982
M-V-U-E.

169
00:11:00,982 --> 00:11:03,860
Minimum variance unbiased estimate.

170
00:11:03,860 --> 00:11:07,980
And it's something that
we normally strive for

171
00:11:07,980 --> 00:11:13,060
if not require when we invent
ways of getting estimates.

172
00:11:13,060 --> 00:11:17,540
Now, it's true that we could
choose any function of the sample,

173
00:11:17,540 --> 00:11:18,870
as our estimator of theta hash.

174
00:11:18,870 --> 00:11:22,180
I could choose something that has
nothing to do with the sample at all,

175
00:11:22,180 --> 00:11:22,820
if I wanted to.

176
00:11:22,820 --> 00:11:26,960
I could choose any single theta hat, or

177
00:11:26,960 --> 00:11:32,130
any single value of y,
of g of y if I wanted to.

178
00:11:32,130 --> 00:11:36,020
But we have to be careful
because if estimators don't have

179
00:11:36,020 --> 00:11:37,440
those desirable properties,

180
00:11:37,440 --> 00:11:42,600
then we don't have any guarantees about
the conclusions that we draw from them.

181
00:11:42,600 --> 00:11:44,880
So, typically, we choose g so

182
00:11:44,880 --> 00:11:49,760
that it's the sample version of the, thing
we want but as I illustrated earlier,

183
00:11:49,760 --> 00:11:52,690
that's not just because it seems
like the right thing to do,

184
00:11:52,690 --> 00:11:58,910
it's because of the legacy of the maximum
likelihood maximum likelihood derivation.

185
00:12:00,520 --> 00:12:04,080
When the thing we care about is
a moment of the original distribution.

186
00:12:05,672 --> 00:12:10,362
There's a fair chance that that's
actually the right thing to do, but

187
00:12:10,362 --> 00:12:14,000
it may not be optimal in
the sense of being MVUE.

188
00:12:14,000 --> 00:12:18,522
Here's a key point I think you can think
of that statistic that you choose to

189
00:12:18,522 --> 00:12:22,370
compute, from the sample,
as a form of data reduction.

190
00:12:22,370 --> 00:12:26,760
We've been using in our example,
we've been using a sample of size three,

191
00:12:26,760 --> 00:12:29,810
which is a point in
three dimensional space.

192
00:12:29,810 --> 00:12:32,160
Most samples are much larger than three.

193
00:12:32,160 --> 00:12:35,880
It would be, rather silly, to take
a sample of size three in most cases.

194
00:12:35,880 --> 00:12:38,940
And so, suddenly we're in very
high dimensional spaces and it's

195
00:12:38,940 --> 00:12:43,680
not really possible to draw pictures that
way or to even operate, with those things.

196
00:12:43,680 --> 00:12:48,140
So we're looking for good forms of
dimension reduction to apply to

197
00:12:48,140 --> 00:12:52,920
our sample in order to get good
estimators of the things we care about.

198
00:12:52,920 --> 00:12:58,060
And there is a concept out there called
statistical sufficiency which says that if

199
00:12:58,060 --> 00:13:03,320
the theta hatch statistic we choose
the g function we choose to apply.

200
00:13:04,470 --> 00:13:07,230
Contains all the information
about theta that is contained in

201
00:13:07,230 --> 00:13:11,210
the original Y then theta have is called
the sufficient statistic for theta.

202
00:13:11,210 --> 00:13:14,200
And there is a formal probabilistic
definition in terms of

203
00:13:15,880 --> 00:13:17,950
the PDFs of the various
densities involved, but

204
00:13:17,950 --> 00:13:22,209
I'm not going to put that here,
is easy to look that up.

205
00:13:22,209 --> 00:13:29,540
Now it may be hard,
to solve maximum likelihood equations.

206
00:13:29,540 --> 00:13:34,042
Particularly if you don't know what F is
and even if you do, well of course you

207
00:13:34,042 --> 00:13:38,269
can't solve the maximum likelihood
equation if you don't know what F is.

208
00:13:38,269 --> 00:13:41,390
Even if you do know what F is,
if F is non-Gaussian or

209
00:13:41,390 --> 00:13:45,020
some other very nicely
behaved distribution.

210
00:13:45,020 --> 00:13:48,100
It can be very hard to
solve maximum likelihood.

211
00:13:48,100 --> 00:13:50,820
If the the y's are not independent and

212
00:13:50,820 --> 00:13:56,170
identically distributed, it can be hard to
solve for maximum likelihood because then

213
00:13:56,170 --> 00:13:59,080
you have a hard time saying
what the join distribution is.

214
00:13:59,080 --> 00:14:02,920
So I already said if N is greater
than three it can be hard.

215
00:14:02,920 --> 00:14:04,900
And it can be hard if we're talking

216
00:14:05,920 --> 00:14:09,220
about arbitrary process distribution
parameters of interest.

217
00:14:09,220 --> 00:14:11,400
Not simple things like means or
other moments.

218
00:14:13,160 --> 00:14:17,670
So, that's the end of this module and
in the next module, we're going to look

219
00:14:17,670 --> 00:14:22,190
at some things that you might be able to
do in the absence of knowing what f is.

220
00:14:23,630 --> 00:14:27,730
Here are a couple of references
that I like, a couple of textbooks.

221
00:14:27,730 --> 00:14:31,920
The top on is actually an upper
division undergraduate-level text.

222
00:14:31,920 --> 00:14:33,860
The bottom one is
a first-year graduate text.

223
00:14:33,860 --> 00:14:36,610
They're both quite good.

224
00:14:36,610 --> 00:14:38,590
And I think we said that so,
let's move on.

