1
00:00:01,380 --> 00:00:04,668
Hi.
Welcome back to Part 2 of

2
00:00:04,668 --> 00:00:06,610
Basics of Inference.

3
00:00:08,940 --> 00:00:12,950
Lecture module for the Caltech-JPL Virtual
Summer School of Big Data Analytics.

4
00:00:15,580 --> 00:00:16,545
So in this section what we're

5
00:00:16,545 --> 00:00:20,970
going to talk about is the Central
Limit Theorem Confidence Intervals.

6
00:00:20,970 --> 00:00:23,480
We'll introduce a Bayesian formalism for

7
00:00:23,480 --> 00:00:28,900
inference and then we'll take a step
back and summarize a few things.

8
00:00:28,900 --> 00:00:33,400
The material on the Central Limit Theorem
is based largely on Tom Ferguson's 1996

9
00:00:33,400 --> 00:00:38,890
book called A Course in Large Sample
Theory which is an enormously.

10
00:00:40,140 --> 00:00:41,670
Useful reference.

11
00:00:41,670 --> 00:00:44,140
You do have to know a little bit
of math though, to appreciate it.

12
00:00:45,420 --> 00:00:50,910
So when last we spoke,
at the end of the last module we talked

13
00:00:50,910 --> 00:00:55,210
about the difficulty of computing maximum
likelihood estimates and their sampling

14
00:00:55,210 --> 00:01:00,150
distributions if you don't know what the
underlying probability distribution is.

15
00:01:00,150 --> 00:01:00,810
That would seem to be.

16
00:01:01,920 --> 00:01:05,338
a, a show stopper for, for
doing maximum likelihood.

17
00:01:05,338 --> 00:01:09,745
[SOUND] But there's a miracle out there.

18
00:01:09,745 --> 00:01:11,774
And the miracle is called
The Central Limit Theorem, and

19
00:01:11,774 --> 00:01:13,896
it really is quite something if
you stop and think about it.

20
00:01:13,896 --> 00:01:18,927
[SOUND] The Central Limit Theorem says
that if y1 through yn are a sequence of

21
00:01:18,927 --> 00:01:20,891
iid random variables, like.

22
00:01:20,891 --> 00:01:24,420
That which constitutes a sample.

23
00:01:24,420 --> 00:01:28,940
Each with a same expected value and
the same variance, both finite.

24
00:01:28,940 --> 00:01:31,320
Then the distribution of
a random variable, s of n,

25
00:01:31,320 --> 00:01:33,860
which is essentially a sum.

26
00:01:33,860 --> 00:01:36,710
A rescaled sum of the y's.

27
00:01:36,710 --> 00:01:38,480
Tens to the standard normal,

28
00:01:38,480 --> 00:01:42,470
that's a Gaussian distribution,
as the sample size goes to infinity.

29
00:01:42,470 --> 00:01:43,430
Or gets large.

30
00:01:43,430 --> 00:01:44,370
In other words.

31
00:01:44,370 --> 00:01:46,130
What I've shown down at the bottom there.

32
00:01:46,130 --> 00:01:51,840
If you take the cumulative distribution
function of that random variable sn

33
00:01:53,390 --> 00:01:58,200
and you look at how it goes as
n gets large, it converges to

34
00:01:58,200 --> 00:02:02,460
something that looks like that expression
there in the middle of the bottom line.

35
00:02:02,460 --> 00:02:04,130
That's the Gaussian.

36
00:02:04,130 --> 00:02:07,340
The integral of the Gaussian PDF.

37
00:02:07,340 --> 00:02:11,550
And that really is quite a stunning
fact if you think about it.

38
00:02:11,550 --> 00:02:12,370
There's.

39
00:02:12,370 --> 00:02:15,350
One wonders why that should be true.

40
00:02:15,350 --> 00:02:21,570
And in most statistics courses at
least at the undergraduate level.

41
00:02:21,570 --> 00:02:24,800
No one ever actually does
tell you why this is true or

42
00:02:24,800 --> 00:02:26,660
proves to you why it's true.

43
00:02:26,660 --> 00:02:29,830
And the reason is because it does
require some more sophisticated

44
00:02:29,830 --> 00:02:33,870
mathematics based on something
called moment generating functions.

45
00:02:33,870 --> 00:02:38,170
And you can find that proof if you're
interested in that you can certainly find

46
00:02:38,170 --> 00:02:40,820
that in many places probably.

47
00:02:40,820 --> 00:02:44,100
Probably even in Wikipedia by now and
it's really something to look at,

48
00:02:44,100 --> 00:02:45,450
it's quite, quite extraordinary.

49
00:02:47,292 --> 00:02:51,960
But the result can be written in any
of the following three equivalent ways

50
00:02:51,960 --> 00:02:56,980
it seems the most natural way to write it
is in terms of the sum random variables,

51
00:02:56,980 --> 00:03:01,580
the sum of the y's,
where the y's are the random variables.

52
00:03:01,580 --> 00:03:05,950
And the notation there has a little arrow
which says that the distribution of that

53
00:03:05,950 --> 00:03:11,260
object on the left converges,
in distribution to a Gaussian distribution

54
00:03:11,260 --> 00:03:17,400
with the mean n times the individual
random variable expected

55
00:03:17,400 --> 00:03:21,828
values and it's variance is n times
the variance of the individual variables.

56
00:03:21,828 --> 00:03:22,564
[SOUND] that.

57
00:03:22,564 --> 00:03:31,130
Convergence in distribution thing
could be a very long explanation.

58
00:03:31,130 --> 00:03:35,460
That is essentially convergence of
functions, so if you've had a math

59
00:03:35,460 --> 00:03:39,490
class that talked about what it means for
functions to converge, we're talking about

60
00:03:39,490 --> 00:03:43,350
distributions function,
distributions functions converging.

61
00:03:43,350 --> 00:03:43,890
In the sense.

62
00:03:43,890 --> 00:03:46,120
There are other senses of the,

63
00:03:46,120 --> 00:03:50,980
of convergence and at a certain point
I'm going to stop by bothering to

64
00:03:50,980 --> 00:03:54,540
distinguish between them because
it's not important at this level.

65
00:03:54,540 --> 00:03:55,530
But do be aware that,

66
00:03:55,530 --> 00:03:59,715
that expression represents a particular
kind of mathematical convergence.

67
00:03:59,715 --> 00:04:03,661
[SOUND] An equivalent formulation
is the expression in the middle.

68
00:04:03,661 --> 00:04:06,921
[SOUND] This is probably the one
you might be most familiar with.

69
00:04:06,921 --> 00:04:12,061
It says that the sample mean converges
to a Gaussian distribution [SOUND] with

70
00:04:12,061 --> 00:04:18,520
the sa, with the mean, or whose expected
value is the expected value that you seek.

71
00:04:18,520 --> 00:04:20,693
And whose variance goes
down with the sample size.

72
00:04:20,693 --> 00:04:24,870
That may be the one that,
that probably is most familiar.

73
00:04:24,870 --> 00:04:27,180
And the last one down there at the bottom.

74
00:04:27,180 --> 00:04:29,540
It's a little mysterious looking.

75
00:04:29,540 --> 00:04:34,190
But all it is,
is the same version as on the line above,

76
00:04:34,190 --> 00:04:39,600
but centered in the sense of
the thing in the parentheses is

77
00:04:39,600 --> 00:04:44,920
simply the deviation of y bar from
the thing that it's targeting.

78
00:04:44,920 --> 00:04:48,200
And then we're blowing it up by
a factor a square root of n.

79
00:04:48,200 --> 00:04:51,460
And the reason we want to do that is
because that random variable will have

80
00:04:51,460 --> 00:04:53,230
a mean of zero and

81
00:04:53,230 --> 00:04:59,234
it will have a variance that looks like
the variance of the original observations.

82
00:04:59,234 --> 00:05:02,850
And we're going to come
back to using that kind of

83
00:05:02,850 --> 00:05:06,000
a random variable that transformation of
a random variable a little bit later.

84
00:05:07,570 --> 00:05:10,210
And so it might be good to get
used to looking at it right now.

85
00:05:11,580 --> 00:05:14,700
Let's see, that limiting distribution
on the right is sometimes called

86
00:05:14,700 --> 00:05:17,080
the asymptotic distribution
of the statistic.

87
00:05:19,910 --> 00:05:21,770
There is also a central limit theorem for

88
00:05:21,770 --> 00:05:24,650
independent but not identically
distributed random variables.

89
00:05:24,650 --> 00:05:27,500
It's call the Lindeberg-Feller
central limit theorem.

90
00:05:27,500 --> 00:05:31,520
And it has some extra
conditions that have to be met.

91
00:05:31,520 --> 00:05:33,700
I'm not even going to
talk about what that is.

92
00:05:33,700 --> 00:05:34,930
You can look that up in a book.

93
00:05:34,930 --> 00:05:39,940
Just know that it exists and you can go
find it if you're actually dealing with

94
00:05:39,940 --> 00:05:47,080
a situation where you have a sample that
is not iid, but it's just id, independent.

95
00:05:48,480 --> 00:05:52,160
There's a Central Limit Theorem for
random vectors, not surprisingly, and

96
00:05:52,160 --> 00:05:54,960
it looks like the expression here.

97
00:05:54,960 --> 00:05:59,670
I'm showing only that centered and
scaled version.

98
00:05:59,670 --> 00:06:05,390
And there is central limit theorem for
functions of random variables and vectors

99
00:06:05,390 --> 00:06:09,890
which I'm showing down at the bottom and
this is a consequence of the fact that.

100
00:06:11,240 --> 00:06:16,450
When I transform by the function G I have
to account for that in the covariance,

101
00:06:16,450 --> 00:06:20,320
because my random variable is already
centered, it doesn't impact the main.

102
00:06:21,420 --> 00:06:24,140
This theorem is ha, is called,
at least in Tom Ferguson's book,

103
00:06:24,140 --> 00:06:25,270
it's called Cramer's Theorem.

104
00:06:26,910 --> 00:06:28,350
The Central Limit Theorem and

105
00:06:28,350 --> 00:06:32,326
Cramer's Theorem premieres are extremely
useful because many estimators end up

106
00:06:32,326 --> 00:06:36,650
being functions of sums or averages
of iid random variables or vectors.

107
00:06:36,650 --> 00:06:37,760
The sample variance for

108
00:06:38,890 --> 00:06:43,900
instance, is the function it is
a function of two different.

109
00:06:43,900 --> 00:06:49,240
Of a transformation of a sum and
using combining the Central Limit Theorem

110
00:06:49,240 --> 00:06:52,950
with Cramer's Theorem, we can obtain
the fact that the centered and

111
00:06:52,950 --> 00:07:02,080
scaled sample variance has a Gaussian
distribution with mean zero and variance.

112
00:07:02,080 --> 00:07:02,870
As I've shown it there.

113
00:07:04,040 --> 00:07:08,120
That mu sub y four is
the fourth central moment of y.

114
00:07:08,120 --> 00:07:12,840
Okay the sample correlation
coefficient for

115
00:07:12,840 --> 00:07:17,390
y variant random vectors actually
obeys the central limit theorem too.

116
00:07:17,390 --> 00:07:22,610
There wasn't enough room on the page to
show you what the asymptotic variance was.

117
00:07:22,610 --> 00:07:25,390
Of this, of this expression.

118
00:07:25,390 --> 00:07:30,190
Like everything else it can be looked
up but it's too ugly to write down and

119
00:07:30,190 --> 00:07:34,890
the point is that the sample
correlation coefficient

120
00:07:34,890 --> 00:07:40,080
does converge to something we know
as the correlation gets large.

121
00:07:41,260 --> 00:07:43,510
I'm sorry as the sample size gets large.

122
00:07:43,510 --> 00:07:48,100
So other kinds of statistics for which the
central limit there are holds are sample

123
00:07:48,100 --> 00:07:53,150
quantiles, rank statistics which
are things like if I take a sample and

124
00:07:53,150 --> 00:07:54,590
I want to look at say the maximum or

125
00:07:54,590 --> 00:07:58,090
the minimum or the third largest element,
those are statistics.

126
00:07:58,090 --> 00:07:59,510
They're called order statistics.

127
00:08:00,615 --> 00:08:02,760
Chi-squared statistics, remember that?

128
00:08:02,760 --> 00:08:05,130
Observe minus expected
squared over expected.

129
00:08:06,610 --> 00:08:08,940
Extrema, maxima or minima and many others.

130
00:08:08,940 --> 00:08:12,610
So it holds in a surprisingly
high number of cases.

131
00:08:12,610 --> 00:08:16,060
It also holds for
dependence sequences of random variables.

132
00:08:16,060 --> 00:08:17,570
Which is maybe a little surprising.

133
00:08:21,140 --> 00:08:24,430
I was, I think I'll just take
a minute here to tell you.

134
00:08:24,430 --> 00:08:27,140
What a stationary
m-dependence sequence is.

135
00:08:27,140 --> 00:08:32,960
It's kind of a daunting name, but
m-dependence simply refers to the fact

136
00:08:32,960 --> 00:08:37,970
that blocks of variables
are sort of travel together.

137
00:08:37,970 --> 00:08:38,850
They move together.

138
00:08:40,050 --> 00:08:42,960
Independent sets separated by.

139
00:08:42,960 --> 00:08:48,470
It length, separated by m in indices
are independent of each other,

140
00:08:48,470 --> 00:08:50,260
but they're dependent within.

141
00:08:50,260 --> 00:08:54,620
And a stationary distribution
of a sequence means that,

142
00:08:54,620 --> 00:08:56,030
that distribution, that the mean and

143
00:08:56,030 --> 00:09:00,500
the variance of that distribution do not
change as a function of your position with

144
00:09:00,500 --> 00:09:04,520
the, within the time series,
within the sequence of random variables.

145
00:09:04,520 --> 00:09:06,300
So, let's not worry too much about that.

146
00:09:08,130 --> 00:09:09,275
Here's the CLT for

147
00:09:09,275 --> 00:09:13,800
dependent sequences of independent,
of stationary independent sequences.

148
00:09:13,800 --> 00:09:18,810
And again, I,
let's not dwell too much on math here.

149
00:09:18,810 --> 00:09:20,930
But everything on the page
should look okay.

150
00:09:20,930 --> 00:09:25,270
In fact, if the y's were independent,
instead of dependent.

151
00:09:25,270 --> 00:09:29,920
The second to last line in the equation
would only have the very first term on

152
00:09:29,920 --> 00:09:31,310
the right side of the equal sign,

153
00:09:31,310 --> 00:09:35,180
because the variances would simply
add if everything was independent.

154
00:09:35,180 --> 00:09:38,590
But if they're not independent, they're
partially dependent out to some degree,

155
00:09:38,590 --> 00:09:41,880
then you have to cope with all
the covariances that arise from

156
00:09:41,880 --> 00:09:43,230
that dependence.

157
00:09:43,230 --> 00:09:46,273
And let's just call that whole thing on
the right hand side of the equal sign

158
00:09:46,273 --> 00:09:47,350
times squared.

159
00:09:47,350 --> 00:09:50,910
But we certainly can say,
that the sample mean obtained from

160
00:09:50,910 --> 00:09:53,620
a stationary independent sequence,
the centered and

161
00:09:53,620 --> 00:09:57,600
scaled version of that statistic,
does converge to a Gaussian distribution,

162
00:09:57,600 --> 00:10:00,300
with the right mean and
with a variance that we can know.

163
00:10:01,300 --> 00:10:02,309
That is assuming that.

164
00:10:03,312 --> 00:10:07,270
We can know all the terms in
the middle parts of that point.

165
00:10:07,270 --> 00:10:12,360
Now, maximum likelihood estimates do have
a really nice property which is that

166
00:10:12,360 --> 00:10:15,550
they're centered and scaled versions.

167
00:10:15,550 --> 00:10:18,054
They make if theta had is
a maximum likelihood estimate,

168
00:10:18,054 --> 00:10:21,158
subtract off the target theta so
now we are looking at the deviation.

169
00:10:21,158 --> 00:10:24,420
And then re-scale it by multiplying
it by the square root of n.

170
00:10:24,420 --> 00:10:29,680
That thing converges in distribution to a
Gaussian distribution with a mean of zero,

171
00:10:29,680 --> 00:10:31,670
which means unbiased.

172
00:10:31,670 --> 00:10:35,080
And a variance that we can figure out.

173
00:10:35,080 --> 00:10:36,920
the, the key is I'm going to,
I'm going to,

174
00:10:36,920 --> 00:10:39,080
I'm going to rain on
that parade in a second.

175
00:10:39,080 --> 00:10:41,040
But we can figure it out
if we know what f is.

176
00:10:41,040 --> 00:10:45,140
That think on the right, i of theta
is called the Fisher Information.

177
00:10:45,140 --> 00:10:48,820
And it is a measure of the information
content, of the rand in the,

178
00:10:48,820 --> 00:10:52,320
in the random variable y,
about the vector theta.

179
00:10:52,320 --> 00:10:56,440
Oh, I'm sorry, about the quantity theta,
the parameter theta.

180
00:10:56,440 --> 00:11:00,300
So if you, we, let's define something that

181
00:11:01,300 --> 00:11:05,660
that expression at the top line of
the equal sign in the middle, I'm sorry,

182
00:11:05,660 --> 00:11:07,830
the top line of the eq,
first equation in the middle.

183
00:11:09,400 --> 00:11:11,950
I think that's the Greek letter psi,

184
00:11:11,950 --> 00:11:15,640
if I'm not mistaken, psi of y comma
theta is called the score function.

185
00:11:16,700 --> 00:11:20,300
And it is simply the derivative
of the log likelihood.

186
00:11:20,300 --> 00:11:23,670
And remember it was that derivative of
the log likelihood that we set equal to

187
00:11:23,670 --> 00:11:26,900
zero and solved for
to get the maximum likelihood estimate.

188
00:11:26,900 --> 00:11:29,250
So we've actually seen that thing before

189
00:11:30,460 --> 00:11:35,320
and if we let y be random in
that expression, then the,

190
00:11:35,320 --> 00:11:38,990
what's called the Fisher Information
is just the variance of the score.

191
00:11:40,530 --> 00:11:45,460
And this, this may be starting
to go off into a bunch of

192
00:11:45,460 --> 00:11:49,768
words that have been packed together
that don't mean anything to you anymore.

193
00:11:49,768 --> 00:11:55,280
But as we've seen a couple
of times already the,

194
00:11:55,280 --> 00:11:59,060
the punch line of the story is that we
can know what this is in principle.

195
00:12:00,100 --> 00:12:03,390
And in fact the inverse of the Fisher
information is something called

196
00:12:03,390 --> 00:12:07,300
the Cramér–Rao lower bound which
is the minimum variance that any

197
00:12:07,300 --> 00:12:08,970
estimator can have.

198
00:12:08,970 --> 00:12:13,560
So what this is telling us is that
maximum likelihood estimates are what's

199
00:12:13,560 --> 00:12:17,950
called asymptotically optical which means
if the sample size is large enough and

200
00:12:17,950 --> 00:12:20,610
you have the luxury of knowing what f is.

201
00:12:20,610 --> 00:12:23,540
So that you can compute these things,
then you know,

202
00:12:23,540 --> 00:12:24,970
you're doing the best you can possibly do.

203
00:12:26,190 --> 00:12:28,770
But then there's the catch, right?

204
00:12:28,770 --> 00:12:30,600
Which is that in most
practical situations,

205
00:12:30,600 --> 00:12:34,890
we may not know what f is,
but we'll come back to that.

206
00:12:34,890 --> 00:12:38,450
Let's talk about confidence, let's do
a quick detour into a confidence interval.

207
00:12:38,450 --> 00:12:44,020
Remember we had the sampling
distribution of our statistic and

208
00:12:44,020 --> 00:12:47,610
the sampling distribution of our statistic
basically quantifies everything we

209
00:12:49,110 --> 00:12:53,210
know about the behavior of
the statistic if it were to

210
00:12:53,210 --> 00:12:56,740
have been calculated repeatedly
over multiple random samples.

211
00:12:56,740 --> 00:12:57,850
So.

212
00:12:57,850 --> 00:13:00,850
Something like a maximum likelihood
estimate I've now, I've,

213
00:13:00,850 --> 00:13:04,990
I've drawn the multiple samples in
the cube with the green dots, I've written

214
00:13:04,990 --> 00:13:08,880
down all the different thetas I might have
computed from each of those samples, and

215
00:13:08,880 --> 00:13:11,930
I've written I've just wrote down

216
00:13:11,930 --> 00:13:16,135
a particularly nasty looking
probability density function.

217
00:13:16,135 --> 00:13:19,020
For theta hat here.

218
00:13:19,020 --> 00:13:21,430
Just to make the point that not
everything in the world is Gaussian or

219
00:13:21,430 --> 00:13:26,990
even symmetric, and what we know,
if we knew what the distribution

220
00:13:26,990 --> 00:13:31,660
was of at least the centered version of
our statistic, which would be theta hat,

221
00:13:31,660 --> 00:13:36,970
which is computed from y, I'm now calling
it theta hat of y instead of g of y.

222
00:13:36,970 --> 00:13:40,660
And I look at how it deviates from
the true value that where trying to

223
00:13:40,660 --> 00:13:42,310
estimate theta.

224
00:13:42,310 --> 00:13:47,270
Then I know that with probability 0.95,
by definition that random variable falls

225
00:13:47,270 --> 00:13:51,910
between the number l of 0.025 and
the number u of 0.975 and

226
00:13:53,140 --> 00:13:58,980
if I know that then I can reverse
engineer and expression that looks like.

227
00:13:58,980 --> 00:14:03,930
The first thing in the equation
under the first bullet, which is

228
00:14:03,930 --> 00:14:10,120
designed to tell us something about
how likely it is that our interval,

229
00:14:10,120 --> 00:14:15,450
computed from our sample y, actually
contains the true value that we're after.

230
00:14:15,450 --> 00:14:21,210
And people are sometimes tempted to regard
that as a statement about the probability

231
00:14:21,210 --> 00:14:25,230
of obtaining a particular value of theta,
which is the true parameter.

232
00:14:25,230 --> 00:14:28,541
And if we're going to be good
classical statisticians we can't,

233
00:14:28,541 --> 00:14:30,800
we don't want to say it that way.

234
00:14:30,800 --> 00:14:33,610
Turns out if we're Bayesians we're going
to be allowed to say it that way and

235
00:14:33,610 --> 00:14:34,829
so we'll get to that in a second.

236
00:14:35,950 --> 00:14:42,460
But the statement, the only thing that's
random in that probability statement

237
00:14:42,460 --> 00:14:47,540
is y, and therefore the upper and
lower limits of the confidence interval.

238
00:14:47,540 --> 00:14:52,810
So, that's just a quick introduction
there you probably remember that if our

239
00:14:52,810 --> 00:14:56,460
statistic follows a Gaussian distribution,
with mean zero and variance one.

240
00:14:56,460 --> 00:14:59,026
[SOUND] Then that lower limits
ends up being minus 1.96,

241
00:14:59,026 --> 00:15:01,603
which is what you'd look up in a table,
or get r to tell you.

242
00:15:01,603 --> 00:15:03,844
And the upper limit ends
up being plus 1.96.

243
00:15:03,844 --> 00:15:09,397
But this is a much more general concept
[SOUND] than just the Gaussian situation.

244
00:15:09,397 --> 00:15:11,591
[SOUND] Okay.

245
00:15:11,591 --> 00:15:16,750
So, now let's talk about Bayes,
Bayesian's again, Bayes Theorem again.

246
00:15:16,750 --> 00:15:19,450
You remember Baye's theorem from
one of the probability modules.

247
00:15:20,500 --> 00:15:23,630
And I'd like to reiterate that
Baye's theorem is just math and

248
00:15:23,630 --> 00:15:24,520
like any other math.

249
00:15:26,030 --> 00:15:31,650
Whether it's a good thing to do or a bad
thing to do to invoke it depends on how

250
00:15:31,650 --> 00:15:35,530
confident you feel about the assumptions
that are necessary to invoke it.

251
00:15:35,530 --> 00:15:38,170
So up to this point,
we've been good frequentists.

252
00:15:38,170 --> 00:15:40,880
We've been good classical statisticians
who think about probability as

253
00:15:40,880 --> 00:15:42,020
long run relative frequency.

254
00:15:43,140 --> 00:15:47,670
And we've treated everything having to do
with, we've treated theta as a fixed but

255
00:15:47,670 --> 00:15:51,080
unknown quantity, a kind of an ordinary
variable in the algebraic sense.

256
00:15:52,080 --> 00:15:56,450
And our inference was based on playing
this sort of sensitivity game with theta.

257
00:15:56,450 --> 00:15:59,260
Saying how does the probability
of observing the sample I got,

258
00:15:59,260 --> 00:16:00,520
which I call the likelihood.

259
00:16:00,520 --> 00:16:04,990
How does that change, as I sort of turn
the knob on potential values of theta?

260
00:16:04,990 --> 00:16:07,010
So, there's nothing random about theta.

261
00:16:07,010 --> 00:16:07,690
There was only,

262
00:16:07,690 --> 00:16:11,265
sort of the sensitivity of the likelihood
to different choices of theta.

263
00:16:12,770 --> 00:16:15,710
But if you're a Bayesian,
you can go a little bit further.

264
00:16:15,710 --> 00:16:16,710
And.

265
00:16:16,710 --> 00:16:22,390
I'll, I'll try to tell a, a small
story straight, sometimes I don't get

266
00:16:22,390 --> 00:16:25,890
the story straight and it ends up being
embarrassing, but I'll take the risk.

267
00:16:25,890 --> 00:16:30,400
I would like to play a game and
what I'm going to do is I'm going to

268
00:16:30,400 --> 00:16:33,630
flip a coin and before I flip the coin,
I'm going to ask you

269
00:16:33,630 --> 00:16:38,350
what you think the probability is that the
head, that heads will come up on the coin.

270
00:16:38,350 --> 00:16:43,000
And you'll probably tell
me that it's a half and

271
00:16:43,000 --> 00:16:46,780
then I flip the coin and
let's say it does come up a head.

272
00:16:46,780 --> 00:16:53,878
Well, the coin, the random variable
has been realized and at that point

273
00:16:53,878 --> 00:16:59,389
the probability of it being a head is one,
because it was a head.

274
00:17:01,350 --> 00:17:04,120
Now we'll play the game again, and
what I'm going to do is I'm going to

275
00:17:04,120 --> 00:17:07,950
flip the coin, but I'm going to put
my hand over it so you can't see it.

276
00:17:07,950 --> 00:17:10,950
And now I'm going to ask you what
you think the probability is that

277
00:17:10,950 --> 00:17:12,630
the coin came up a head.

278
00:17:12,630 --> 00:17:15,810
And you'll probably tell me it's a half,
again.

279
00:17:15,810 --> 00:17:16,310
Because.

280
00:17:17,310 --> 00:17:21,520
The fact is, that even though
the coin is either a head or

281
00:17:21,520 --> 00:17:24,730
its not at that point,
you still don't know.

282
00:17:24,730 --> 00:17:29,360
And you're willing to use probability to
express the fact that you don't know.

283
00:17:29,360 --> 00:17:33,010
So, if you're okay with that,
then you might be a Bayesian.

284
00:17:33,010 --> 00:17:34,340
In fact, you probably are a Bayesian.

285
00:17:35,590 --> 00:17:42,230
And if you feel that way then you probably
also feel that it's okay to treat theta,

286
00:17:42,230 --> 00:17:44,820
that unknown parameter
that we're interested in

287
00:17:44,820 --> 00:17:49,100
as a random variable instead of
some fixed but unknown quantity.

288
00:17:49,100 --> 00:17:54,380
And rather than play the sensitivity game,
We will use

289
00:17:54,380 --> 00:17:57,100
probability to describe the con, we only,

290
00:17:57,100 --> 00:18:02,980
we will use the conditional probability
distribution of theta, given y

291
00:18:02,980 --> 00:18:09,100
to describe what we think the the behavior
of random variable theta is.

292
00:18:10,130 --> 00:18:13,820
And in order to do that,
we have to assert a marginal distribution.

293
00:18:13,820 --> 00:18:17,960
For theta, because if you remember
back to the probability module,

294
00:18:17,960 --> 00:18:20,170
when we first wrote down Bayes theorem.

295
00:18:20,170 --> 00:18:22,920
You can see it over here on
the right in grey letters.

296
00:18:22,920 --> 00:18:25,030
At the bottom of the page
I rewrote it there.

297
00:18:25,030 --> 00:18:28,600
It's the probability of b given
a times the probability of

298
00:18:28,600 --> 00:18:29,960
a over the probability of b.

299
00:18:29,960 --> 00:18:31,330
That's all Bayes theorem says.

300
00:18:31,330 --> 00:18:37,560
And if random variable let's
say what I'm interested in,

301
00:18:37,560 --> 00:18:41,180
is random variable theta, the conditional
distribution of theta given y.

302
00:18:42,270 --> 00:18:46,680
I don't know what that is, but
maybe I have a much better idea,

303
00:18:46,680 --> 00:18:49,190
of what the conditional distribution of y,
given theta is.

304
00:18:50,340 --> 00:18:52,770
So, if I'm drawing from
a Gaussian distribution.

305
00:18:53,770 --> 00:18:57,360
And I don't know what
the expected value is.

306
00:18:57,360 --> 00:18:58,544
I don't know what mu is.

307
00:18:58,544 --> 00:19:04,623
I don't know what mu what,
I'm sorry, what's what y given,

308
00:19:04,623 --> 00:19:10,910
what theta given, I'm sorry, if I'm
drawing from a Gaussian distribution.

309
00:19:12,590 --> 00:19:16,760
I don't, and I don't know what
the expected value is, rather than

310
00:19:16,760 --> 00:19:21,490
simply twiddling with potential values
of theta that maximize the likelihood,

311
00:19:21,490 --> 00:19:24,450
I'm going to actually use
Bayes theorem to compute the,

312
00:19:24,450 --> 00:19:30,300
what's called the posterior distribution
of mu given the sample that I got.

313
00:19:30,300 --> 00:19:33,340
And that was not particularly well done,
but I'm just going to.

314
00:19:33,340 --> 00:19:34,940
Don't buy it.

315
00:19:34,940 --> 00:19:36,000
Let it be.
So

316
00:19:36,000 --> 00:19:40,790
here's my cartoon of being a Bayesian and
I'm doing Bayesian confidence intervals.

317
00:19:42,440 --> 00:19:45,290
So the I think

318
00:19:45,290 --> 00:19:49,380
the cartoon starts with a little
graphic all the way over on the left.

319
00:19:49,380 --> 00:19:53,210
Which is what I'm going to postulate or
what I'm going to declare to

320
00:19:53,210 --> 00:19:58,020
be what's called the prior distribution or
the marginal distribution of theta.

321
00:19:58,020 --> 00:20:01,380
Now what I'm doing, remember what I'm
doing is I'm basically assuming theta and

322
00:20:01,380 --> 00:20:02,450
y are jointly distributed.

323
00:20:04,254 --> 00:20:05,660
And therefore I can,

324
00:20:05,660 --> 00:20:11,120
I could get the I have this notion
of a marginal distribution of theta.

325
00:20:12,240 --> 00:20:16,950
And I have a distribution of y given
theta, because let's say I take a random

326
00:20:16,950 --> 00:20:22,730
draw from that distribution f sub theta, I
get a particular value of theta and then I

327
00:20:22,730 --> 00:20:29,780
generate a bunch of potential samples from
a population that has that value of theta.

328
00:20:29,780 --> 00:20:32,530
Well, if I'd had picked a different
value to theta to start with,

329
00:20:32,530 --> 00:20:36,110
from the distribution,
I would get a different set of samples,

330
00:20:36,110 --> 00:20:38,460
like I do in the se,
say the second cube there.

331
00:20:38,460 --> 00:20:41,920
And for each potential realization
of the random variable theta,

332
00:20:41,920 --> 00:20:45,750
I might get a different set of
sample of green dots there.

333
00:20:45,750 --> 00:20:50,180
And would get a different
little sampling distribution.

334
00:20:51,440 --> 00:20:52,320
Of theta.

335
00:20:52,320 --> 00:20:56,090
I'm sorry, of sampling distribution
of theta, that's correct.

336
00:20:57,380 --> 00:21:01,740
So essentially what we're doing as
Bayesian's is we're saying, we're just

337
00:21:01,740 --> 00:21:06,320
using conditional probability to say I'm
going to average over all possible choices

338
00:21:06,320 --> 00:21:11,430
of theta, with the weights for
that averaging being give by fs of theta.

339
00:21:11,430 --> 00:21:15,650
And I'm going to use that, to get
this posterior distribution of theta,

340
00:21:15,650 --> 00:21:18,320
after I've seen the data y or
the sample y.

341
00:21:20,090 --> 00:21:25,180
Just one little note is that, here's
a place where we might definitely want to

342
00:21:25,180 --> 00:21:29,310
use a sufficient statistic, instead
of working with the original sample,

343
00:21:29,310 --> 00:21:30,700
even though I've drawn the picture as if.

344
00:21:32,150 --> 00:21:34,850
As if I could work with
the original sample here.

345
00:21:34,850 --> 00:21:39,180
If I'm a Bayesian, I can write down
a Bayesian confidence interval,

346
00:21:39,180 --> 00:21:42,890
which will take the form of being
a probability statement about theta.

347
00:21:42,890 --> 00:21:45,570
And by the way, it's also
a probability statement about l and

348
00:21:45,570 --> 00:21:49,173
u, because those two things
are functions of the observed data y.

349
00:21:49,173 --> 00:21:53,320
Okay, all right, so

350
00:21:53,320 --> 00:21:56,720
I'll make a few editorial comments
now in the way of a summary,

351
00:21:56,720 --> 00:21:59,760
which is that it all boils down to how
you want to model your unknown parameter.

352
00:21:59,760 --> 00:22:03,770
Is it a random variable, or
is it a fixed but unknown value?

353
00:22:03,770 --> 00:22:07,210
And the, I guess the anecdote I

354
00:22:07,210 --> 00:22:10,490
like to give is we've probably
all done the calculus problem.

355
00:22:10,490 --> 00:22:13,500
Where we're told that
a bath tub is filling up

356
00:22:13,500 --> 00:22:16,630
at a certain rate because water is
pouring in from the spigot, and it's

357
00:22:16,630 --> 00:22:20,876
draining out at a certain rate because the
plug, the stopper isn't in properly, and

358
00:22:20,876 --> 00:22:25,130
we would like to know how long it'll
take for the bath tub to overflow.

359
00:22:25,130 --> 00:22:27,480
And if that's really your bath tub.

360
00:22:27,480 --> 00:22:30,720
Then you darn well care whether
the assumptions about how

361
00:22:30,720 --> 00:22:33,780
fast the water's coming in and
how fast it's going out are right.

362
00:22:35,200 --> 00:22:36,280
And it's the same thing here.

363
00:22:36,280 --> 00:22:40,080
You know, you have to think about the
consequences of making a wrong choice, or

364
00:22:40,080 --> 00:22:42,900
the consequences of getting
a bad estimate in this case.

365
00:22:42,900 --> 00:22:48,330
As to whether you want to be a Bayesian or
a Frequentist about how you're doing it.

366
00:22:48,330 --> 00:22:51,990
Do you think you have reliable
information about theta,

367
00:22:51,990 --> 00:22:57,480
that you can bring to bear on the problem
by asserting a prior using Bayes theorem?

368
00:22:57,480 --> 00:22:57,980
Or don't you,

369
00:22:57,980 --> 00:23:01,720
do you want to fall back to sort of
the sensitivity version of the problem?

370
00:23:01,720 --> 00:23:05,620
So in my opinion, the Bayesian formalism
is more complete, more flexible, and

371
00:23:05,620 --> 00:23:09,750
lends itself to conditional modeling,
and therefore is actually.

372
00:23:09,750 --> 00:23:13,330
In a lot of cases at least in
scientific applications, where we

373
00:23:13,330 --> 00:23:18,240
think we know something about how things
work, a really good way to model things.

374
00:23:18,240 --> 00:23:21,710
But whether you are a Frequentist or
Bayesian, you still have to know or

375
00:23:21,710 --> 00:23:26,540
assume things about the distribution in
order to use any of these techniques.

376
00:23:26,540 --> 00:23:29,550
And that's where, you know,
that's where I think that,

377
00:23:29,550 --> 00:23:33,470
that People become disappointed and
disillusioned with formal

378
00:23:33,470 --> 00:23:37,200
statistics because that's,
a lot of the time that's not true.

379
00:23:37,200 --> 00:23:40,540
And so in the next two modules what I'd
like to do is talk to you about a couple

380
00:23:40,540 --> 00:23:47,440
new procedure that have come
on the scene in the last,

381
00:23:47,440 --> 00:23:51,280
it's actually been 20 or 30 years,
so they're not actually that new.

382
00:23:52,320 --> 00:23:55,520
That might help us out in
the situation where we can't use

383
00:23:55,520 --> 00:23:59,330
the Central Limit Theorem to help us out
with what the form of the density is, or

384
00:23:59,330 --> 00:24:00,950
what the form of the sampling
distribution is.

385
00:24:02,310 --> 00:24:05,170
So that's where we'll go next Let's see.

386
00:24:05,170 --> 00:24:07,716
I did, I had these references here,

387
00:24:07,716 --> 00:24:13,030
I believe I've shown both of them before,
or all three of them before, so.

388
00:24:14,060 --> 00:24:17,080
I'll leave you to explore
those references on your own.

