1
00:00:01,340 --> 00:00:05,100
Hi and welcome back to
the JPL-Caltech Virtual Summer School and

2
00:00:05,100 --> 00:00:06,810
Big Data Analytics.

3
00:00:06,810 --> 00:00:11,974
I'm Amy Braverman from the Jet Propulsion
Laboratory, and in this module I'd like to

4
00:00:11,974 --> 00:00:17,206
talk to you about the Bootstrap, which
is a statistical, technique for [SOUND]

5
00:00:17,206 --> 00:00:21,970
estimating or approximating sampling
distributions of complex statistics.

6
00:00:21,970 --> 00:00:25,330
[SOUND] As you might recall
from the previous module,

7
00:00:25,330 --> 00:00:29,950
whether you're a Frequentist or a
Bayesian, you need to know something about

8
00:00:29,950 --> 00:00:34,300
the sampling distribution of
the statistic that you're interested in.

9
00:00:34,300 --> 00:00:39,243
And how it depends on, the true or
realized value of the population parameter

10
00:00:39,243 --> 00:00:43,504
that you're trying to estimate,
or learn something about.

11
00:00:43,504 --> 00:00:48,895
and, we've looked at a bunch of [SOUND]
analytical methods in the last module.

12
00:00:48,895 --> 00:00:53,383
Which we can use when
we know the underlying

13
00:00:53,383 --> 00:00:58,880
probability density
function of the process.

14
00:00:58,880 --> 00:01:01,170
We call that the process distribution.

15
00:01:01,170 --> 00:01:05,070
Or when we can invoke the Central Limit
Theorem and we have a very large sample.

16
00:01:05,070 --> 00:01:09,458
[SOUND] however, those things are,
in the real world,

17
00:01:09,458 --> 00:01:12,563
sometimes hard to have at our disposal.

18
00:01:12,563 --> 00:01:17,050
We sometimes don't know, whether we
can use the Central Limit Theorem.

19
00:01:17,050 --> 00:01:20,690
We sometimes don't know
whether what the form of

20
00:01:20,690 --> 00:01:25,710
our underlying probability process
distribution is, or the sampling

21
00:01:25,710 --> 00:01:30,050
distribution of the statistic we're
interested in either for that matter.

22
00:01:30,050 --> 00:01:34,200
So the question is can we get information.

23
00:01:34,200 --> 00:01:36,798
About the distribution
of our statistics and

24
00:01:36,798 --> 00:01:42,390
how it might depend on the true
perimeter we care about without making

25
00:01:42,390 --> 00:01:48,180
those assumptions and I want to apologized
for the abuse of notation that,

26
00:01:48,180 --> 00:01:52,260
is embodied in my little expression for
f sub Theta hat.

27
00:01:53,330 --> 00:01:55,230
Bar capital Theta there.

28
00:01:55,230 --> 00:01:59,420
I said way back at the beginning
how important notation is, and

29
00:01:59,420 --> 00:02:03,110
that we would represent random
variables with capital letters and

30
00:02:03,110 --> 00:02:06,050
their realized values with
lower case letters and

31
00:02:06,050 --> 00:02:12,250
I would seemed to have violated that
by using Theta hat for both purposes.

32
00:02:12,250 --> 00:02:13,060
In that expression.

33
00:02:13,060 --> 00:02:19,110
But, Theta hat is so
traditionally the, notation used for

34
00:02:19,110 --> 00:02:25,200
a sample's statistics that I
can't tear myself away from it.

35
00:02:25,200 --> 00:02:28,320
So, in this module we're going to
look at a technique called Bootstrap.

36
00:02:28,320 --> 00:02:33,670
In fact, two particular versions of
the Bootstrap, I would like to say that.

37
00:02:33,670 --> 00:02:39,440
I made, great use of some notes from
Professor Aniron, Anirban Gupta,

38
00:02:39,440 --> 00:02:44,710
DasGupta at Purdue University
in the Department of Statistics.

39
00:02:44,710 --> 00:02:46,240
And Professor Charles Geyer in the School

40
00:02:46,240 --> 00:02:49,170
of Statistics at
the University of Minnesota.

41
00:02:49,170 --> 00:02:53,840
And a, book that I'll give you a reference
to, called Politis, Romano, and Wolf.

42
00:02:53,840 --> 00:02:55,360
1999's book's subsampling.

43
00:02:55,360 --> 00:03:03,610
So the Bootstrap, the ordinary
Bootstrap is quite old by now.

44
00:03:03,610 --> 00:03:08,872
I think it was first, popularized
in about 1979 by Brad Efron and

45
00:03:08,872 --> 00:03:11,330
Rob Tibshirani from Stanford.

46
00:03:12,730 --> 00:03:17,900
And they were, looking for a way to, solve
this problem of how can we learn about

47
00:03:17,900 --> 00:03:23,170
the sampling distribution of a statistic
when we don't want to make assumptions?

48
00:03:23,170 --> 00:03:25,930
This is what's called
a non-parametric method.

49
00:03:25,930 --> 00:03:31,090
Parametric referring to having to
make assumptions about parameters, or

50
00:03:31,090 --> 00:03:34,317
about the form of a distribution function,
or a density function, or a mass function.

51
00:03:34,317 --> 00:03:36,679
[SOUND] so.
[SOUND] Let's say we have an iid sample,

52
00:03:36,679 --> 00:03:40,072
Y1 through YN, as we had before,
from a distribution F.

53
00:03:40,072 --> 00:03:43,017
[SOUND] Earlier we called that
distribution F sub X, but

54
00:03:43,017 --> 00:03:44,623
I'm just going to call it F now.

55
00:03:44,623 --> 00:03:49,596
[SOUND] And we have a statistic, which
as before, is a function of the sample.

56
00:03:49,596 --> 00:03:52,239
And we want to know its
sampling distribution.

57
00:03:53,700 --> 00:03:58,030
First thing to notice is that the sample
itself actually defines a distribution.

58
00:03:58,030 --> 00:04:01,340
We call that the empirical
distribution function of the sample.

59
00:04:01,340 --> 00:04:06,640
That puts mass 1 over N on each of the
realized values that belong to the sample.

60
00:04:06,640 --> 00:04:11,920
So let's call this, empirical
distribution function F hat sub N.

61
00:04:11,920 --> 00:04:15,850
And I'm careful to put the N on there
because that distribution changes as I

62
00:04:15,850 --> 00:04:18,030
have more, and
more elements of the sample.

63
00:04:18,030 --> 00:04:23,254
So, the graphic on the left
is just a plot of an F hat,

64
00:04:23,254 --> 00:04:28,700
sub N for a simple example where
I drew a sample of size five.

65
00:04:30,040 --> 00:04:32,810
I think I got these, I might have gotten
these from Gaussian I'm not sure.

66
00:04:32,810 --> 00:04:38,790
So there they are, there's five numbers
there and you'll recall that for discrete

67
00:04:38,790 --> 00:04:44,580
distributions the cumulative distribution
function is a step function as it is here,

68
00:04:44,580 --> 00:04:49,000
because essentially what I'm going to do
is I'm going to sample from the sample.

69
00:04:49,000 --> 00:04:52,250
And therefore, the only values I
can get are these five values.

70
00:04:54,630 --> 00:04:58,787
So, [SOUND] let's define a resample,
as a sample of size N,

71
00:04:58,787 --> 00:05:03,957
which was the original sample size,
drawn from F hat N with replacement.

72
00:05:03,957 --> 00:05:07,712
And let's call those realizations
Y star 1 through Y star N

73
00:05:07,712 --> 00:05:10,330
are those random variables.

74
00:05:10,330 --> 00:05:14,990
So notice I am sampling, with replacement
which means every time I draw something I

75
00:05:14,990 --> 00:05:18,980
put it back so I can get the same
thing twice and I'm going to

76
00:05:18,980 --> 00:05:23,814
draw a sample of the same size as the
original sample, I call that a resample.

77
00:05:23,814 --> 00:05:28,240
On the resample I'm going to
compute the statistic Theta hat.

78
00:05:28,240 --> 00:05:31,298
And I'm going to call that
Theta hat star because I

79
00:05:31,298 --> 00:05:34,750
used the value Y1 star
through Y1 star to obtain it.

80
00:05:34,750 --> 00:05:39,820
I'm going to do that,
B times where B is a large number, and

81
00:05:39,820 --> 00:05:45,520
you may ask me how big B is and there's
actually a longer answer to that than you,

82
00:05:45,520 --> 00:05:47,780
than you might, think but
we'll get to that.

83
00:05:48,860 --> 00:05:53,710
And we will estimate the cumulative
distribution function of

84
00:05:53,710 --> 00:05:54,740
our statistic, Theta hat.

85
00:05:54,740 --> 00:05:57,440
Which is computed from
the original sample.

86
00:05:58,470 --> 00:06:00,950
By a cumulative distribution
function I'm calling H Boot.

87
00:06:00,950 --> 00:06:02,800
So now I've switched from F to H.

88
00:06:02,800 --> 00:06:06,770
I've kept the N in there to show that
it's going to be different depending on

89
00:06:06,770 --> 00:06:09,360
how large the original sample is.

90
00:06:09,360 --> 00:06:13,640
And I subscripted it with Boot,
to indicate it's a Bootstrap distribution.

91
00:06:13,640 --> 00:06:17,840
And all it is, is the,
probability that a random draw

92
00:06:20,230 --> 00:06:22,940
from the sample itself, that's Theta star.

93
00:06:25,390 --> 00:06:27,170
Realizes a value that's less than or
equal to t.

94
00:06:28,250 --> 00:06:32,080
And I'm going to approximate that with
a Monte Carlo estimate where after

95
00:06:32,080 --> 00:06:34,940
doing this capital B times I
simply count the number of

96
00:06:34,940 --> 00:06:38,320
my Theta star B's that were less than two.

97
00:06:38,320 --> 00:06:41,710
I add them up and
I divide by that big num, that big B.

98
00:06:41,710 --> 00:06:45,380
And there you have it, it's a Monte Carlo
estimate of that probability for

99
00:06:45,380 --> 00:06:46,670
that particular value of t.

100
00:06:47,690 --> 00:06:51,850
You may or may not have seen the notation
one, with an argument there,

101
00:06:51,850 --> 00:06:53,870
that's called an indicator function, and

102
00:06:53,870 --> 00:06:57,770
it simply takes on the value one if,
it's true and zero if it's not.

103
00:06:59,200 --> 00:07:03,450
Okay, so here's a simple example,
here's my original sample that's copied

104
00:07:03,450 --> 00:07:06,473
from the previous page, and let's say
that I'm interested in the median.

105
00:07:07,953 --> 00:07:09,780
I'm interested in the median and
by the way,

106
00:07:09,780 --> 00:07:14,830
that's a particularly nasty choice and in
the examples that I show you, we're going

107
00:07:14,830 --> 00:07:20,810
to see just how well this doesn't work
as well as how well it does work for

108
00:07:20,810 --> 00:07:23,960
a choice of, of, of something that's
a little complicated like the median.

109
00:07:23,960 --> 00:07:25,810
If I showed you this for the mean.

110
00:07:25,810 --> 00:07:28,930
It would all come out beautifully,
and we'd all go home and

111
00:07:28,930 --> 00:07:31,980
say that's nice, but
wonder what we might have learned.

112
00:07:31,980 --> 00:07:37,480
So, the true median of
this sample is 2.82.

113
00:07:37,480 --> 00:07:42,840
And then I, draw B additional
resamples where you can see that now,

114
00:07:42,840 --> 00:07:45,860
in some cases, I have gotten
the same value more than once.

115
00:07:45,860 --> 00:07:49,770
And on each resample I compute the median.

116
00:07:51,100 --> 00:07:56,850
And the, PDF or PMF of g is going to be
approximated by the histogram of those

117
00:07:56,850 --> 00:08:03,562
Theta hat stars, as the CDF is going to be
approximated by that Monte Carlo average.

118
00:08:03,562 --> 00:08:04,737
[SOUND] Okay.
So, now we're at

119
00:08:04,737 --> 00:08:06,540
the really important part.

120
00:08:08,850 --> 00:08:12,780
We really want the true CDF of Theta hat.

121
00:08:12,780 --> 00:08:14,960
But we only have one value of Theta hat,

122
00:08:14,960 --> 00:08:18,130
namely the one computed
from the original sample.

123
00:08:18,130 --> 00:08:20,470
We don't know what the true CDF is.

124
00:08:20,470 --> 00:08:22,560
Let's call it H sub N of t.

125
00:08:24,310 --> 00:08:29,810
We can, by this,
by this procedure that we just described,

126
00:08:29,810 --> 00:08:34,520
compute what H Boot sub N is,
in the way that we described.

127
00:08:34,520 --> 00:08:37,402
And our hope is that H Boot is close to H,
because if

128
00:08:37,402 --> 00:08:42,720
we're going to make inferences based on
H Boot and claim they apply to Theta hat.

129
00:08:42,720 --> 00:08:45,070
Then that would have to be true.

130
00:08:45,070 --> 00:08:47,050
So, the question is, is that true?

131
00:08:47,050 --> 00:08:52,230
And to, get after that point,
we might think about the following things.

132
00:08:52,230 --> 00:08:56,710
There are really sort of two main sources
of error in making that, approximation.

133
00:08:57,740 --> 00:09:02,050
Probably the really important one
is treating those resample draws.

134
00:09:02,050 --> 00:09:05,300
As if they were draws from the original
population rather than draws from

135
00:09:05,300 --> 00:09:08,230
the empirical distribution of
the one sample we actually obtained.

136
00:09:08,230 --> 00:09:14,000
The second source of error is the Monte
Carlo approximation of the probability.

137
00:09:14,000 --> 00:09:17,730
But that can be driven down by
making B sufficiently large.

138
00:09:17,730 --> 00:09:22,110
It's the other one
that's really important.

139
00:09:22,110 --> 00:09:25,280
Okay, so Now we do know that,

140
00:09:25,280 --> 00:09:31,110
F hat the empirical distribution function
of a sample of size N, converges to,

141
00:09:31,110 --> 00:09:35,460
the true distribution function F
as the sample size gets large.

142
00:09:35,460 --> 00:09:37,730
In other worse, as N goes to infinity.

143
00:09:37,730 --> 00:09:40,230
And I specifically didn't put
a little d over top of that,

144
00:09:40,230 --> 00:09:42,070
although I probably could have.

145
00:09:42,070 --> 00:09:45,060
But I believe that we want to be
a little more careful there so

146
00:09:45,060 --> 00:09:47,230
I'm just going to leave that
kind of vague and loose.

147
00:09:48,270 --> 00:09:53,470
And refer you to books to determine just
what sense that convergence is in for

148
00:09:53,470 --> 00:09:56,900
the purposes of the rest of this lecture
your common sense is as good as any.

149
00:09:58,250 --> 00:10:02,240
So if that's true then the wise
star ends are getting more and

150
00:10:02,240 --> 00:10:04,270
more like the original Y ends so
that's good.

151
00:10:05,670 --> 00:10:07,680
Does that alone guarantee,

152
00:10:07,680 --> 00:10:12,460
that the approximation of HN
by H Boot is good enough?

153
00:10:15,040 --> 00:10:15,680
Not really, no.

154
00:10:15,680 --> 00:10:17,290
Unfortunately not.

155
00:10:17,290 --> 00:10:19,290
It depends on a bunch of things.

156
00:10:19,290 --> 00:10:21,570
It depends on the nature of that
transformation g, in other words,

157
00:10:21,570 --> 00:10:23,710
on the form of Theta hat.

158
00:10:23,710 --> 00:10:27,800
And on some of the underlying properties
of the true distribution function F.

159
00:10:27,800 --> 00:10:32,820
The main point I want to make here, we'll
go into that last point in a minute, but

160
00:10:32,820 --> 00:10:37,650
the main point is, when we say
the Bootstrap works, this is what we mean.

161
00:10:37,650 --> 00:10:40,280
We mean that H Boot there,

162
00:10:40,280 --> 00:10:43,600
is what's called a statistically
consistent estimate.

163
00:10:43,600 --> 00:10:44,550
Of H sub N.

164
00:10:44,550 --> 00:10:51,700
In other words the Bootstrap distribution,
converges for N going to infinity

165
00:10:51,700 --> 00:10:58,090
to the true distribution so, hopefully,
that comes as a shock to you.

166
00:10:59,540 --> 00:11:03,260
Because to say something works
doesn't mean it always does work.

167
00:11:03,260 --> 00:11:06,990
In this sets, that sounds,
sounds like a, contradiction.

168
00:11:06,990 --> 00:11:11,670
But, wh, when we say the Bootstrap or
any of these methods work, what me mean,

169
00:11:11,670 --> 00:11:15,550
in fact, for all that matter we mean,
the central limit theorem,

170
00:11:15,550 --> 00:11:19,010
they work because they converge
the right thing as N gets large.

171
00:11:19,010 --> 00:11:21,640
We haven't said how large N needs to be.

172
00:11:21,640 --> 00:11:25,900
So we need to keep that in mind that
if we do an experiment with a finite N

173
00:11:27,140 --> 00:11:30,050
we have to worry a little bit about how
big that N is and whether it's big enough.

174
00:11:31,700 --> 00:11:34,540
Okay so now I'm going to, going to
finish off this module with a bunch of

175
00:11:34,540 --> 00:11:38,610
examples here that hopefully I can
lead you through in a clear way.

176
00:11:40,090 --> 00:11:42,500
So the example I want to talk about.

177
00:11:42,500 --> 00:11:47,655
Is estimating the median
from a sample of size 3,000.

178
00:11:50,290 --> 00:11:54,890
For the the median value of
a Chi-squared distribution.

179
00:11:54,890 --> 00:11:58,060
So what's shown on the left here
is a Chi-squared distribution.

180
00:11:58,060 --> 00:12:00,650
I used R to simply plot this.

181
00:12:00,650 --> 00:12:03,620
And the true median of a Chi-squared
distribution with three degrees of

182
00:12:03,620 --> 00:12:07,760
freedom, which is what I used here so

183
00:12:07,760 --> 00:12:12,692
that it would be,
pleasingly skewed is 2.366.

184
00:12:12,692 --> 00:12:16,707
If I take one sample of 3000, from a
Chi-squared distribution where everything

185
00:12:16,707 --> 00:12:21,480
is independent, I've made a histogram
of what's, of what's in that sample.

186
00:12:21,480 --> 00:12:23,880
And I've shown you where the median is and
so for

187
00:12:23,880 --> 00:12:28,210
that one sample that I took,
obviously the median from the sample,

188
00:12:28,210 --> 00:12:31,550
is not the same as the median
of the true distribution.

189
00:12:31,550 --> 00:12:36,160
So I think we all already knew that,
but I wanted to make that point, clear.

190
00:12:36,160 --> 00:12:40,630
So simply assuming that 2.415 is.

191
00:12:40,630 --> 00:12:42,670
The true median would not be a good idea.

192
00:12:43,779 --> 00:12:48,670
So now, lets play another game
that we can play because we

193
00:12:48,670 --> 00:12:51,130
know what the true distribution is.

194
00:12:51,130 --> 00:12:53,350
What is the true distribution
of the sample median?

195
00:12:53,350 --> 00:12:57,349
Well if I take 10,000 samples
of size 3,000 from that

196
00:12:57,349 --> 00:12:59,570
Chi-square distribution with
three degrees of freedom.

197
00:13:00,660 --> 00:13:04,120
And I compute the median for
each sample and I make a histogram.

198
00:13:04,120 --> 00:13:06,570
I'll plot it and
I get what's on the left there.

199
00:13:06,570 --> 00:13:10,620
And I put a Kernel density smoother on
there to make it look a little smoother.

200
00:13:11,800 --> 00:13:15,900
And I'm going to call that the true
distribution of the sample median for

201
00:13:15,900 --> 00:13:19,165
the purposes of what we're
talking about here because,

202
00:13:19,165 --> 00:13:23,040
10,000 is a pretty big number and
that's probably fairly close.

203
00:13:23,040 --> 00:13:27,080
So what I would like is,
that Theta hat is centered at

204
00:13:27,080 --> 00:13:32,220
the right value 2.365 which it
appears to be because I believe

205
00:13:32,220 --> 00:13:36,810
that was the value from the earlier page,
and that it has small variance.

206
00:13:36,810 --> 00:13:38,770
And it certainly does appear
to have small variance here.

207
00:13:38,770 --> 00:13:39,270
That's nice.

208
00:13:40,550 --> 00:13:45,000
Now it turns out, that, it's actually
easier to look at this centered and

209
00:13:45,000 --> 00:13:48,800
scaled version of the sample median,
rather than the raw version.

210
00:13:48,800 --> 00:13:52,800
And principally, that's because what I'd
really like to see straight away, is that,

211
00:13:52,800 --> 00:13:55,880
that distribution is
centered at the right value.

212
00:13:55,880 --> 00:13:58,450
And that it has a small variance and
that variance number,

213
00:13:58,450 --> 00:14:02,690
the way it's, shown there on the left,
is so small that I might not be able to,

214
00:14:02,690 --> 00:14:06,710
see it well enough, of course I
could have used more decimal spaces.

215
00:14:06,710 --> 00:14:09,740
But when I start comparing this to
other ways of estimating the sampling

216
00:14:09,740 --> 00:14:13,210
distribution of the sample median
it's going to be more convenient to

217
00:14:13,210 --> 00:14:17,360
look at this, centered and
scaled version of the statistic.

218
00:14:17,360 --> 00:14:18,510
Which I've plotted on the right and

219
00:14:18,510 --> 00:14:21,960
you can see the distribution pretty much
looks exactly the same except now it

220
00:14:21,960 --> 00:14:28,210
appears to be centered at nearly zero,
not quite exactly zero, but nearly zero.

221
00:14:28,210 --> 00:14:33,040
And it has a variance of 7.118 which
doesn't really mean a whole lot right at

222
00:14:33,040 --> 00:14:35,140
the moment, but it will in a minute.

223
00:14:35,140 --> 00:14:39,960
So, this is the distribution that
I assert we, this is the target.

224
00:14:39,960 --> 00:14:42,560
If we could reproduce this
distribution then we'd be happy.

225
00:14:45,440 --> 00:14:47,400
Okay so
now let's look at the ordinary Bootstrap.

226
00:14:48,890 --> 00:14:56,320
If I take 3,000 sample of size 3,000 and
compute a sample median for each one.

227
00:14:59,970 --> 00:15:00,810
From the, I'm sorry.

228
00:15:00,810 --> 00:15:05,860
From the one sample that I have,
we take the one sample that we have and

229
00:15:05,860 --> 00:15:10,430
we take 3,000 resamples and
compute the median for each one.

230
00:15:10,430 --> 00:15:12,440
Make a histogram of that distribution.

231
00:15:12,440 --> 00:15:15,250
What we get is the thing on the left and
it's centered and

232
00:15:15,250 --> 00:15:17,760
scaled version's the thing on the right.

233
00:15:17,760 --> 00:15:20,250
Now, one thing you might notice
is that in the centered and

234
00:15:20,250 --> 00:15:22,420
scaled version on the right.

235
00:15:22,420 --> 00:15:24,690
What I really would have
liked to have had there,

236
00:15:24,690 --> 00:15:28,390
where I have Theta hat is q50, q.50.

237
00:15:28,390 --> 00:15:33,740
But I don't know what q.50 is, so I am
actually approximating the centered and

238
00:15:33,740 --> 00:15:36,560
scaled distribu, Bootstrap
distribution of the sample median.

239
00:15:37,660 --> 00:15:42,580
With something that's not quite exactly
it because I substituted Theta hat.

240
00:15:42,580 --> 00:15:44,120
The value of the sampling median for

241
00:15:44,120 --> 00:15:49,450
the sample that I actually got,
in in place of q50, but hopefully this

242
00:15:49,450 --> 00:15:54,040
is a fairly good approximation,
at least in terms of the variability and

243
00:15:54,040 --> 00:15:59,130
the bias of the sampling distribution of
the sample median under the Bootstrap.

244
00:16:01,808 --> 00:16:03,120
okay, so if we look over on the right.

245
00:16:04,190 --> 00:16:09,750
We can see that this distribution is
centered at 0.257 which is rather

246
00:16:09,750 --> 00:16:14,975
far off of zero, or so it would seem,
and it has a variance of 9.063.

247
00:16:14,975 --> 00:16:18,630
So we could say that the performance
of the Bootstrap, in this case,

248
00:16:18,630 --> 00:16:22,060
I don't know whether you think
of that as being not so good or.

249
00:16:23,220 --> 00:16:26,880
Good it depends what, sort of what
you need for your application.

250
00:16:26,880 --> 00:16:30,630
But let's just remember these numbers
because this does appear to be biased and

251
00:16:30,630 --> 00:16:39,360
the variance is too high,
by a factor of two out of seven.

252
00:16:39,360 --> 00:16:43,380
Okay, so.
Not great but Better than nothing,

253
00:16:43,380 --> 00:16:44,580
I think we can say that.

254
00:16:45,620 --> 00:16:51,240
The Bootstrap does, work in a, in the
sense of being statistically consistent

255
00:16:51,240 --> 00:16:57,970
where things are, nice and you've probably
heard nice in your math classes before.

256
00:16:57,970 --> 00:17:01,670
Nice in this case means we're
working with a nicely behaved.

257
00:17:01,670 --> 00:17:06,886
Transformation g, or
form of Theta hat like the mean.

258
00:17:06,886 --> 00:17:11,074
And if the true process distribution
has a finite variance, or

259
00:17:11,074 --> 00:17:16,316
has some other, nice features about it and
I've listed some of them here.

260
00:17:16,316 --> 00:17:18,429
now, you know, we may want to do this for

261
00:17:18,429 --> 00:17:22,870
things like sample skewness and
of the sample correlation coefficient.

262
00:17:24,070 --> 00:17:29,100
Without knowing what F is, how we would
ever know that is have four, or six, or

263
00:17:29,100 --> 00:17:32,710
eight finite moments and
a positive variance I'm not sure.

264
00:17:32,710 --> 00:17:34,580
For positive variance it's probably okay.

265
00:17:34,580 --> 00:17:37,390
We can probably look at our sample,
and get a feel for that.

266
00:17:38,690 --> 00:17:41,010
But these other things
sort of take us back to.

267
00:17:42,160 --> 00:17:43,270
The problem we started with,

268
00:17:43,270 --> 00:17:46,900
which is if we don't need these,
if we don't know these things,

269
00:17:46,900 --> 00:17:50,630
then we might be playing with fire here
if we draw important conclusions from it.

270
00:17:50,630 --> 00:17:53,980
I will point out the last one,
which is that sample quantiles of

271
00:17:53,980 --> 00:17:57,960
which the median is one,
this does tend to work for the sample,

272
00:17:57,960 --> 00:18:02,899
quantiles but n has to be really big,
possible much, much bigger than 3,000.

273
00:18:04,610 --> 00:18:09,900
So if you weren't,
depressed enough by that already,

274
00:18:09,900 --> 00:18:13,250
then here's some more things to
be depressed about the Bootstrap.

275
00:18:14,290 --> 00:18:15,640
Which is here's when it doesn't work.

276
00:18:16,710 --> 00:18:19,000
These badly behaved.

277
00:18:19,000 --> 00:18:23,070
Transformations or sample statistic
calculations where, things aren't

278
00:18:23,070 --> 00:18:29,080
continuous or differentiable we said that
it does tend to work for the quantiles but

279
00:18:29,080 --> 00:18:33,510
N has to be very large because the, the
median is not a continuous function and

280
00:18:33,510 --> 00:18:35,810
N has to be very,
very big before we can overcome that.

281
00:18:37,680 --> 00:18:40,930
Things like sampling on the uniform
distribution, where the, the upper or

282
00:18:40,930 --> 00:18:44,530
lower limit of that uniform distribution
is the parameter we care about,

283
00:18:44,530 --> 00:18:46,840
that's tough.

284
00:18:46,840 --> 00:18:49,970
And rather unfortunately,
the, things go wrong and

285
00:18:49,970 --> 00:18:52,450
the same sorts of situations where
the central endotherm goes wrong.

286
00:18:53,510 --> 00:18:54,360
Now what can we do about it?

287
00:18:55,360 --> 00:18:59,768
There, there has been a new
development relatively recently,

288
00:18:59,768 --> 00:19:04,328
meaning in the last 20 years,
where we have something new called the M

289
00:19:04,328 --> 00:19:09,340
out of N Bootstrap which can fix some
of these problems for some situations.

290
00:19:09,340 --> 00:19:11,740
And I think,
we're going to look at that now.

291
00:19:11,740 --> 00:19:15,250
The M out of N Bootstrap is,
just like the ordinary Bootstrap.

292
00:19:15,250 --> 00:19:19,550
Except that we are not going to
resample N things with replacement,

293
00:19:19,550 --> 00:19:22,880
where N was the original sample size.

294
00:19:22,880 --> 00:19:26,410
We're going to sample fewer than N things.

295
00:19:26,410 --> 00:19:28,240
We're going to sample M things.

296
00:19:28,240 --> 00:19:32,640
And the condition is, that M needs
to be big but it can't be too big.

297
00:19:32,640 --> 00:19:36,210
It needs to, the ration M over N needs to
go to zero as both M and N go to infinity.

298
00:19:36,210 --> 00:19:42,370
And other than that, there's not a lot
of guidance as to, what M should be.

299
00:19:42,370 --> 00:19:45,320
So that's the main question here,
is how do you choose it?

300
00:19:45,320 --> 00:19:49,280
Rules of thumb, that obey this rule
are things like N to the half and

301
00:19:49,280 --> 00:19:50,930
N to the two-thirds.

302
00:19:50,930 --> 00:19:54,260
I've seen one suggestion that
2 times N to the half power,

303
00:19:54,260 --> 00:19:56,320
2 times the square root
of N is a good number.

304
00:19:56,320 --> 00:20:01,310
I, played around in the examples that
I'm going to show you a little bit and

305
00:20:01,310 --> 00:20:03,840
I was certainly able to find
a bunch of M's that didn't work.

306
00:20:04,910 --> 00:20:08,810
I found one,
that sort of did work if you recall,

307
00:20:08,810 --> 00:20:12,050
I didn't emphasize it but
the, sampling distribution,

308
00:20:12,050 --> 00:20:16,428
in other words the form of this histogram
and the density curve over plotted on it.

309
00:20:16,428 --> 00:20:20,810
The ordinary Bootstrap was pretty ugly
looking it seemed to have a couple of

310
00:20:20,810 --> 00:20:24,570
modes and
it really didn't look very pleasing.

311
00:20:24,570 --> 00:20:25,700
Here are the results for

312
00:20:25,700 --> 00:20:29,610
the M out of N Bootstrap I think I'll
direct your attention immediately to

313
00:20:29,610 --> 00:20:33,660
the right side of the page because
that's the one we can easily compare.

314
00:20:33,660 --> 00:20:38,370
And what we'll see is that while that
second mode hasn't been eliminated it

315
00:20:38,370 --> 00:20:40,440
has been somewhat reduced.

316
00:20:40,440 --> 00:20:43,490
So the distribution is looking
a little more symmetric.

317
00:20:43,490 --> 00:20:48,000
Which is a good thing, but
maybe not quite symmetric enough.

318
00:20:48,000 --> 00:20:52,440
The bias has been reduced to 0.111
here and the variance has been

319
00:20:52,440 --> 00:20:55,700
reduced considerably in fact,
the variance is now getting really.

320
00:20:55,700 --> 00:20:59,448
A bit closer to the 7.11,
I think it was, or, I believe it was

321
00:20:59,448 --> 00:21:05,460
about 7.11 was the variance of the true
sampling distribution of the median.

322
00:21:05,460 --> 00:21:06,790
So that's something of an improvement,

323
00:21:06,790 --> 00:21:08,790
at least in terms of the variance,
and the bias.

324
00:21:10,720 --> 00:21:15,360
Now I will say that that's just a cartoon
example, there are a zillion Bootstrap.

325
00:21:15,360 --> 00:21:17,410
Methods out there, designed for

326
00:21:17,410 --> 00:21:21,570
specific situations and
with specific kinds of improvements.

327
00:21:21,570 --> 00:21:24,470
And here's a list of just a few of them.

328
00:21:24,470 --> 00:21:28,060
probably, the one I'll direct
your attention to are these two,

329
00:21:28,060 --> 00:21:30,540
two top items on the right side.

330
00:21:30,540 --> 00:21:35,080
The block Bootstraps,
which are Bootstraps for dependent data.

331
00:21:35,080 --> 00:21:40,530
Where essentially instead of drawing,
individual items with replacement

332
00:21:40,530 --> 00:21:44,610
from the original sample to create
a resample, you draw things in blocks.

333
00:21:45,640 --> 00:21:47,590
So you're attempting to preserve,

334
00:21:47,590 --> 00:21:52,650
some amount of dependence within block
by drawing whole blocks at a time.

335
00:21:52,650 --> 00:21:55,830
But of course, the catch is there you
have to know what block size to use, and

336
00:21:55,830 --> 00:21:58,160
that is not an easy problem either.

337
00:21:58,160 --> 00:22:00,980
So you can check out in
any number of places.

338
00:22:00,980 --> 00:22:08,000
You can, search on these terms, and you'll
find a lot of information about these.

339
00:22:08,000 --> 00:22:11,790
Here are, sort of,
three books on the subject.

340
00:22:11,790 --> 00:22:15,088
An Introduction to the Bootstrap
by Brad Efron and Ron Tibshirani.

341
00:22:15,088 --> 00:22:18,406
[SOUND] That was almost the original book.

342
00:22:18,406 --> 00:22:20,980
That's a good place to learn things.

343
00:22:22,990 --> 00:22:26,370
I believe the methods that are discussed
in there are the basis of an R package.

344
00:22:26,370 --> 00:22:27,680
It's either Boot or Bootstrap.

345
00:22:28,712 --> 00:22:34,190
There's a more sophisticated book,
called Resampling Methods for

346
00:22:34,190 --> 00:22:38,350
Dependent Data which discusses the problem
of using the Bootstrap on depended data.

347
00:22:38,350 --> 00:22:39,280
It's quite mathematical.

348
00:22:39,280 --> 00:22:42,490
And the subsampling book which we're
going to talk about in the next module,

349
00:22:42,490 --> 00:22:43,980
we're going to talk about that technique.

350
00:22:43,980 --> 00:22:45,868
It's called subsampling.

351
00:22:45,868 --> 00:22:47,833
That's quite highly mathematical but

352
00:22:47,833 --> 00:22:51,640
truth be told, I found
the discussion in the first chapter.

353
00:22:51,640 --> 00:22:56,080
To be really interesting it described the
Bootstrap and then described this other

354
00:22:56,080 --> 00:23:00,090
method called subsampling and contrasted
them and I thought was pretty useful.

355
00:23:00,090 --> 00:23:03,910
So now we'll look at this
method called subsampling and

356
00:23:03,910 --> 00:23:07,600
see how we can do, subsampling is
related to the amount of in Bootstraps.

357
00:23:07,600 --> 00:23:08,240
We'll how we can do.

