1
00:00:00,000 --> 00:00:02,079
And so I'm gonna talk about bootstrapping
now.

2
00:00:02,079 --> 00:00:05,064
And show how the bootstrapping algorithm
works.

3
00:00:05,064 --> 00:00:10,361
To do things like computing a standard
error for, say, the estimate of value of

4
00:00:10,361 --> 00:00:12,072
risk.
All right.

5
00:00:12,072 --> 00:00:17,082
So talking about the boot strap.
The boot strap is a very modern,

6
00:00:18,004 --> 00:00:24,021
innovation in statistics, and it's part of
the computer revelation in statistical

7
00:00:24,021 --> 00:00:27,071
analysis.
So twenty years ago, when you, wuh, you

8
00:00:27,071 --> 00:00:33,003
know, back when I was studying statistics
in college, I was actually an

9
00:00:33,003 --> 00:00:37,031
undergraduate Stat major at Berkeley.
I was a Stat-Econ double major.

10
00:00:37,031 --> 00:00:43,431
And, during that time actually the
bootstrap was just being developed so this

11
00:00:43,431 --> 00:00:47,609
was the middle 1980's.
And most of the theorems that were proved

12
00:00:47,609 --> 00:00:52,522
that justified the use of the bootstrap
was done in the say early 1990's.

13
00:00:52,522 --> 00:00:55,697
So it's a very modern statistical
technique.

14
00:00:55,697 --> 00:01:00,665
And something that is only feasible
because we now have fast, cheap computers,

15
00:01:00,665 --> 00:01:02,554
right.
So, you know, when I was a.

16
00:01:02,554 --> 00:01:07,222
I hate to say this when I started college
in 1982, you know, we were still.

17
00:01:07,222 --> 00:01:11,006
I mean in terms of computer technology, I
was the first.

18
00:01:11,006 --> 00:01:14,646
And let me see.
When I was taking my Computer Programming

19
00:01:14,646 --> 00:01:18,662
Class that was the first year that we
moved away from punch cards.

20
00:01:18,662 --> 00:01:22,900
They used to do computer programming by
not, you know, typing things into a

21
00:01:22,900 --> 00:01:25,962
terminal but literally you, you punched
holes in cards.

22
00:01:25,962 --> 00:01:28,870
And that, those it was the computer
instructions.

23
00:01:28,870 --> 00:01:33,154
And then you fed these cards into a
machine and then that compiled your

24
00:01:33,154 --> 00:01:36,748
program [laugh].
And if you made a mistake, you had to go

25
00:01:36,748 --> 00:01:39,800
and repunch a card.
Okay, so, I mean, I missed that by one

26
00:01:39,800 --> 00:01:40,610
year.
[laugh].

27
00:01:40,610 --> 00:01:43,226
So, I mean, and, and in the lec-, you
know?

28
00:01:43,226 --> 00:01:46,949
I mean it's amazing to see what, you know,
the advances in computer technology.

29
00:01:46,949 --> 00:01:50,723
If you go from essentially the beginning
of the 1980's until now.

30
00:01:50,723 --> 00:01:54,796
I mean, it's, it's quite incredible.
You know, what we're able to do

31
00:01:54,796 --> 00:01:58,940
essentially on your iPhone in terms of
computing, you know, is much, much more

32
00:01:58,940 --> 00:02:02,932
powerful than you know, most people could
have done on a, you know, a lot of

33
00:02:02,932 --> 00:02:06,681
computing platforms.
And so the bootstrap relies on being able

34
00:02:06,681 --> 00:02:12,066
to essentially simulate many, many, many
values you know, in a computer simulation.

35
00:02:12,066 --> 00:02:17,499
And because computers are fast and cheap
now, and we have programs like R, doing

36
00:02:17,499 --> 00:02:22,434
the bootstrap is, is completely trivial.
And to a large extent, the bootstrap

37
00:02:22,434 --> 00:02:27,190
eliminates the need to do all this, you
know, very technical complicated

38
00:02:27,190 --> 00:02:30,057
mathematics.
Used to be when you took statistics

39
00:02:30,057 --> 00:02:34,655
courses you just learned all these
formulas and have to grind out all this

40
00:02:34,655 --> 00:02:37,486
stuff.
Well you don't have to do that anymore.

41
00:02:37,486 --> 00:02:42,402
If you want things like a standard error,
you don't have to memorize a formula.

42
00:02:42,402 --> 00:02:46,357
All you need to do in, in practice is run
the bootstrap.

43
00:02:46,357 --> 00:02:52,317
So bootstrapping, one of the two great
computer simulation innovations in

44
00:02:52,317 --> 00:02:55,572
statistics.
The other being Markov Chain Monte Carlo

45
00:02:55,572 --> 00:03:00,594
Methods, which is a simulation technique
used in what's called basing statistics.

46
00:03:00,594 --> 00:03:03,951
The motivation behind bootstrapping is
like I said.

47
00:03:03,951 --> 00:03:09,080
Before modern computers doing statistical
analysis involved mathematics, and

48
00:03:09,080 --> 00:03:14,295
probability theory, to the derive formulas
for standard errors and confidence

49
00:03:14,295 --> 00:03:17,017
intervals.
Very often these formulas are

50
00:03:17,017 --> 00:03:22,492
approximations, their nasty looking
formula and often these approximations

51
00:03:22,492 --> 00:03:27,096
rely on the central limit theorem and
require having very large samples.

52
00:03:27,096 --> 00:03:30,069
Okay.
And the innovation is, with modern

53
00:03:30,069 --> 00:03:36,041
computers and statistical software like R,
The bootstrap, which is a, a subset of

54
00:03:36,041 --> 00:03:41,093
what's called a resampling method, can be
used to produce standard errors in

55
00:03:41,093 --> 00:03:45,036
confidence intervals without the use of
formulas.

56
00:03:46,009 --> 00:03:52,040
And they're often more reliable than the
formulas because they don't rely on a

57
00:03:52,040 --> 00:03:56,087
bunch of the assumptions that the formulas
are based on.

58
00:03:57,019 --> 00:04:02,090
So, the advantages of bootstrapping is you
know, many fewer assumptions about the

59
00:04:02,090 --> 00:04:06,013
underlying statistical behavior of the
data.

60
00:04:06,013 --> 00:04:11,035
For example, you don't, need to rely on
everything being normally distributed.

61
00:04:11,035 --> 00:04:16,099
The bootstrap works in a general context.
The bootstrap often gives you greater

62
00:04:16,099 --> 00:04:20,043
accuracy.
You don't have to have a large sample in

63
00:04:20,043 --> 00:04:24,069
order to use the bootstrap.
You can have a sample size of five.

64
00:04:24,069 --> 00:04:30,025
And the bootstrap works just the same way
as if you have a sample size of 20,000.

65
00:04:30,051 --> 00:04:34,778
So, you don't need to rely on the central
limit theorem, for example, to justify

66
00:04:34,778 --> 00:04:39,030
what you're doing.
And the boot strap has great generality.

67
00:04:39,030 --> 00:04:44,047
If you want to construct a bootstrap
standard error for the mean, and you want

68
00:04:44,047 --> 00:04:49,394
to construct a bootstrap standard error
for valiant risk, you do the same

69
00:04:49,394 --> 00:04:55,022
technique for both of the quantities.
There's not one bootstrap for the mean and

70
00:04:55,022 --> 00:04:58,055
another bootstrap for the valiant risk and
so on.

71
00:04:58,055 --> 00:05:03,045
The same bootstrapping algorithm works for
anything that you want to do.

72
00:05:03,045 --> 00:05:07,068
So it, it has a great generality.
Okay.

73
00:05:07,068 --> 00:05:13,027
So let's illustrate bootstrapping in the
constant expected return model.

74
00:05:13,027 --> 00:05:17,015
Now, so we have continuously compounded
returns.

75
00:05:17,015 --> 00:05:21,004
Have a mean and an error, which is our
random news.

76
00:05:21,004 --> 00:05:26,093
The random news is IID normal with mean
zero and a volatility sigma squared.

77
00:05:27,036 --> 00:05:32,037
So we have two parameters to estimate, the
mean and the variance, alright?

78
00:05:32,037 --> 00:05:37,086
And we wanna construct a standard air for
the mean, and a standard air for the

79
00:05:37,086 --> 00:05:40,070
variance.
And a 95 percent confidence interval for

80
00:05:40,070 --> 00:05:44,077
the mean and a 95 percent confidence
interval for the variants, okay?

81
00:05:44,077 --> 00:05:49,031
So, we have an observed sample.
So this is the data we download from

82
00:05:49,031 --> 00:05:52,297
Yahoo.
So, let's say we have 100 options, 100

83
00:05:52,297 --> 00:05:57,569
monthly returns on Microsoft or Starbucks,
and our goal is to compute the standard

84
00:05:57,569 --> 00:06:00,886
air of the mean, the standard air of the
volatility.

85
00:06:00,886 --> 00:06:05,528
And we wanna compute 95 percent confidence
intervals for, mu and sigma.

86
00:06:05,528 --> 00:06:08,072
Now.
The previous lecture, we derived analytic

87
00:06:08,072 --> 00:06:11,012
formulas for these standard errors.
Okay?

88
00:06:11,012 --> 00:06:15,009
We went through the math.
And it turns out that the standard error

89
00:06:15,009 --> 00:06:19,084
for the mean is the estimate of the
volatility / the square root of the sample

90
00:06:19,084 --> 00:06:21,064
size.
This was an exact result.

91
00:06:21,064 --> 00:06:26,057
And then we had an approximation to the
standard error based on the central limit

92
00:06:26,057 --> 00:06:31,056
theorem, that says, the standard error for
the volatility is the volatility / by two

93
00:06:31,056 --> 00:06:32,088
the sample size.
Okay?

94
00:06:32,088 --> 00:06:37,033
And if we want to construct a 95 percent
confidence interval, we can take estimate

95
00:06:37,033 --> 00:06:39,043
+ and - two the standard error.
Okay?

96
00:06:39,095 --> 00:06:43,099
Now, we just happen to know these
formulas, so we can use them.

97
00:06:43,099 --> 00:06:48,088
Now what we wanna do now is, use the
bootstrap to construct a bootstrap

98
00:06:48,088 --> 00:06:53,005
standard error for the mean.
A bootstrap standard error for the

99
00:06:53,005 --> 00:06:56,016
volatility.
And a bootstrap confidence interval.

100
00:06:59,083 --> 00:07:04,065
So how does bootstrapping work?
Turns out bootstrapping is very, very

101
00:07:04,065 --> 00:07:08,038
close to the Monte Carlo simulation that
we just did, okay.

102
00:07:08,038 --> 00:07:12,087
We ran a Monte Carlo simulation of the
constant expected return model.

103
00:07:12,087 --> 00:07:17,063
We specified values for the mean and the
volatility, and then we generated

104
00:07:17,063 --> 00:07:22,051
pseudo-data by, you know, using the
computer to generate normal data with the

105
00:07:22,051 --> 00:07:26,094
specified mean and the variance.
Bootstrapping works in a similar way

106
00:07:26,094 --> 00:07:31,648
except we're not going to simulate data
from a normal distribution.

107
00:07:31,648 --> 00:07:36,559
We're gonna resample with replacement from
the original data.

108
00:07:36,559 --> 00:07:40,885
So, we're gonna treat our original data as
if it were.

109
00:07:40,885 --> 00:07:46,370
Well we're gonna treat our original data
as the sort of like balls in a box.

110
00:07:46,370 --> 00:07:51,544
And yet you know in your stat class you
think about, you know, randomly drawing

111
00:07:51,544 --> 00:07:55,790
balls out of a box, the bootstrap works
exactly the same way.

112
00:07:55,790 --> 00:08:01,438
The balls are the observed data, and the
bootstrap sample is randomly picking

113
00:08:01,438 --> 00:08:05,745
observations out of our observed sample
and creating a new sample.

114
00:08:05,745 --> 00:08:08,787
Okay?
So, the bootstrapping algorithm, and this

115
00:08:08,787 --> 00:08:16,182
is what's known as non-parametric
bootstrapping.

116
00:08:16,182 --> 00:08:29,860
Alright, so what does the word
nonparametric mean?

117
00:08:29,860 --> 00:08:35,597
Anybody?
Or the, the obvious answer to something

118
00:08:35,597 --> 00:08:39,762
that's non-parametric, it's something
that's not parametric, right?

119
00:08:39,762 --> 00:08:47,453
[laugh] So what does parametric mean?
>> If I tell you I have a parametric model

120
00:08:47,453 --> 00:08:51,666
what does that mean?
>> I told you mounting something.

121
00:08:51,666 --> 00:08:53,397
[laugh] You know.
>> No.

122
00:08:53,397 --> 00:08:57,761
>> So a parametric model is any model that
has parameters, right.

123
00:08:57,761 --> 00:09:02,239
So the Constant Expected Return model has
two parameters.

124
00:09:02,239 --> 00:09:06,682
It has a mean parameter and it has a
variance parameter.

125
00:09:06,682 --> 00:09:12,997
So if you can write down a mathematical
formula for a model that has parameters,

126
00:09:12,997 --> 00:09:18,039
you have a parametric model.
And if something is non-parametric, well,

127
00:09:18,039 --> 00:09:20,381
you, you don't have a model.
Okay?

128
00:09:20,381 --> 00:09:26,281
Something that's non-parametric is just
based on the observed data.

129
00:09:26,281 --> 00:09:34,538
So in nonparametric bootstrapping, we're
not gonna generate our data from this

130
00:09:34,538 --> 00:09:39,537
model.
We're going to resample from the observed

131
00:09:39,537 --> 00:09:44,402
data.
So, what we're gonna do is we're gonna

132
00:09:44,402 --> 00:09:51,254
create capital B, bootstrap samples we're,
where we're gonna sample with replacement

133
00:09:51,254 --> 00:09:55,902
from the observed data.
And each bootstrap sample, is going to

134
00:09:55,902 --> 00:10:00,742
have capital T observations.
That's the same number of observations as

135
00:10:00,742 --> 00:10:05,275
the original sample.
>> And, I'm gonna denote a bootstrap

136
00:10:05,275 --> 00:10:12,585
sample, like this so in my first bootstrap
sample is gonna be say R11 to R1t so I

137
00:10:12,585 --> 00:10:19,367
have T observations and this, this one
here represents the, the first sam,

138
00:10:19,367 --> 00:10:24,203
bootstrap sample.
And the star means that its, its a

139
00:10:24,203 --> 00:10:29,560
randomly re-sampled observation from the
original data.

140
00:10:29,560 --> 00:10:36,641
Then remember our, our original sample is
R1 upto R capital T.

141
00:10:36,641 --> 00:10:40,401
So when we look at this first observation
here.

142
00:10:40,401 --> 00:10:45,745
So what we want to do is we want to think
of these observations like balls in an

143
00:10:45,745 --> 00:10:46,758
urn.
And we just.

144
00:10:46,758 --> 00:10:49,738
Stick our hand in and we pick out an
observation.

145
00:10:49,738 --> 00:10:54,646
And so we might pick out observation 22.
And that becomes this observation here.

146
00:10:54,646 --> 00:10:59,504
Then we do sampling with replacement.
So we take that observation, we put it

147
00:10:59,504 --> 00:11:03,944
back in the sample, right?
And then, you know, we reach in and we

148
00:11:03,944 --> 00:11:07,509
pull out a new observation.
Say it's observation 55.

149
00:11:07,509 --> 00:11:12,290
That becomes the second observation here.
And then we put it back in here.

150
00:11:12,290 --> 00:11:15,774
And then we reach in again, we pull out
another one.

151
00:11:15,774 --> 00:11:18,451
Until we get capital T observations.
Okay?

152
00:11:18,451 --> 00:11:23,634
So that's the idea of random sampling from
the observed, sample.

153
00:11:23,634 --> 00:11:28,204
Yeah.
>> What are, what one?

154
00:11:28,204 --> 00:11:32,867
>> Yeah.
So here, the first one represents

155
00:11:32,867 --> 00:11:35,313
bootstrap.
Sample.

156
00:11:35,313 --> 00:11:40,682
One, right?
And this is, so we're gonna create, say, a

157
00:11:40,682 --> 00:11:45,450
hundred samples of size P.
And this is my first sample.

158
00:11:45,450 --> 00:11:51,541
And so notice that there's a one subject
here that represents the, the first

159
00:11:51,541 --> 00:11:55,216
sample.
The capital B sub-script here, that my

160
00:11:55,216 --> 00:11:57,411
last bootstrap sample.
Okay?

161
00:11:57,411 --> 00:12:02,043
Now the second sub-script represents
observation numbers.

162
00:12:02,043 --> 00:12:09,472
So this is, sample observation number one,
sample ob, observation number two.

163
00:12:09,472 --> 00:12:13,483
Yeah, of course.
[inaudible], 'cause your sampling with

164
00:12:13,483 --> 00:12:16,232
replace.
So the question is, in a bootstrap sample,

165
00:12:16,232 --> 00:12:20,844
can I have duplicate observations?
And yes, you can, because you're sampling

166
00:12:20,844 --> 00:12:24,211
with replacement.
So you're not gonna rely on one sample.

167
00:12:24,211 --> 00:12:27,291
You're gonna have many, many, many
bootstrap samples.

168
00:12:27,291 --> 00:12:32,441
It could be that one of these you could
have, you know, the same observation could

169
00:12:32,441 --> 00:12:36,747
end up here 50 times, right?
That's just like flipping a coin and you

170
00:12:36,747 --> 00:12:40,529
get 50 heads, right?
It can happen, but it's very unlikely that

171
00:12:40,529 --> 00:12:41,678
it will.
But, yeah.

172
00:12:41,678 --> 00:12:45,948
>> How many permutations of the
observations can you make?

173
00:12:45,948 --> 00:12:48,466
>> T factorial.
>> T factorial, right.

174
00:12:48,466 --> 00:12:52,014
So if T is 100 you have 100 factorial
possible samples.

175
00:12:52,014 --> 00:12:56,197
That's a lot, right.
But if you have five observations you only

176
00:12:56,197 --> 00:13:00,769
have five factorial, right.
So the, the, the disadvantage of the

177
00:13:00,769 --> 00:13:04,712
bootstrap with a small sample size is you
only have a.

178
00:13:04,712 --> 00:13:08,394
Fine, you have really a finite number of
permutations.

179
00:13:08,394 --> 00:13:11,983
Right?
So even though the bootstrap works, it, it

180
00:13:11,983 --> 00:13:17,526
doesn't work as well if you get a sampled
size of three or four for example.

181
00:13:17,526 --> 00:13:22,305
So alright, so I so we just random
sampling with replacement from a

182
00:13:22,305 --> 00:13:28,564
population and we create our one pseudo
sample and, and the last sample.

183
00:13:28,564 --> 00:13:31,518
Now.
The rationale and the intuition about

184
00:13:31,518 --> 00:13:36,568
re-sampling from the observed data, is
well me again, think about what you're

185
00:13:36,568 --> 00:13:40,443
doing with statistics.
You don't know the model, that the

186
00:13:40,443 --> 00:13:42,391
underlying true model.
Right?

187
00:13:42,391 --> 00:13:45,945
But you observe the data.
The data comes from whatever that model

188
00:13:45,945 --> 00:13:49,029
is.
So if your, wanna create a representative

189
00:13:49,029 --> 00:13:54,587
sample, well re-sampling from the observed
data is a very intuitive way to create a

190
00:13:54,587 --> 00:13:57,700
representative sample from your unknown
model.

191
00:13:57,700 --> 00:14:04,626
So, So the distribution of each boot-strap
sample, is exactly the same as the

192
00:14:04,626 --> 00:14:08,036
probability distribution of your
observations.

193
00:14:08,036 --> 00:14:14,637
And so the idea of the bootstrap, when you
do the sampling to create your sample, you

194
00:14:14,637 --> 00:14:20,948
want the sampling mechanism that you're
going to use to pull observations out of

195
00:14:20,948 --> 00:14:26,497
the data to be the same as your underlying
assumptions about the probabilistic

196
00:14:26,497 --> 00:14:31,236
behaviour of your data.
So if you believe your data to be a random

197
00:14:31,236 --> 00:14:36,411
sample from a population.
Then when you do the bootstrap, you want a

198
00:14:36,411 --> 00:14:39,882
random sample from your observed,
observations.

199
00:14:39,882 --> 00:14:45,676
If you think your observations are a
covariant, stationary time series but

200
00:14:45,676 --> 00:14:49,360
that.
The each observation is uncorrelated with

201
00:14:49,360 --> 00:14:53,101
the rest.
Then you can still random sample from the

202
00:14:53,101 --> 00:14:58,942
population cause that would preserve the
uncorrellatedness in the data.

203
00:14:58,942 --> 00:15:07,031
If your sample observations are correlated
with one another, then when you sample

204
00:15:07,031 --> 00:15:11,340
from the bootstrap.
From your observed data you want to

205
00:15:11,340 --> 00:15:15,579
preserve that correlation.
So in that case you're going to sample in

206
00:15:15,579 --> 00:15:18,529
blocks.
So if you think R1 is correlating with R2,

207
00:15:18,529 --> 00:15:23,352
then your gonna wanna sample two adjacent
observations at a time for example.

208
00:15:23,352 --> 00:15:28,992
And so that the bootstrap would preserve
Whatever as- assumptions that you have in,

209
00:15:28,992 --> 00:15:32,498
in, in the data.
You have an autocorrelated series.

210
00:15:32,498 --> 00:15:37,986
The whole idea to do the bootstrap
correctly is you want the bootstrap to be

211
00:15:37,986 --> 00:15:41,289
able to capture the autocorrelation in the
data.

212
00:15:41,289 --> 00:15:45,781
And the procedure known as block
bootstrapping, that is boostrapping

213
00:15:45,781 --> 00:15:51,736
adjacent observations, and so you randomly
take out blocks of the data, will preserve

214
00:15:51,736 --> 00:15:54,772
the correlation structure in, in the
sample.

215
00:15:54,772 --> 00:16:01,458
So we're just gonna use the simple, Random
bootstrapping for the stuff in this

216
00:16:01,458 --> 00:16:06,511
course, but, you know, as we'll see in R,
they have functions for doing the block

217
00:16:06,511 --> 00:16:11,847
bootstrapping and other things like that.
Cuz when you're bootstrapping, so, here

218
00:16:11,847 --> 00:16:15,599
the sample is size T or it can be a sample
of size M.

219
00:16:15,599 --> 00:16:20,195
Your bootstrap sample is always the same
size as your observed sample.

220
00:16:20,195 --> 00:16:25,813
Right, cuz again you want to capture, I
mean the idea is, you are using the

221
00:16:25,813 --> 00:16:30,371
bootstrap to compute the, say a standard
error for the mean.

222
00:16:30,371 --> 00:16:34,483
Your mean estimate was computed on a
sample of size T.

223
00:16:34,483 --> 00:16:39,722
So, when you use the bootstrap.
You're gonna use the bootstrap to compute

224
00:16:39,722 --> 00:16:43,587
say a standard error.
You want the bootstrap sample to reflect

225
00:16:43,587 --> 00:16:46,357
exactly the same size sample as, as your
data.

226
00:16:46,357 --> 00:16:51,464
Its because you are using the bootstrap to
evaluate a statistic that is based on a

227
00:16:51,464 --> 00:16:56,691
sample of size T, so you want each of your
bootstrap samples to have exactly the same

228
00:16:56,691 --> 00:16:59,672
size.
If it has a different size, then it's not

229
00:16:59,672 --> 00:17:03,507
going to correspond to, you know, the
sample that you observe.

230
00:17:03,507 --> 00:17:08,066
The bootstrap sample that you observe is,
is randomly drawn from this.

231
00:17:08,066 --> 00:17:13,839
The distribution of the bootstrap sample
is exactly the same as the distribution of

232
00:17:13,839 --> 00:17:17,320
the observed data.
If the observed data is normal, your

233
00:17:17,320 --> 00:17:19,652
bootstrap sample is normal.
If your.

234
00:17:19,652 --> 00:17:22,734
Observe sample follows a gamma
distribution.

235
00:17:22,734 --> 00:17:26,022
Your bootstrap sample will follow a gamma
distribution.

236
00:17:26,022 --> 00:17:32,011
So the whole point about the bootstrap is,
whatever the distribution of your data is,

237
00:17:32,011 --> 00:17:36,589
your bootstrap sample will follow the same
distribution as that data.

238
00:17:36,589 --> 00:17:41,211
So it doesn't rely on assuming things are
normally distributed and so on.

239
00:17:41,211 --> 00:17:45,139
I don't know.
I once heard a statistician saying that

240
00:17:45,139 --> 00:17:50,779
the bootstrap is the closest thing to
magic in statistics that you'll find.

241
00:17:50,779 --> 00:17:56,518
Because, you know, you sort of do this and
you think, you know, wow this is so easy.

242
00:17:56,518 --> 00:18:01,063
But it turns out to be so powerful.
And, you know, and it's, it is.

243
00:18:01,063 --> 00:18:07,243
It's one of these where it was a, it was a
truly revolutionary concept in the area of

244
00:18:07,243 --> 00:18:09,397
statistics.
To make, you know?

245
00:18:09,397 --> 00:18:14,439
Our lives so much easier [laugh].
And we don't have to learn all of these

246
00:18:14,439 --> 00:18:18,407
horrible formulas anymore.
Alright so, so boo random sampling for a

247
00:18:18,407 --> 00:18:21,412
population.
Okay, so hopefully you understand what

248
00:18:21,412 --> 00:18:25,260
that means.
So once we have these bootstrap samples.

249
00:18:25,260 --> 00:18:29,492
What do we do with it?
Each bootstrap sample we, we calculate

250
00:18:29,492 --> 00:18:32,353
whatever statistic of interest that we
want.

251
00:18:32,353 --> 00:18:37,046
So if you want to compute a standard air
for the mean, than on each bootstrap

252
00:18:37,046 --> 00:18:40,765
sample we calculate the mean on the
bootstrap sample.

253
00:18:40,765 --> 00:18:45,700
And so then we'll have capital B.
Estimates of our mean.

254
00:18:45,700 --> 00:18:47,611
Okay.
And so now we have.

255
00:18:47,611 --> 00:18:52,453
So any statistical wants this could be we
can compute the mean.

256
00:18:52,453 --> 00:18:57,636
We can compute the standard deviation.
We can compute the value at risk.

257
00:18:57,636 --> 00:19:02,874
Anything that we can compute on the sample
is, could be theta, right?

258
00:19:02,874 --> 00:19:07,412
So, so now we have capital b value of our
statistics.

259
00:19:07,412 --> 00:19:10,511
Okay?
Now, what do you do with these capital B

260
00:19:10,511 --> 00:19:14,183
values of the statistic?
Well, we use the bootstrap distribution.

261
00:19:14,183 --> 00:19:17,860
We could, if, you know, for example,
suppose we wanna know, what is the

262
00:19:17,860 --> 00:19:21,660
probability distribution of our estimate?
Well, that's f of theta hat.

263
00:19:21,660 --> 00:19:24,706
We can just look at the histogram of the
bootstrap.

264
00:19:24,706 --> 00:19:28,793
That will give us an estimate of the
probability curve of our estimator.

265
00:19:28,793 --> 00:19:31,689
Right?
Does it follow a normal distribution?

266
00:19:31,689 --> 00:19:35,349
Well, if this looks like a normal
distribution, then it does.

267
00:19:35,349 --> 00:19:39,298
If it doesn't look like a normal
distribution, then it doesn't.

268
00:19:39,298 --> 00:19:44,327
If we wanna es-, evaluate bias of a
statistic, we look at the mean of the

269
00:19:44,327 --> 00:19:49,310
bootstrap relative to the sample mean that
we calculate from the actual data.

270
00:19:49,310 --> 00:19:52,228
So we can use the bootstrap to estimate
bias.

271
00:19:52,228 --> 00:19:59,689
If we wanna know the standard devia-, the
standard error of an estimator.

272
00:19:59,689 --> 00:20:05,624
Screen saver.
If you what to know a standard error of an

273
00:20:05,624 --> 00:20:12,292
estimator we just calculate the standard
deviation of theta hat on bootstrap

274
00:20:12,292 --> 00:20:13,346
samples.
'Kay?

275
00:20:13,346 --> 00:20:18,801
So, the two things that people are most
interested in, is an estimate of bias.

276
00:20:18,801 --> 00:20:24,153
So, a bootstrap estmate of bias.
So, I'll use a subscript boot to represent

277
00:20:24,153 --> 00:20:27,741
something calculated from the bootstrap
sample.

278
00:20:27,741 --> 00:20:32,898
So, what is, what is bias?
It's the expected value of the estimator

279
00:20:32,898 --> 00:20:36,561
minus the truth.
The bootstrap estimator of the bias.

280
00:20:36,561 --> 00:20:42,476
You take the bootstrap mean and you
subtract off the sample estimate and

281
00:20:42,476 --> 00:20:47,711
that's the bootstrap's estimate of the
bias of your estimator.

282
00:20:47,711 --> 00:20:51,605
Okay?
So it's bootstrap mean minus your sample

283
00:20:51,605 --> 00:20:54,931
mean.
If I want a bootstrap estimate of the

284
00:20:54,931 --> 00:20:58,562
standard error, well, what is the standard
error?

285
00:20:58,562 --> 00:21:05,551
The standard error of an estimate is the
standard deviation of the bootstrap values

286
00:21:05,551 --> 00:21:09,944
of your statistic.
So the sample standard deviation is one

287
00:21:09,944 --> 00:21:13,521
over the number of bootstrap samples minus
one.

288
00:21:13,521 --> 00:21:20,030
And we look at the Bootstrap estimate of
your statistic, minus the.

289
00:21:20,030 --> 00:21:25,260
Bootstrap mean, and then you square it,
and you sum over all the bootstrap.

290
00:21:25,260 --> 00:21:29,710
This is just the sample standard deviation
of your bootstrap values, 'kay?

291
00:21:29,710 --> 00:21:34,097
Notice that, 'kay, the bootstrap standard
error is just the sample standard

292
00:21:34,097 --> 00:21:37,268
deviation across the bootstrap.
This is trivial to compute.

293
00:21:37,268 --> 00:21:40,043
It doesn't matter how complicated your
theta hat is.

294
00:21:40,043 --> 00:21:45,239
For example, theta hat could be our value
at risk, and our bootstrap standard error

295
00:21:45,239 --> 00:21:49,369
of value at risk, is just a sample
standard deviation of value at risk

296
00:21:49,369 --> 00:21:54,642
computed on each of the bootstrap samples.
Alright.

297
00:21:54,642 --> 00:22:00,044
Now, what if we want to construct a 95
percent confidence interval using the

298
00:22:00,044 --> 00:22:05,049
bootstrap?
So, here there are sort of two approaches

299
00:22:05,049 --> 00:22:11,041
to take.
Remember, Last we, last time, if an s, we

300
00:22:11,041 --> 00:22:16,032
were talking about constructing a
proximate confidence interval as being

301
00:22:16,032 --> 00:22:19,054
estimate plus or minus two times the
standard error.

302
00:22:19,054 --> 00:22:24,026
That's justified if the underlying
probability curve of the estimator isn't

303
00:22:24,026 --> 00:22:27,078
normal.
Well, in the bootstrap, you can look at

304
00:22:27,078 --> 00:22:31,038
the histogram of your bootstrap values of
theta.

305
00:22:31,038 --> 00:22:37,032
If the bootstrap histogram looks like a
normal curve, then you can use our nice,

306
00:22:37,032 --> 00:22:43,048
easy formula estimate, plus and minus 2x
the bootstrap standard error, to compute a

307
00:22:43,048 --> 00:22:45,059
95 percent confidence interval.
Okay?

308
00:22:45,059 --> 00:22:50,085
If the bootstrap doesn't look like a
normal curve, so if the bootstrap

309
00:22:50,085 --> 00:22:55,074
histogram is very skewed, then this is
gonna be not very accurate.

310
00:22:55,074 --> 00:23:02,057
A more accurate confidence interval is
gonna be based on the quantiles of the

311
00:23:02,057 --> 00:23:05,029
bootstrap distribution.
So a 95%.

312
00:23:05,029 --> 00:23:09,050
This is called a percentile confidence
interval.

313
00:23:09,050 --> 00:23:14,076
You're gonna take the 2.5 percent lower
quantile, and the 97.5 percent upper

314
00:23:14,076 --> 00:23:18,044
quantile.
So the, this range contains 95 percent of

315
00:23:18,044 --> 00:23:21,024
the bootstrap observations.
Okay.

316
00:23:21,024 --> 00:23:26,042
That would be your estimate of your 95
percent confidence interval.

317
00:23:26,042 --> 00:23:34,042
And so again very easy to compute because
these, these are just sample quantiles

318
00:23:34,042 --> 00:23:41,031
from your Bootstrap values.
Now, let us talk little bit about actually

319
00:23:41,031 --> 00:23:46,060
doing the Bootstrap in R.
So, you have two choices, brute force,

320
00:23:46,060 --> 00:23:49,092
that is programming the boot strap by
hand.

321
00:23:49,092 --> 00:23:56,002
And doing the boot strap by hand by brute
force, the coding is exactly like the

322
00:23:56,002 --> 00:24:01,004
Monte Carlo simulation.
You write a fore loop and, and then you

323
00:24:01,004 --> 00:24:05,052
just resample from your observations
inside the fore loop.

324
00:24:05,052 --> 00:24:11,026
There is a R package called Boot.
As you'll find out, there's an R package

325
00:24:11,026 --> 00:24:17,079
for doing almost everything imaginable.
There are, in fact, almost 4,000 different

326
00:24:17,079 --> 00:24:18,097
R packages.
Okay.

327
00:24:18,097 --> 00:24:23,020
There's almost like 50 new R packages a
day, that show up.

328
00:24:23,020 --> 00:24:28,096
And one of the, the, the Ya know, sort of,
the most difficult things associated with

329
00:24:28,096 --> 00:24:33,099
r is just, sort of, keeping up with, you
know, what's all that, that's available.

330
00:24:33,099 --> 00:24:38,096
The boot package r is very good.
It's written by, two of the inventors of

331
00:24:38,096 --> 00:24:43,006
the bootstrap.
And, it's very reliable and, and easy to

332
00:24:43,006 --> 00:24:43,089
use.
Alright.

333
00:24:43,089 --> 00:24:49,031
So if we do the brute force version of
bootstraping, what do we have to do?

334
00:24:49,031 --> 00:24:53,033
So, you sample with replacement from the
original data.

335
00:24:53,033 --> 00:24:58,090
And, there's a nice arch command called
sample that allows you to random sample

336
00:24:58,090 --> 00:25:01,086
from a data matrix, or anything, any
object.

337
00:25:01,086 --> 00:25:07,015
So, if you have your data in the data
frame, or your data in the matrix, then

338
00:25:07,015 --> 00:25:11,038
you can create random samples from that
matrix using sample.

339
00:25:11,068 --> 00:25:19,002
And you do this capital B times.
And then, you wanna compute a statistic

340
00:25:19,002 --> 00:25:24,033
from the bootstrap, from each of these
bootstrap samples.

341
00:25:24,033 --> 00:25:28,038
Now, typically, we, we, capital B is set =
to 999.

342
00:25:28,038 --> 00:25:34,055
So usually, you know?
And the reason why you'd use 999 is, Well,

343
00:25:34,055 --> 00:25:39,052
one of the reasons is, when you calculate
quantiles in order to get, you know, sort

344
00:25:39,052 --> 00:25:43,006
of an exact quantile, you need an odd
number of observations.

345
00:25:43,006 --> 00:25:47,085
So if you want the median of a sample, you
can get a median, exact, but, you can get

346
00:25:47,085 --> 00:25:52,064
exactly the center of the distribution if
you have nine observations or some odd

347
00:25:52,064 --> 00:25:55,077
number of observations.
If you have an even number of

348
00:25:55,077 --> 00:25:58,038
observations, you can't actually get
halfway.

349
00:25:58,038 --> 00:26:02,063
So very often when you're doing
bootstrapping, that's why you see an odd

350
00:26:02,063 --> 00:26:05,024
number of values, 99, 999, something like
that.
