1
00:00:00,850 --> 00:00:06,620
Welcome to the very last module
in the Inference and Uncertainty

2
00:00:06,620 --> 00:00:10,370
section of the JPL Caltech Virtual
Summer School on Big Data Analytics.

3
00:00:11,740 --> 00:00:16,298
This section is about
a technique called Subsampling.

4
00:00:16,298 --> 00:00:19,338
And I'll remind you with
this little outline slide.

5
00:00:20,772 --> 00:00:24,001
Where we are in the discussion.

6
00:00:24,001 --> 00:00:27,633
And this is the same introductory or
outline slide that we had at

7
00:00:27,633 --> 00:00:31,604
the beginning of the last section,
in which we discussed the bootstrap.

8
00:00:31,604 --> 00:00:33,430
And both the bootstrap and

9
00:00:33,430 --> 00:00:39,240
subsampling methods are nonparametric
techniques for approximating or

10
00:00:39,240 --> 00:00:44,552
estimating or understanding the sampling
distribution of a statistic that

11
00:00:44,552 --> 00:00:50,106
we want to use to estimate a process
distribution parameter of interest.

12
00:00:50,106 --> 00:00:55,400
When we don't know the analytical form or
the parametric form of

13
00:00:55,400 --> 00:01:01,493
the underlying probability distributions
that gave rise to the sample.

14
00:01:01,493 --> 00:01:05,391
And the, that is often the case
in a lot of applied works so

15
00:01:05,391 --> 00:01:09,292
these methods are not
surprisingly very popular and

16
00:01:09,292 --> 00:01:15,487
perhaps potentially very useful for
some of the things that we're doing here.

17
00:01:15,487 --> 00:01:19,391
I will not repeat
the asterisk at the bottom,

18
00:01:19,391 --> 00:01:23,311
which is an apology for
using some notation.

19
00:01:23,311 --> 00:01:25,928
I did that in the last module so
we'll just live with that.

20
00:01:25,928 --> 00:01:28,830
So this section is reasonably short.

21
00:01:28,830 --> 00:01:32,668
We're just going to introduce
subsampling and then I'm going to

22
00:01:32,668 --> 00:01:37,627
make a few comments and give you a few
cautions about using any of these methods.

23
00:01:37,627 --> 00:01:41,967
The material in this section is based
substantially on notes provided by

24
00:01:41,967 --> 00:01:46,308
Professor DasGupta in the Department
of Statistics at Purdue University and

25
00:01:46,308 --> 00:01:50,381
Professor Charlie Geyer at the School
of Statistics at the University of

26
00:01:50,381 --> 00:01:53,118
Minnesota and
on the book by Politis, Romano and

27
00:01:53,118 --> 00:01:57,696
Wolf called Subsampling, which I'll give
you the full reference to at the end.

28
00:01:57,696 --> 00:02:03,043
So remember in the bootstrap we
talked about drawing resamples from

29
00:02:03,043 --> 00:02:05,682
the one sample we actually have.

30
00:02:05,682 --> 00:02:10,852
Now we're going to do something slightly
different called drawing a subsample from

31
00:02:10,852 --> 00:02:12,905
the one sample we actually have.

32
00:02:12,905 --> 00:02:17,908
And a subsample is a series of
draws from the one sample we

33
00:02:17,908 --> 00:02:21,325
actually have without replacement.

34
00:02:21,325 --> 00:02:24,717
A resample was a draw with replacement,
and

35
00:02:24,717 --> 00:02:29,300
it was either a draw of the same
size as the original sample,

36
00:02:29,300 --> 00:02:33,792
namely capital N,
in the case of the ordinary bootstrap or

37
00:02:33,792 --> 00:02:38,984
a draw of some smaller number of
observations which we called m.

38
00:02:38,984 --> 00:02:41,776
But in either case in the bootstrap
it was with replacement and

39
00:02:41,776 --> 00:02:44,855
now we're drawing without replacement and
we're only drawing m.

40
00:02:44,855 --> 00:02:51,100
So let's call the subsample
v sub one to v sub m and

41
00:02:51,100 --> 00:02:58,620
the crucial difference here may not
seem obvious is that when we drew

42
00:02:58,620 --> 00:03:04,080
the resample with replacement, we were not
drawing from the original distribution f,

43
00:03:04,080 --> 00:03:07,870
we were drawing from the empirical
distribution of the sample f hats of n.

44
00:03:08,890 --> 00:03:12,460
Here we are drawing from
the true distribution f.

45
00:03:12,460 --> 00:03:17,280
And if you think about that, you think
you're drawing a subsample of a sample.

46
00:03:18,540 --> 00:03:22,440
If you had just gone out this morning
to draw the subsample directly, or

47
00:03:22,440 --> 00:03:26,310
a sample of the size of the subsample
directly from the population,

48
00:03:26,310 --> 00:03:28,990
you could've gotten this very collection.

49
00:03:28,990 --> 00:03:31,770
And you may be thinking to yourself,
but wait a minute.

50
00:03:31,770 --> 00:03:35,490
The only values I can get are the ones
that were in the original sample so

51
00:03:35,490 --> 00:03:37,330
how can that possibly be true?

52
00:03:37,330 --> 00:03:42,710
Well, the answer to that is that
everything we're doing here is

53
00:03:42,710 --> 00:03:48,360
treating the sample itself as random,
so all the probabilistic statements that

54
00:03:48,360 --> 00:03:52,169
we're making are implicitly averaging over
all possible samples we could have gotten.

55
00:03:53,190 --> 00:03:58,050
So now you're taking
a sample from a sample

56
00:03:58,050 --> 00:04:03,370
where the sample from which the secondary
sample is taken could be any

57
00:04:03,370 --> 00:04:06,649
one of a number of things that you could
have gotten from the true distribution f.

58
00:04:08,135 --> 00:04:10,700
So that tiny change of

59
00:04:10,700 --> 00:04:16,220
sampling without replacement rather than
sampling with replacement is huge here.

60
00:04:17,460 --> 00:04:22,200
So we now have v one to v m a sample
from the true distribution f.

61
00:04:22,200 --> 00:04:27,440
And we compute theta hat star as we did
with the bootstrap for each resample.

62
00:04:27,440 --> 00:04:32,410
We do that capital b times to obtain
a collection of theta hat stars.

63
00:04:32,410 --> 00:04:35,790
Which are estimates of the statistic of

64
00:04:35,790 --> 00:04:38,955
interest based on the resample
sorry on the subsamples.

65
00:04:40,230 --> 00:04:44,210
And we're going to estimate the CDF of
theta hat by this thing that I'm now

66
00:04:44,210 --> 00:04:51,840
calling l sub m l sub,
sub N comma sub, sub comma n.

67
00:04:51,840 --> 00:04:52,857
You know what I mean.

68
00:04:52,857 --> 00:04:55,121
And I'm going to talk right at the moment,

69
00:04:55,121 --> 00:04:58,996
now I'm going to talk only about
the centered and rescaled statistic.

70
00:04:58,996 --> 00:05:03,492
And I'm centering again here
around the value of the statistic,

71
00:05:03,492 --> 00:05:07,510
the value of theta hat computed
from the original sample.

72
00:05:08,730 --> 00:05:13,000
And I am looking at the distribution of
the theta hat stars and I'm going to do

73
00:05:13,000 --> 00:05:16,180
the same kind of Monte Carlo approximation
that we did with the bootstrap.

74
00:05:16,180 --> 00:05:21,406
So, again our question is do,
does that distribution function

75
00:05:21,406 --> 00:05:26,170
converge to the true distribution
function of theta hat?

76
00:05:27,310 --> 00:05:32,680
And the reason subsampling is
kind of another sort of minor

77
00:05:32,680 --> 00:05:39,220
miracle here is that this actually only
requires two relatively weak conditions.

78
00:05:39,220 --> 00:05:44,510
It requires that we know
that the centered and

79
00:05:44,510 --> 00:05:48,080
scaled version of theta hat converges, and

80
00:05:48,080 --> 00:05:53,030
here I'm going to say the word converges
in probability, which is one particular

81
00:05:53,030 --> 00:05:59,220
sense of convergence to some distribution
function that isn't pathological.

82
00:05:59,220 --> 00:06:03,350
In other words, it isn't point
mass at some one potential value.

83
00:06:03,350 --> 00:06:08,640
And that the scaling exponent there
alpha which tells us you might call

84
00:06:08,640 --> 00:06:12,780
this the rate, the rate at which n has
to get large in order to make this work.

85
00:06:14,200 --> 00:06:16,530
Is something just that it exists.

86
00:06:16,530 --> 00:06:19,310
We don't have to know which either
of these things are we just have to

87
00:06:19,310 --> 00:06:21,010
know that they exist.

88
00:06:21,010 --> 00:06:25,880
So both things seem like
relatively weak and

89
00:06:25,880 --> 00:06:29,769
general sorts of assumptions that we
might actually be justified in making.

90
00:06:30,890 --> 00:06:35,230
As with the m out of n bootstrap we
have to choose the subsample size and

91
00:06:35,230 --> 00:06:38,170
it has to be alike with
the M out of N bootstrap.

92
00:06:38,170 --> 00:06:42,840
M has to be big but
not too big relative to N.

93
00:06:42,840 --> 00:06:48,265
And if this is true you,
you can prove that

94
00:06:48,265 --> 00:06:54,210
the distribution function, the subsampling
distribution function converges

95
00:06:54,210 --> 00:06:57,780
in probability to the true distribution
function as N goes to infinity.

96
00:06:57,780 --> 00:07:00,950
Now there are a million,
I think there's a whole page of

97
00:07:00,950 --> 00:07:04,260
technical conditions that go
with this mathematically, and

98
00:07:04,260 --> 00:07:07,869
if you want to see them you can go to page
43 of the Politis, Romano and Wolf book.

99
00:07:10,510 --> 00:07:14,450
It's actually very interesting but
I think that the essence of it

100
00:07:14,450 --> 00:07:17,850
is if you're willing to assume that those
assumptions are not too restrictive, and

101
00:07:17,850 --> 00:07:22,690
in my opinion they aren't, then you can
just go ahead with this in this way.

102
00:07:22,690 --> 00:07:27,790
So what this means is that the subsampling
technique will quote unquote

103
00:07:27,790 --> 00:07:32,000
work in the assumption sense in
situations where the bootstrap won't.

104
00:07:32,000 --> 00:07:32,710
So it will work for

105
00:07:32,710 --> 00:07:37,700
stationary and dependent sequences and
extremes in a lot of

106
00:07:37,700 --> 00:07:42,440
places where the bootstrap either doesn't
work well or is really hard to use.

107
00:07:42,440 --> 00:07:44,970
So the catch once again
you have to choose M and

108
00:07:44,970 --> 00:07:46,580
you don't know exactly how to do that.

109
00:07:46,580 --> 00:07:48,230
You might have to play some games,

110
00:07:48,230 --> 00:07:51,730
so here's a cartoon
illustration of subsampling.

111
00:07:51,730 --> 00:07:55,860
For a really simple example of,
sample of sample of,

112
00:07:55,860 --> 00:08:01,810
original sample of size 5 like we had with
the bootstrap and a subsample size of 3.

113
00:08:01,810 --> 00:08:05,390
So here are the top line here is

114
00:08:05,390 --> 00:08:11,500
the original sample 2.82 is the original
median of the original sample.

115
00:08:11,500 --> 00:08:17,020
And then I've show v one to be vb as
the subsamples and their corresponding

116
00:08:17,020 --> 00:08:21,770
medians and I've replaced the empirical
distribution function on the left with

117
00:08:21,770 --> 00:08:26,540
a true distribution function of a pie
square three in order to make the point.

118
00:08:26,540 --> 00:08:27,998
And alluded to this already,

119
00:08:27,998 --> 00:08:31,218
yes it's true that it all depends
on the sample we actually have.

120
00:08:31,218 --> 00:08:37,048
But we are in the sense of of asserting
that this methodology quote unquote

121
00:08:37,048 --> 00:08:43,923
works the there's an averaging process
going on over all possible samples.

122
00:08:43,923 --> 00:08:48,165
And maybe this is a good place to say,
that no matter what you are doing,

123
00:08:48,165 --> 00:08:53,048
if you just got unlucky and got a lousy
sample that isn't really representative of

124
00:08:53,048 --> 00:08:56,820
the true process, then you've got trouble.

125
00:08:56,820 --> 00:08:58,240
And, you know,

126
00:08:58,240 --> 00:09:01,430
there's just, there's no way around
that in any statistical methodology.

127
00:09:01,430 --> 00:09:04,410
If all you know is the sample you have,
then you pretty much have to

128
00:09:04,410 --> 00:09:07,980
rely on that sample as having been more or
less typical.

129
00:09:07,980 --> 00:09:12,580
Now you can do something about
the uncertainty that you

130
00:09:12,580 --> 00:09:13,860
quote with your estimate.

131
00:09:13,860 --> 00:09:16,390
And that is in part what
these methods are doing.

132
00:09:16,390 --> 00:09:19,360
So, here are the same kind
of results we showed for

133
00:09:19,360 --> 00:09:23,600
the bootstrap in the M out of N bootstrap,
only for the subsampling version of it.

134
00:09:24,800 --> 00:09:27,260
I'll direct your attention
to the graph on the right.

135
00:09:27,260 --> 00:09:32,030
The PDF there, the estimated PDF of
the centered and scaled statistic,

136
00:09:32,030 --> 00:09:36,300
looks very similar to what we got
from the M out of N bootstrap and

137
00:09:36,300 --> 00:09:42,240
that's probably not surprising because
I set M and N and B all to be the same.

138
00:09:42,240 --> 00:09:46,810
And the technique is really rather
similar, it's just the only

139
00:09:46,810 --> 00:09:50,770
difference is pulling with replacement
versus pulling without replacement.

140
00:09:52,560 --> 00:09:56,680
And what we can see is that the bias
here has been reduced way down,

141
00:09:56,680 --> 00:09:58,170
it's now only .015.

142
00:09:58,170 --> 00:10:01,120
So it looks like that one difference,
resampling with and

143
00:10:01,120 --> 00:10:04,170
without replacement made
a difference to the bias.

144
00:10:04,170 --> 00:10:08,260
It also made a difference to the variance,
the variance has now become too small.

145
00:10:08,260 --> 00:10:12,994
So that's a danger, we did better on
the bias but worse on the variance.

146
00:10:12,994 --> 00:10:17,230
so, let me summarize
the results in case you

147
00:10:17,230 --> 00:10:20,450
can't remember all those numbers from
the previous pages which I can't either.

148
00:10:21,950 --> 00:10:26,970
So the top line of this table
is what the right answer is.

149
00:10:26,970 --> 00:10:31,210
At least according to my Monte Carlo
approximation with 10,000 trials.

150
00:10:31,210 --> 00:10:33,090
We would like it if our centered and

151
00:10:33,090 --> 00:10:38,020
scaled distribution had an expected value
of minus 0.49 and a variance of 7.118,

152
00:10:38,020 --> 00:10:42,010
and we see that the ordinary
bootstrap does not do very well.

153
00:10:42,010 --> 00:10:44,760
It does not do as well as
the M out of N bootstrap.

154
00:10:44,760 --> 00:10:48,740
The M out of N bootstrap actually
outperforms the subsampling with respect

155
00:10:48,740 --> 00:10:51,220
to the variance, but
not with respect to the main.

156
00:10:51,220 --> 00:10:56,480
So the question is, which one of these
things would be the right thing to use and

157
00:10:56,480 --> 00:11:00,820
the answer really depends on what your
problem is and on what you care about.

158
00:11:00,820 --> 00:11:04,220
If you're going to choose between
having a less biased estimate or

159
00:11:04,220 --> 00:11:07,270
a less variable estimate,
that's really a question, a,

160
00:11:07,270 --> 00:11:09,760
a decision has to be made in
the context of a particular problem.

161
00:11:11,010 --> 00:11:12,750
So I'm going to,

162
00:11:14,290 --> 00:11:19,050
come to the last slide here of this
whole set of lectures on inference and

163
00:11:19,050 --> 00:11:23,340
uncertainty and I'm going to quote and
paraphrase Dr. Charlie Geyer and

164
00:11:23,340 --> 00:11:26,380
I really like this because I think this
is actually pretty easy to remember.

165
00:11:26,380 --> 00:11:29,350
It says, the right thing is
samples of the right size,

166
00:11:29,350 --> 00:11:32,770
which is capital N from
the right distribution.

167
00:11:32,770 --> 00:11:36,080
The Romano and Politis thing which
is subsampling is samples of

168
00:11:36,080 --> 00:11:38,750
the wrong size from
the right distribution.

169
00:11:38,750 --> 00:11:42,590
The bootstrap thing is samples of the
right size from the wrong distribution.

170
00:11:42,590 --> 00:11:45,600
Both wrong things are wrong,
we would like to do the right thing but

171
00:11:45,600 --> 00:11:49,190
we can't because we really only
have one sample at our disposal.

172
00:11:49,190 --> 00:11:50,540
So which wrong thing do we want to do?

173
00:11:50,540 --> 00:11:53,880
And that goes to the comment of it
really depends on what your problem is.

174
00:11:53,880 --> 00:11:56,520
Whether you care more about bias or
more about variance.

175
00:11:56,520 --> 00:12:00,030
And I want to close the entire
set of little lectures by

176
00:12:00,030 --> 00:12:04,050
saying the following thing, that I
think these methods, subsampling and

177
00:12:04,050 --> 00:12:07,850
bootstrapping are really the way to go for
modern data analysis.

178
00:12:07,850 --> 00:12:09,861
They are computationally intensive.

179
00:12:09,861 --> 00:12:14,081
And certainly if you're talking about
sophisticated machine learning algorithms,

180
00:12:14,081 --> 00:12:16,966
which may take a long time to
run in the first place, might be

181
00:12:16,966 --> 00:12:21,033
a little difficult to do them, but I think
these things are on the right track.

182
00:12:21,033 --> 00:12:25,747
The caution is that what it means to
work is an isotonic guarantee and

183
00:12:25,747 --> 00:12:30,381
it's a mathematical guarantee
that requires assumptions that we

184
00:12:30,381 --> 00:12:33,230
don't necessarily know to be true.

185
00:12:33,230 --> 00:12:38,600
So, what I like to do is whenever I have a
real problem, I try to cook up a synthetic

186
00:12:38,600 --> 00:12:42,280
problem that contains the elements of
the real problem that I care about.

187
00:12:42,280 --> 00:12:45,140
And then I do a bunch of experiments
to try to see how much difference

188
00:12:46,200 --> 00:12:47,980
making different assumptions makes.

189
00:12:47,980 --> 00:12:52,730
And how sensitive the method itself is to
being wrong about certain assumptions.

190
00:12:52,730 --> 00:12:56,740
And that can help to decide which wrong
thing we want to do and perhaps what

191
00:12:56,740 --> 00:13:02,080
cautions we might want to issue along with
any conclusions that we draw particularly

192
00:13:02,080 --> 00:13:08,070
when it comes to people who might
be making policy decisions based

193
00:13:08,070 --> 00:13:14,820
on the conclusions that, that that we
have drawn from the data that we have.

194
00:13:14,820 --> 00:13:17,670
So with that I'd like to say thanks and
thanks for paying attention.

195
00:13:18,810 --> 00:13:22,270
And we'll see you in
the exercises session I believe.

