1
00:00:00,590 --> 00:00:02,630
Hi there, my name is Amy Braverman.

2
00:00:02,630 --> 00:00:05,840
I'm a statistician at
the Jet Propulsion Laboratory, and

3
00:00:05,840 --> 00:00:10,960
welcome to the JPL Caltech Virtual
Summer School on Big Data Analytics.

4
00:00:10,960 --> 00:00:12,980
And the modules here are on inference and
uncertainty.

5
00:00:14,540 --> 00:00:17,210
As I said, I'm a statistician, and so

6
00:00:17,210 --> 00:00:20,359
I'm very concerned with these two topics,
and I want to tell you a little bit about.

7
00:00:21,560 --> 00:00:24,400
How I think they're relevant for
big data analytics.

8
00:00:27,320 --> 00:00:32,160
So before we begin, I'd like to
just say a few things the goal of

9
00:00:32,160 --> 00:00:34,360
the modules is to present
statistical inference and

10
00:00:34,360 --> 00:00:37,900
it's role data analytics in
the simplest terms possible.

11
00:00:37,900 --> 00:00:40,710
There's obviously no way in
this short amount of time.

12
00:00:42,060 --> 00:00:46,760
That we would be doing the equivalent
of a full semester course in even

13
00:00:46,760 --> 00:00:50,560
a subset of the topics that actually
belong to inference and uncertainty.

14
00:00:50,560 --> 00:00:52,530
So the way I approached this
was to look back at a number

15
00:00:53,760 --> 00:00:56,950
of textbooks that I'm fond of and
to ask myself what I thought the key.

16
00:00:58,950 --> 00:01:05,410
Items were that I would want to impart
to you during this, set of lectures.

17
00:01:05,410 --> 00:01:10,280
And what that means is I tried to pick the
absolute fewest number of topics because I

18
00:01:10,280 --> 00:01:16,110
don't want to overwhelm anyone and
to present in a very heuristic, and

19
00:01:16,110 --> 00:01:18,839
intuitive way with
the minimum amount of math.

20
00:01:19,870 --> 00:01:23,160
That it is possible to even
discuss this stuff with.

21
00:01:23,160 --> 00:01:26,440
I do assume that everybody has a basic
knowledge of at least a little bit of

22
00:01:26,440 --> 00:01:31,100
calculus like knows what a derivative
is and you know what a limit is and

23
00:01:31,100 --> 00:01:35,510
I am going to, with apologies to my
mathematically oriented colleagues,

24
00:01:35,510 --> 00:01:38,460
I'm going to present this information.

25
00:01:38,460 --> 00:01:42,820
In a very conversational way, and without
stating all the possible conditions one

26
00:01:42,820 --> 00:01:47,340
might really want to check if one was
to approach this as a math problem.

27
00:01:47,340 --> 00:01:51,070
These lectures are not intended to
be either comprehensive or thorough.

28
00:01:51,070 --> 00:01:55,390
They are intended to be something of
a survey of what I regard as the key.

29
00:01:55,390 --> 00:01:59,420
Topics from, within this subject matter.

30
00:01:59,420 --> 00:02:01,950
And that said I have
three broad topic areas.

31
00:02:01,950 --> 00:02:07,340
I have review of basic probability, basic
concepts of inference, and an introduction

32
00:02:07,340 --> 00:02:11,160
to two popular nonparametric
procedures for performing inference.

33
00:02:12,560 --> 00:02:14,630
That you may or may not be familiar with.

34
00:02:14,630 --> 00:02:19,820
Some of the things that I discuss
in the basic probability lecture

35
00:02:19,820 --> 00:02:21,700
are therefore completeness.

36
00:02:21,700 --> 00:02:25,330
We will use some of them as we proceed
into the basic concepts of inference.

37
00:02:26,810 --> 00:02:32,500
And some of the basic material in the
probability section will not appear again,

38
00:02:32,500 --> 00:02:37,090
but it would be hard to leave it out
in any discussion of basic probability.

39
00:02:37,090 --> 00:02:41,600
So with that said, I'd like to move
on to the introduction where I'd

40
00:02:41,600 --> 00:02:45,760
like to try to give you a broad, broad
picture of what I'm trying to get at here.

41
00:02:46,840 --> 00:02:54,220
And I'm going to also, I define a few
terms for you that will be good to know.

42
00:02:54,220 --> 00:02:56,090
And that we will probably
end up using again later.

43
00:02:58,720 --> 00:03:01,510
Okay.
So this is my basic cartoon of what

44
00:03:01,510 --> 00:03:03,930
goes in statistical inference.

45
00:03:03,930 --> 00:03:05,780
We have an unknown population.

46
00:03:05,780 --> 00:03:09,260
Sometimes we might want to call
that a process, particularly.

47
00:03:09,260 --> 00:03:16,678
If we are looking at scientific,
scientific disciplines where we're talking

48
00:03:16,678 --> 00:03:22,510
about modeling physical,
physical, processes or,

49
00:03:22,510 --> 00:03:27,220
phenomena, and we understand that
process to be generating, Observations.

50
00:03:28,330 --> 00:03:31,690
And some of those observations
we actually get to

51
00:03:31,690 --> 00:03:33,910
see through the process of sampling.

52
00:03:33,910 --> 00:03:37,490
So the way I've drawn this, I've drawn
a little probability distribution,

53
00:03:37,490 --> 00:03:40,750
a cartoon probability distribution,
for the unknown population or process.

54
00:03:42,590 --> 00:03:45,280
Because we're going to use
probability models to model.

55
00:03:46,530 --> 00:03:50,280
How that process or
population behaves or looks like.

56
00:03:50,280 --> 00:03:54,090
And the, my little sampling cartoon
there is supposed to look like

57
00:03:54,090 --> 00:04:00,160
a screen that lets certain of those items
that are produced by the process through

58
00:04:00,160 --> 00:04:03,900
into the sample that you actually get
to see, or that we actually get to see.

59
00:04:03,900 --> 00:04:05,550
And typically what we do is we.

60
00:04:06,640 --> 00:04:11,740
Estimate something or draw a conclusion
from the sample that we actually have.

61
00:04:11,740 --> 00:04:14,920
And then we try to make a statement about

62
00:04:17,030 --> 00:04:20,620
some feature of the unknown process or
population that we

63
00:04:20,620 --> 00:04:23,040
care about based on what we learned
by interrogating the sample.

64
00:04:23,040 --> 00:04:26,530
And that's why written inference
with quantified uncertainty down

65
00:04:26,530 --> 00:04:27,100
there at the bottom.

66
00:04:29,930 --> 00:04:32,870
So we say that sampling supplies
us with realizations from

67
00:04:32,870 --> 00:04:35,310
the probability model that
describes the population.

68
00:04:36,410 --> 00:04:38,470
And what we have to recognize here,

69
00:04:38,470 --> 00:04:42,770
is that in the one sample that we actually
got, we could have gotten another sample.

70
00:04:42,770 --> 00:04:45,300
And we might have drawn
a different conclusion.

71
00:04:45,300 --> 00:04:49,540
From that other sample and
it is the that possibility,

72
00:04:49,540 --> 00:04:53,960
the possibility of getting more than
one different conclusion that leads to

73
00:04:53,960 --> 00:04:56,640
there being uncertainty in
whatever it is that we infer.

74
00:04:58,830 --> 00:05:02,630
Okay, so here's my little cartoon
picture of what's going on there.

75
00:05:02,630 --> 00:05:04,790
This is what we would really
like to be able to do.

76
00:05:04,790 --> 00:05:08,620
We would like to be able to collect
many samples from our population,

77
00:05:08,620 --> 00:05:12,730
compute on each one of them, and then look
at how different the results are, and

78
00:05:12,730 --> 00:05:16,540
I've put that over here on the right
with this little histogram with

79
00:05:16,540 --> 00:05:20,590
a density curve plotted on top of it, and
we call that the sampling distribution of

80
00:05:20,590 --> 00:05:23,850
the statistic that we computed,
and that will give us some idea.

81
00:05:23,850 --> 00:05:26,220
Of how different things could have been.

82
00:05:26,220 --> 00:05:32,440
Now, in, in the real world, we often
don't get to have more than one sample.

83
00:05:32,440 --> 00:05:34,220
We only get the one.

84
00:05:34,220 --> 00:05:36,970
So we're going to have come up with
some ways of understanding what that

85
00:05:36,970 --> 00:05:41,150
sampling distribution looks like
in order to make our inference.

86
00:05:41,150 --> 00:05:46,240
I will say since this is a summer school
on big data analytics that having our one

87
00:05:46,240 --> 00:05:51,450
sample be extremely large
actually gives us a way

88
00:05:51,450 --> 00:05:55,520
to obtain additional samples by
sampling from the one sample we have.

89
00:05:55,520 --> 00:05:58,150
And we'll come back to that at
the end of this little section and

90
00:05:58,150 --> 00:06:00,150
talk about that a little bit.

91
00:06:00,150 --> 00:06:04,650
Because that maybe a key for
leveraging the information in

92
00:06:04,650 --> 00:06:07,250
a massive data center or
bid data, as we're calling it.

93
00:06:09,210 --> 00:06:12,560
So we want to infer the prop, the
characteristics of the true probability

94
00:06:12,560 --> 00:06:14,360
model that describe our population.

95
00:06:14,360 --> 00:06:16,050
And we want to infer it from.

96
00:06:16,050 --> 00:06:20,000
Let's go back to the old fashion way,
from just one sample that we have.

97
00:06:20,000 --> 00:06:26,010
And we need to find a way to understand
not what the computed value is but

98
00:06:26,010 --> 00:06:29,550
where it sits in that notional
sampling distribution.

99
00:06:29,550 --> 00:06:32,620
And there are generally two
schools of thought on this.

100
00:06:32,620 --> 00:06:36,900
Statisticians tend to fall into either
the frequentists, sometimes also called,

101
00:06:36,900 --> 00:06:38,580
classical statistician.

102
00:06:38,580 --> 00:06:40,110
Camp, or the Bayesian Camp.

103
00:06:40,110 --> 00:06:43,300
And many of you may have
heard those terms before.

104
00:06:43,300 --> 00:06:45,070
And we will come back to that.

105
00:06:45,070 --> 00:06:47,740
I'm pretty much going to stick to
the frequentest point of view in

106
00:06:47,740 --> 00:06:52,530
these lectures, because, like I said, I'm
trying to be as uncluttered as possible.

107
00:06:52,530 --> 00:06:55,240
But I will say something about
Bayesian inference later on.

108
00:06:57,200 --> 00:06:59,100
So, now what I'd like to do is define,

109
00:06:59,100 --> 00:07:04,830
what I think are a few key terms that you,
may or may not have heard before.

110
00:07:05,980 --> 00:07:10,390
An observational unit is an object
about which we want to know something.

111
00:07:10,390 --> 00:07:13,460
So if I was interested in, my.

112
00:07:13,460 --> 00:07:17,120
Was interested in all the students
at Caltech, let's say.

113
00:07:17,120 --> 00:07:21,170
That an observational student
would be one student at Caltech.

114
00:07:21,170 --> 00:07:24,640
A variable is a quantity we
measure on an observational unit.

115
00:07:24,640 --> 00:07:29,300
So, let's say it's the height of
all the students at Caltech that I

116
00:07:29,300 --> 00:07:30,680
would like to know about.

117
00:07:30,680 --> 00:07:33,850
So I would measure height
on the observational units.

118
00:07:33,850 --> 00:07:37,530
The population is the entire
collection of observational units,

119
00:07:37,530 --> 00:07:41,780
that would be all students at Caltech,
and we should at

120
00:07:41,780 --> 00:07:45,900
this point distinguish between finite
populations and infinite populations.

121
00:07:45,900 --> 00:07:48,960
The example I just used,
finite populations.

122
00:07:48,960 --> 00:07:52,970
Would be all the students who are here
at Caltech today or all the U.S.

123
00:07:52,970 --> 00:07:54,570
citizens that are alive today.

124
00:07:54,570 --> 00:07:58,310
Because in principle we could actually
figure out how many of them there.

125
00:07:58,310 --> 00:08:00,490
And there would be
a countable number of them.

126
00:08:00,490 --> 00:08:04,570
But there's also a concept of an infinite
population, where we might talk about all

127
00:08:04,570 --> 00:08:08,320
citizens of the United States
that ever were or ever will be.

128
00:08:08,320 --> 00:08:10,780
Or all Caltech students there ever were or
ever will be.

129
00:08:10,780 --> 00:08:14,040
I'm not sure that's actually infinite,
but, you get my meaning.

130
00:08:15,610 --> 00:08:19,400
And the concept of an infinite population
is actually especially useful when we

131
00:08:19,400 --> 00:08:24,940
think about physical mechanisms, things
like surface temperature on the Earth.

132
00:08:24,940 --> 00:08:26,720
Surface temperature where you might ask.

133
00:08:26,720 --> 00:08:28,349
You might ask what's
the observational unit.

134
00:08:29,420 --> 00:08:30,170
Is it a point location?

135
00:08:30,170 --> 00:08:35,970
Or is it my block or my census tracked or

136
00:08:35,970 --> 00:08:40,830
some other larger aggregation of space,
that defines the observational unit.

137
00:08:40,830 --> 00:08:42,970
And there are many ways
to think about that.

138
00:08:42,970 --> 00:08:44,630
And if we really wanted to get down to it,

139
00:08:44,630 --> 00:08:47,810
and I know many of my physics
friends like to think this way.

140
00:08:47,810 --> 00:08:49,520
Space would be continuous.

141
00:08:49,520 --> 00:08:54,630
And we would be thinking about the process
that describes temperature at any

142
00:08:54,630 --> 00:08:58,190
arbitrary location in a,
let's say a two dimensional space.

143
00:08:59,300 --> 00:09:01,330
So that notion of
an infinite population and

144
00:09:01,330 --> 00:09:03,850
the notion of using
a probability distribution.

145
00:09:03,850 --> 00:09:08,790
To describe that infinite population
actually comes fairly close to the kind of

146
00:09:08,790 --> 00:09:10,080
thing that we often want to do in science.

147
00:09:11,710 --> 00:09:12,670
okay.
So

148
00:09:12,670 --> 00:09:17,149
a parameter is a quantity computed
from all the units in the population.

149
00:09:17,149 --> 00:09:19,820
All right.

150
00:09:19,820 --> 00:09:25,120
A sample is a subset of the population and
we may obtain it purposefully or.

151
00:09:26,230 --> 00:09:31,290
Coincidentally and a statistic is
a quantity computed from that sample.

152
00:09:31,290 --> 00:09:34,260
And note that the statistic
itself will be something that's

153
00:09:34,260 --> 00:09:36,830
random because the sample is random.

154
00:09:36,830 --> 00:09:39,760
The sample might have been different
therefore the calculated value of

155
00:09:39,760 --> 00:09:41,170
the statistic might have been different.

156
00:09:42,270 --> 00:09:45,150
I have this little side bar here about
an observational study versus and

157
00:09:45,150 --> 00:09:47,070
experiment because I think
that's an important distinction.

158
00:09:49,180 --> 00:09:53,540
Back in the old days when statistics
was young and one of the original

159
00:09:53,540 --> 00:09:58,220
applications was to understand
the impact of different treatments on,

160
00:09:58,220 --> 00:10:00,380
let's say, how well your corn grows.

161
00:10:01,650 --> 00:10:03,840
People did what were called experiments.

162
00:10:03,840 --> 00:10:06,270
They would divide up their
plot in a particular way and

163
00:10:06,270 --> 00:10:07,850
they would add fertilizer and

164
00:10:07,850 --> 00:10:12,010
water in different combinations, and then
they would measure the height of the corn.

165
00:10:12,010 --> 00:10:15,180
And try to determine which is the best
combination of fertilizer and

166
00:10:15,180 --> 00:10:17,210
water to get the best corn crop.

167
00:10:17,210 --> 00:10:22,090
That's an experiment because the
observational units were actually being

168
00:10:22,090 --> 00:10:25,320
changed and affected by that experiment.

169
00:10:25,320 --> 00:10:29,280
So I'm calling that si,
circumstance an experiment.

170
00:10:29,280 --> 00:10:33,450
In the kind of stuff that we do at JPL,
we have a some what different situation.

171
00:10:33,450 --> 00:10:38,240
We, let's say we fly a satellite
that looks down and measures or

172
00:10:38,240 --> 00:10:41,860
observes some characteristics
of the Earth's surface.

173
00:10:42,870 --> 00:10:47,080
The satellite flies at a prescribed
orbit that is not under our control.

174
00:10:47,080 --> 00:10:50,150
It may have been under control of
the people who designed the mission but.

175
00:10:50,150 --> 00:10:52,910
At the point where we enter the picture
it's not under our control.

176
00:10:52,910 --> 00:10:56,500
And typically that satellite flies in
an orbit that has some regularity.

177
00:10:57,530 --> 00:10:59,810
So, that's what we call
an observational study.

178
00:10:59,810 --> 00:11:06,690
We do not have the option of actually
experimenting on the experimental units or

179
00:11:06,690 --> 00:11:09,510
on the observational units
that we want to understand.

180
00:11:09,510 --> 00:11:11,590
We simply observe what's there.

181
00:11:11,590 --> 00:11:16,180
So that's an important,
an important distinction to keep in mind.

182
00:11:16,180 --> 00:11:19,800
Let me now say something about exploratory
versus confirmatory analysis because I

183
00:11:19,800 --> 00:11:25,570
think this is where, I'd like to
distinguish between this module and

184
00:11:25,570 --> 00:11:27,200
some of the other modules
that you might have heard,

185
00:11:27,200 --> 00:11:32,410
particularly from my machine
learning colleagues, you may

186
00:11:32,410 --> 00:11:35,730
have heard the terminology
exploratory vs confirmatory analysis.

187
00:11:35,730 --> 00:11:43,040
Exploratory data analysis sometimes called
EDA is a popular term, and that pertains

188
00:11:43,040 --> 00:11:46,760
to taking the data that you have, as I've
shown it here, the data from the sample.

189
00:11:47,820 --> 00:11:50,960
And exploring it as the name says.

190
00:11:50,960 --> 00:11:55,150
Maybe making plots, looking at it,
computing some summary statistics, or

191
00:11:55,150 --> 00:11:59,300
perhaps even doing something extremely
sophisticated like using a machine

192
00:11:59,300 --> 00:12:03,980
learning algorithm such as the one Dave
Thompson might have talked about to you.

193
00:12:03,980 --> 00:12:06,870
But in any case you
are applying that algorithm to

194
00:12:06,870 --> 00:12:09,290
the data you have in front of you.

195
00:12:09,290 --> 00:12:13,430
And that can be a daunting tak because
those data sets can be large and

196
00:12:13,430 --> 00:12:19,550
there's a lot of clever things that can be
done there but in the end at some point I

197
00:12:19,550 --> 00:12:25,080
think that our objective will be to draw a
conclusion about the population from which

198
00:12:25,080 --> 00:12:29,370
those data were obtained and that's the
job of confirmatory analysis or inference.

199
00:12:30,390 --> 00:12:34,530
Where we have to understand
the uncertainty of the result that we

200
00:12:34,530 --> 00:12:36,940
obtain from the one sample that we have.

201
00:12:36,940 --> 00:12:41,640
So let me go here,
EDA illuminates structures, patterns,

202
00:12:41,640 --> 00:12:43,650
relationships, and so forth in the sample.

203
00:12:43,650 --> 00:12:48,930
EDA is often necessary to formulate
hypotheses about the unknown population.

204
00:12:48,930 --> 00:12:52,770
One of our major problems in
big data is that we often don't

205
00:12:52,770 --> 00:12:56,570
actually know what's in the data, we don't
really know much about it at all, and

206
00:12:56,570 --> 00:12:59,490
in order to form a hypothesis that
we might want to test later, or

207
00:12:59,490 --> 00:13:03,010
to know what's important,
we have to go through the process of

208
00:13:03,010 --> 00:13:06,230
understanding what's in at least
the sample in front of us.

209
00:13:06,230 --> 00:13:09,610
In confirmatory data analysis, we use
the tools of statistical inference to

210
00:13:09,610 --> 00:13:12,760
make a definitive probabilistic
statements about the population,

211
00:13:12,760 --> 00:13:15,030
based on what we learned from the sample.

212
00:13:15,030 --> 00:13:20,060
And I'll throw out the notion that there
are two tools of statistical inference.

213
00:13:20,060 --> 00:13:22,970
That turns out that these
are really flip sides of

214
00:13:22,970 --> 00:13:26,040
the same thing hypothesis testing.

215
00:13:26,040 --> 00:13:29,200
And estimation and I'm not going to
talk about hypothesis testing in

216
00:13:29,200 --> 00:13:30,660
these lectures, just about estimation.

217
00:13:33,250 --> 00:13:38,540
So, let's get back to massive datasets or
big data and I allude at the beginning of

218
00:13:38,540 --> 00:13:44,650
this module, sometimes the sample
that we have is too big to treat

219
00:13:44,650 --> 00:13:47,890
like our friends from a hundred years
ago that would like to treat a sample.

220
00:13:47,890 --> 00:13:50,980
Sometimes it's even too big
to get on to our computer.

221
00:13:50,980 --> 00:13:57,210
And so there comes a question about
how to interrogate that sample for

222
00:13:57,210 --> 00:13:59,950
EDA purposes and
then how to make inferences from it.

223
00:14:01,430 --> 00:14:04,360
There are in general two strategies for
that.

224
00:14:04,360 --> 00:14:05,050
One is bigger,

225
00:14:05,050 --> 00:14:10,380
better, faster algorithms and
machine learning provides a lot of those.

226
00:14:10,380 --> 00:14:13,950
And the other strategy is
to make the data smaller,

227
00:14:13,950 --> 00:14:18,830
where you might sample again,
from the big data sample that you have.

228
00:14:18,830 --> 00:14:22,180
Or you might choose to
apply some algorithm that

229
00:14:23,660 --> 00:14:25,990
reduces it in some way to
a more manageable form.

230
00:14:27,130 --> 00:14:30,620
And, I will the machine learning
algorithms to the machine learning

231
00:14:30,620 --> 00:14:33,570
experts, and I'm not really going to
talk about data reduction, although,

232
00:14:33,570 --> 00:14:34,860
we're allude to it again later.

233
00:14:35,880 --> 00:14:39,440
So finally,
let's talk about the question of

234
00:14:39,440 --> 00:14:42,840
whether massive data sets
are populations or samples.

235
00:14:42,840 --> 00:14:44,510
They're kind of both actually.

236
00:14:45,690 --> 00:14:51,230
They are samples in the sense that they
are a random selection of some kind

237
00:14:51,230 --> 00:14:56,690
from a larger population or process but
they're also like populations in that

238
00:14:56,690 --> 00:15:00,370
we have to do something else, we have
to sample from them again, we have to.

239
00:15:00,370 --> 00:15:02,740
We know that we can
interrogate them completely and

240
00:15:02,740 --> 00:15:07,060
understand them completely in the way we
might have treated the smaller samples.

241
00:15:07,060 --> 00:15:10,930
So, here we have an option and I,
as I said I alluded to this earlier.

242
00:15:10,930 --> 00:15:14,840
To perhaps sample again from
our big data sample and now we

243
00:15:14,840 --> 00:15:18,790
could sample many times and we could
actually make that histogram on the right.

244
00:15:18,790 --> 00:15:24,380
By computing the statistic of interest
over those samples-the secondary samples.

245
00:15:24,380 --> 00:15:26,150
And, wouldn't it be nice if we could,

246
00:15:26,150 --> 00:15:31,440
some how, compute the thing we really
wanted to know on the big data sample, and

247
00:15:31,440 --> 00:15:34,119
then look at the sampling distribution
that we have on the right.

248
00:15:35,140 --> 00:15:40,330
And from that relationship, somehow infer
what the relationship between the computed

249
00:15:40,330 --> 00:15:44,400
value from the big data sample was and
the true population.

250
00:15:44,400 --> 00:15:46,210
And people are working on methods for

251
00:15:46,210 --> 00:15:49,070
doing things like that and
thinking about that.

252
00:15:49,070 --> 00:15:53,750
I will not go that far today, but
it's out there in the literature.

253
00:15:53,750 --> 00:15:56,490
There's something developed up
at Berkley called the bag of

254
00:15:56,490 --> 00:16:00,130
little boot straps which is
starting to get in this direction.

255
00:16:00,130 --> 00:16:01,970
And you may want to go
have a look at that.

256
00:16:01,970 --> 00:16:06,880
So with that I think we'll
move on to the first module on

257
00:16:06,880 --> 00:16:11,580
probability which is we will just do
a brief review of probability theory.

