1
00:00:00,950 --> 00:00:04,720
Hi there, I'm Amy Braverman from JPL and

2
00:00:04,720 --> 00:00:09,190
this is,
a module on a review of basic probability.

3
00:00:09,190 --> 00:00:13,730
In preparation, for
a discussion of inference and

4
00:00:13,730 --> 00:00:16,010
uncertainty in big data analytics.

5
00:00:17,400 --> 00:00:19,470
So here's the outline
of this short section,.

6
00:00:20,940 --> 00:00:24,020
Some of this maybe beneath many of you.

7
00:00:24,020 --> 00:00:25,720
But maybe not.

8
00:00:25,720 --> 00:00:29,560
So just proceed and see what happens.

9
00:00:29,560 --> 00:00:31,110
I'm going to talk about
the following topics.

10
00:00:31,110 --> 00:00:32,290
What is probability?

11
00:00:32,290 --> 00:00:34,280
Sample spaces and events.

12
00:00:34,280 --> 00:00:38,810
The axioms of probability and
some of those corollaries of those axioms.

13
00:00:38,810 --> 00:00:41,229
And joint and conditional probabilities.

14
00:00:43,280 --> 00:00:46,920
I think we all have an idea in our
minds about what probability is.

15
00:00:46,920 --> 00:00:50,310
Probability it will rain,
probability if I throw to dice,

16
00:00:50,310 --> 00:00:53,440
I'll get two sixes,
probability that the free way will be

17
00:00:53,440 --> 00:00:55,400
jammed this morning when I
get on it to go to work.

18
00:00:55,400 --> 00:00:58,710
Or the probability that
global warming is real.

19
00:01:00,940 --> 00:01:05,610
Historically, and classically
probability has been defined as long run

20
00:01:05,610 --> 00:01:10,790
relative frequency which conforms, I
think, to those examples that I just gave.

21
00:01:11,870 --> 00:01:14,300
The canonical example is flipping a coin.

22
00:01:14,300 --> 00:01:18,130
If I flip a coin one time,
what's the probability I get a head?

23
00:01:18,130 --> 00:01:19,580
Well, it's a half.

24
00:01:19,580 --> 00:01:24,170
And one might say that the reason we
think that is because, I believe that if

25
00:01:24,170 --> 00:01:28,290
I flipped a coin many, many, many times,
I should get about heads half the time.

26
00:01:28,290 --> 00:01:33,440
And I believe that because I believe that
the coin is physically built to come up

27
00:01:33,440 --> 00:01:35,480
either head or tail with equal.

28
00:01:36,930 --> 00:01:38,110
Long run relative frequency.

29
00:01:40,110 --> 00:01:45,440
probability, fortunately,
is also a mathematical object.

30
00:01:45,440 --> 00:01:50,050
We can say that it's a set function in
that it assigns a number, in this case,

31
00:01:50,050 --> 00:01:52,270
a number between zero and one to a set.

32
00:01:54,990 --> 00:01:56,430
So.

33
00:01:56,430 --> 00:01:59,560
Let's, we're going to need
a couple of definitions before we,

34
00:01:59,560 --> 00:02:01,500
before we proceed with that.

35
00:02:01,500 --> 00:02:04,180
Let's define a trial,
to be a measurement or

36
00:02:04,180 --> 00:02:07,110
observation of some outcome,
of some phenomenon.

37
00:02:08,300 --> 00:02:11,850
The set of all possible outcomes of
a trial is called the sample space and

38
00:02:11,850 --> 00:02:14,850
it's traditionally denoted by a capital S.

39
00:02:14,850 --> 00:02:18,650
Examples are the possible
sexes of a newborn baby.

40
00:02:18,650 --> 00:02:22,740
That would be a boy or a girl or the
minutes, number of minutes I have to wait

41
00:02:22,740 --> 00:02:27,190
to get on the freeway to go to work this
morning and that would be a number between

42
00:02:27,190 --> 00:02:30,820
zero and infinity tending toward infinity
if you live here in Southern California.

43
00:02:32,190 --> 00:02:35,060
An event is a subset of the sample space.

44
00:02:35,060 --> 00:02:41,080
So event e might be the event that
the newborn baby was a boy or

45
00:02:41,080 --> 00:02:45,080
the event that I was able to get on
the freeway in five minutes in less.

46
00:02:48,510 --> 00:02:51,450
So sample spaces and events are sets.

47
00:02:51,450 --> 00:02:56,120
And people are very fond of depicting
sets and set operations with

48
00:02:56,120 --> 00:02:59,750
these things called Venn diagrams,
which probably everyone has seen.

49
00:02:59,750 --> 00:03:01,600
We represent the whole sample space, S,

50
00:03:01,600 --> 00:03:06,050
by the square black rectangle
with the white background.

51
00:03:06,050 --> 00:03:10,300
And we find subsets of that sample
space that represent certain events.

52
00:03:10,300 --> 00:03:12,720
So here I have the even A,
and the event B.

53
00:03:12,720 --> 00:03:15,550
Event A is the red one,
the event B is the blue one.

54
00:03:15,550 --> 00:03:20,080
And I'm going to define certain
operations on those two set.

55
00:03:20,080 --> 00:03:23,310
For example the union of A and
B show in the upper right.

56
00:03:24,750 --> 00:03:28,820
are, all the elements of the sample
space that belong to both A and B.

57
00:03:28,820 --> 00:03:29,750
I'm sorry, A or B.

58
00:03:29,750 --> 00:03:34,120
I make the mistake that I was
about to tell you not to make.

59
00:03:34,120 --> 00:03:39,350
Or means that the, that the event
that you're interested in belongs to.

60
00:03:40,350 --> 00:03:41,570
A or B.

61
00:03:42,730 --> 00:03:44,990
The intersection is
shown in the lower left.

62
00:03:44,990 --> 00:03:47,990
I'm sure you're all familiar with that.

63
00:03:47,990 --> 00:03:50,210
That's the end operator.

64
00:03:50,210 --> 00:03:54,120
That says the element of the sample
space is a member of both A and

65
00:03:54,120 --> 00:03:55,940
B simultaneously.

66
00:03:55,940 --> 00:03:56,870
And over the right,

67
00:03:56,870 --> 00:04:00,750
I tried to depict the give you
an idea of what the complement is.

68
00:04:00,750 --> 00:04:03,200
The complement means not.

69
00:04:03,200 --> 00:04:09,640
And in my little example here,
what I've shown you here, is B and not A.

70
00:04:09,640 --> 00:04:10,880
So hopefully that is clear, and

71
00:04:10,880 --> 00:04:15,460
hopefully that is not the first time
you've ever seen anything like that.

72
00:04:15,460 --> 00:04:18,490
so, set theory offers us many

73
00:04:20,990 --> 00:04:27,280
Logical consequences of the mathematical
formulation of what I just explained.

74
00:04:27,280 --> 00:04:30,900
And there are some laws of set operations,
which are fairly basic,

75
00:04:30,900 --> 00:04:33,220
which you've probably seen before as well.

76
00:04:33,220 --> 00:04:35,040
The Commutative Law,
the Associative Law and

77
00:04:35,040 --> 00:04:38,360
the Distributive Law,
the Distributive Law, excuse me.

78
00:04:38,360 --> 00:04:41,560
And these things at the bottom,
called DeMorgan's Laws.

79
00:04:41,560 --> 00:04:45,950
Which essentially says that
the complement of the union of a bunch of

80
00:04:45,950 --> 00:04:51,180
disjoint sets is equal to
the intersection of the complements.

81
00:04:51,180 --> 00:04:52,890
And similarly, the other way around.

82
00:04:54,270 --> 00:04:57,580
These are one of those things that we
may or may not end up using later but

83
00:04:57,580 --> 00:05:00,759
it has to be said or I would
probably be run out of the business.

84
00:05:01,870 --> 00:05:05,140
Okay.
So there are three axioms of probability.

85
00:05:05,140 --> 00:05:06,800
Only three things that we have to assume.

86
00:05:08,380 --> 00:05:10,770
Probabilities are numbers between zero and
one inclusive.

87
00:05:10,770 --> 00:05:15,620
The probability of the entire
sample space is one meaning

88
00:05:15,620 --> 00:05:17,830
something must happen on your trial.

89
00:05:19,030 --> 00:05:23,400
And Axiom three says that, for
any sequence of disjoint events,

90
00:05:23,400 --> 00:05:27,080
in other words, events which have
an intersection that's empty,

91
00:05:27,080 --> 00:05:30,190
they do not overlap,
in terms of those Venn diagrams.

92
00:05:30,190 --> 00:05:32,820
The probability of
the union of those events,

93
00:05:32,820 --> 00:05:35,980
is equal to the sum of their
individual probabilities.

94
00:05:35,980 --> 00:05:39,600
All other rule of probability
can be derived from these three.

95
00:05:39,600 --> 00:05:41,980
So we could stop there and
if there's anything else you want to know

96
00:05:41,980 --> 00:05:46,620
about probability you could pull out your
pencil and paper and figure it out but

97
00:05:46,620 --> 00:05:52,120
you probably be advised to go to a book,
someones already done it.

98
00:05:52,120 --> 00:05:55,000
Here's a few of the highlights of
the things that you could prove from

99
00:05:55,000 --> 00:05:57,330
those three axioms, but

100
00:05:57,330 --> 00:06:01,600
the probability of [INAUDIBLE] is
simply one minus the probability of A.

101
00:06:01,600 --> 00:06:03,250
If A is a subset of B.

102
00:06:03,250 --> 00:06:08,040
Meaning in terms of those Venn diagrams,
if the event A was the red circle, and

103
00:06:08,040 --> 00:06:11,130
it lay entirely within the bounds
of the red circle that defined B.

104
00:06:11,130 --> 00:06:15,000
Then the probability of A would be less
than or equal to the probability of B.

105
00:06:16,000 --> 00:06:21,200
And, in general, if the probabil, the
probability of A union B, that's A or B.

106
00:06:21,200 --> 00:06:25,260
Is the probability of a plus the
probability of b minus the intersection,

107
00:06:25,260 --> 00:06:27,470
minus the probability of the intersection.

108
00:06:27,470 --> 00:06:32,160
And that is a little bit of
a generalization of the rule about

109
00:06:32,160 --> 00:06:33,050
disjoint events.

110
00:06:33,050 --> 00:06:35,640
Here what we're saying is A, if A and

111
00:06:35,640 --> 00:06:40,120
B are not disjoint,
I have to remember to remove one copy.

112
00:06:40,120 --> 00:06:44,680
Of the intersection, otherwise I would
have double-counted the intersection.

113
00:06:44,680 --> 00:06:47,930
In the fourth point there,
you see the generalization of the,

114
00:06:47,930 --> 00:06:54,010
of the A union B rule when we
have many events, not just two.

115
00:06:54,010 --> 00:06:57,180
And I will leave you to go and
look that up in a book.

116
00:06:57,180 --> 00:07:01,180
If you try to work out that little
example with A just A1 and A2,

117
00:07:01,180 --> 00:07:04,440
you should get the result
on the third line.

118
00:07:07,280 --> 00:07:11,290
And we will move on now to defining
joint and conditional probabilities.

119
00:07:11,290 --> 00:07:13,140
The joint probabilities of events A and

120
00:07:13,140 --> 00:07:16,380
B we already talked about that,
that's the intersection.

121
00:07:16,380 --> 00:07:20,870
That's elements of the sample space that
belong to both events at the same time.

122
00:07:20,870 --> 00:07:24,620
The conditional probability of
A given that B has occurred.

123
00:07:24,620 --> 00:07:28,500
Is simply the probability of
A intersection B divided by

124
00:07:28,500 --> 00:07:30,130
the probability of B.

125
00:07:30,130 --> 00:07:33,530
And the way to visualize that, I think,
with this diagram is to say well that's,

126
00:07:33,530 --> 00:07:36,820
you might think of that as
the area of the intersection, but

127
00:07:36,820 --> 00:07:41,260
not now normalized by the area of the
entire box labelled S, which would be one,

128
00:07:41,260 --> 00:07:44,020
but simply normalized by
the area of the blue circle.

129
00:07:45,190 --> 00:07:47,290
The unconditional probability of A,

130
00:07:47,290 --> 00:07:51,750
is the intersection of the A, with S,
divided by the probability of S.

131
00:07:52,920 --> 00:07:56,770
Since the probability of S is one,
we, the unconditional probability of

132
00:07:56,770 --> 00:08:00,730
A is just A intersectional
with a sample space, itself.

133
00:08:00,730 --> 00:08:01,660
And.

134
00:08:01,660 --> 00:08:05,490
The usual definition of statistical
independence is that the probability of

135
00:08:05,490 --> 00:08:10,280
A intersection B is equal to the product
of the probabilities of A and B.

136
00:08:10,280 --> 00:08:13,570
Now, I would take this moment to,
to make a comment.

137
00:08:13,570 --> 00:08:18,120
I often encounter with, particularly
with my physics friends up at the lab,

138
00:08:19,480 --> 00:08:21,119
different uses of the word independent.

139
00:08:22,130 --> 00:08:29,070
Sometimes the word independence is taken
in an English language sense/ and whenever

140
00:08:29,070 --> 00:08:32,160
we use the word independence in this set
of lectures, and in generally, when you're

141
00:08:32,160 --> 00:08:35,190
talking about probability, you mean this
very precise definition of independence.

142
00:08:36,990 --> 00:08:40,360
So that's something important to keep
in mind, if you are a Physics friend or

143
00:08:40,360 --> 00:08:43,110
you're talking to one or
to anyone else for that matter.

144
00:08:44,370 --> 00:08:46,730
finally, we can see that another
equivalent definition of

145
00:08:46,730 --> 00:08:51,210
statistical independence is that the
conditional probability of A given B is

146
00:08:51,210 --> 00:08:52,640
simply equal to the probability of A.

147
00:08:52,640 --> 00:08:55,610
And that has a sort of
pleasing interpretation,

148
00:08:55,610 --> 00:08:59,010
which is that you learn nothing
from knowing that B occurred.

149
00:09:00,190 --> 00:09:04,370
So if the probability if A given B is
exactly the same as the probability of

150
00:09:04,370 --> 00:09:05,490
A without.

151
00:09:05,490 --> 00:09:07,950
Anything to do with B then we would
say that A and B are independent.

152
00:09:10,780 --> 00:09:13,780
Okay, now we are getting
into the good stuff.

153
00:09:13,780 --> 00:09:19,770
There is something out there called the
law of total probability and all this says

154
00:09:19,770 --> 00:09:25,310
is that if I have the partitioning
of the Sample Space Say my,

155
00:09:25,310 --> 00:09:29,290
say I'm going to going to take away,
what I was calling event A before, and I'm

156
00:09:29,290 --> 00:09:33,710
going to simply divide the sample space
into three different events A1, A2 and A3.

157
00:09:34,930 --> 00:09:38,780
Then, I can reconstruct the probability
of B there, by simply taking

158
00:09:38,780 --> 00:09:42,360
the individual probabilities of
the intersection of each of those As.

159
00:09:42,360 --> 00:09:44,550
Oh, I'm sorry,
the conditional probability.

160
00:09:44,550 --> 00:09:46,510
Of, B given A for each of the As.

161
00:09:46,510 --> 00:09:49,603
And then weighting it by the probability
of each of those individual As.

162
00:09:49,603 --> 00:09:56,178
[SOUND] And this should make
some kind of intuitive sense.

163
00:09:56,178 --> 00:09:59,870
And in fact, that is nothing
more than the definition of

164
00:09:59,870 --> 00:10:03,870
conditional probability that
we saw in the earlier slide.

165
00:10:03,870 --> 00:10:06,174
now.
You've probably all heard of this,

166
00:10:06,174 --> 00:10:08,157
Bayes' Rule or Bayes' Theorem.

167
00:10:08,157 --> 00:10:11,790
Bayes' Theorem is nothing more
than another definition of

168
00:10:11,790 --> 00:10:13,610
conditional probability.

169
00:10:13,610 --> 00:10:18,010
It simply says that the probability
of A given B, can be expressed, well,

170
00:10:18,010 --> 00:10:24,390
we know it's definition is the probability
of A intersection B, which is on the.

171
00:10:24,390 --> 00:10:28,010
An equivalent expression is in
the numerator of the central term there.

172
00:10:28,010 --> 00:10:29,300
Divided by the probability of B.

173
00:10:30,800 --> 00:10:32,290
So, I can leave the numerator alone.

174
00:10:32,290 --> 00:10:35,070
And then I can notice
that I can use the law of

175
00:10:35,070 --> 00:10:38,570
total probability to re
express the probability of B.

176
00:10:38,570 --> 00:10:43,900
In the form on the, in the denominator
of the right side of the equation.

177
00:10:43,900 --> 00:10:47,520
And that's Bayes' rule and
what's good about Bayes' rule or

178
00:10:47,520 --> 00:10:52,200
what's useful about is it allows
to express the probability of

179
00:10:52,200 --> 00:10:56,110
a A given B in terms of
the probability of B given A and

180
00:10:56,110 --> 00:10:59,860
sometimes we know more about the
probability of B given A than we do about

181
00:10:59,860 --> 00:11:04,500
the probability of A given B and here's
probably the one example that I'll take.

182
00:11:05,750 --> 00:11:10,430
From things I do every day which
would be suppose A is the event that

183
00:11:10,430 --> 00:11:15,390
the true CO2 concentration in the
atmosphere over a column of atmosphere,

184
00:11:15,390 --> 00:11:18,160
is greater than 400 parts per million.

185
00:11:18,160 --> 00:11:21,120
And let B be the event that the
Orbiting Carbon Observatory instrument,

186
00:11:21,120 --> 00:11:23,560
which just launched in July.

187
00:11:23,560 --> 00:11:30,950
Observes CO2 concentration of 398 parts
per million, but what we would really

188
00:11:30,950 --> 00:11:35,650
like to know is what the conditional
distribution of A given B is in that case.

189
00:11:35,650 --> 00:11:40,260
What we would like to know, how likely
is it that the true concentration is 400

190
00:11:40,260 --> 00:11:44,980
parts per million if what CO2
observed was 398 parts per million.

191
00:11:44,980 --> 00:11:47,240
But we don't have a good
way of knowing that,

192
00:11:47,240 --> 00:11:52,080
but because we built OCO-2 ourselves and
we know how the instrument works and

193
00:11:52,080 --> 00:11:55,850
we know how the instrument behaves under
different physical conditions, we know

194
00:11:55,850 --> 00:12:01,430
much more about the probability that OCO-2
would observe 398 parts per million.

195
00:12:01,430 --> 00:12:06,520
If the true concentration in
the atmospheric column was 400 parts

196
00:12:06,520 --> 00:12:07,240
per million.

197
00:12:07,240 --> 00:12:13,560
So that's a case where we can use that,
and in fact is used to retrieve OCO2,

198
00:12:13,560 --> 00:12:17,900
CO2 concentration measurements that base
therom is how they do their so called

199
00:12:17,900 --> 00:12:22,188
retrieval, their measurement of CO2 in
the atmosphere, so that may be the little.

200
00:12:22,188 --> 00:12:25,340
Cocktail party tidbit.

201
00:12:25,340 --> 00:12:27,190
Okay, so
at the end of each of these sections.

202
00:12:27,190 --> 00:12:33,460
I'm going to put down some references
books typically that I'll,

203
00:12:33,460 --> 00:12:39,250
I'll be embarrassed to say in most cases
the books I learned these things from.

204
00:12:39,250 --> 00:12:44,010
Or, or, later editions of them, so that if
you wanted to go and look these things.

205
00:12:44,010 --> 00:12:47,180
There are many places you can go,
of course, to read about probability.

206
00:12:47,180 --> 00:12:49,760
A First Cour in probability
by Sheldon Ross is a good,

207
00:12:49,760 --> 00:12:53,790
compact, discussion of these items and
many, many more.

208
00:12:53,790 --> 00:12:54,940
And the classic.

209
00:12:56,280 --> 00:12:59,500
Is William Feller's Volume I and
II Introduction to Probability Theory and

210
00:12:59,500 --> 00:13:01,700
its Applications which
is quite old by now.

211
00:13:01,700 --> 00:13:05,570
But fortunately probability doesn't
change very much, over time.

212
00:13:05,570 --> 00:13:09,550
The Volume I is devoted exclusively
to discreet probabilities,

213
00:13:09,550 --> 00:13:11,380
which we'll get to in a minute.

214
00:13:11,380 --> 00:13:16,760
Volume two is, devoted to, continuous
probabilities, and they're both excellent

215
00:13:16,760 --> 00:13:22,030
and they have great problems in the back
if you like doing that sort of thing.

216
00:13:22,030 --> 00:13:25,710
So in the module now, we'll discuss how
these rules of probability will translate

217
00:13:25,710 --> 00:13:28,100
for settings in which we
model numerical phenomenon.

218
00:13:28,100 --> 00:13:28,970
In other words data.

