1
00:00:05,616 --> 00:00:06,921
[MUSIC]. 
As we talked about Bayesian versus 

2
00:00:06,921 --> 00:00:10,071
Frequentist approaches to two statistics, 
and this cartoon is my attempt to a quick 

3
00:00:10,071 --> 00:00:15,211
summary. 
So the Frequentist approach is concerned 

4
00:00:15,211 --> 00:00:18,750
with using the data and only the data to 
make a decision. 

5
00:00:18,750 --> 00:00:24,840
While the Bayesian approach incorporates 
prior belief as well as the data, okay. 

6
00:00:24,840 --> 00:00:27,380
And that's both a strength and a 
weakness. 

7
00:00:27,380 --> 00:00:31,680
Strength in that, if you have good 
information, you can incorporate it. 

8
00:00:31,680 --> 00:00:35,290
If the weakness in that, you have to 
actually incorporate information and that 

9
00:00:35,290 --> 00:00:40,357
it's somewhat subjective, depending on 
who you're talking to, fine. 

10
00:00:40,357 --> 00:00:47,320
So Bayes' Theorem, breaking it down in a 
little more detail, you've got this 

11
00:00:47,320 --> 00:00:51,720
terminology you can apply to it. 
So the prior, as we said, here in blue, 

12
00:00:51,720 --> 00:00:55,390
is the probability of the hypothesis 
being true before you've collected any 

13
00:00:55,390 --> 00:00:58,220
data, right? 
So what's the, of the, in the global 

14
00:00:58,220 --> 00:01:02,645
population, how many people are above a 
certain height? 

15
00:01:02,645 --> 00:01:06,840
Okay, then the marginal probability is 
what is the probability of collecting 

16
00:01:06,840 --> 00:01:10,355
this data, this particular data 
observations under all possible 

17
00:01:10,355 --> 00:01:16,200
hypotheses, okay. 
And then the likelihood is the 

18
00:01:16,200 --> 00:01:21,154
probability of collecting this data given 
that our hypothesis is actually true. 

19
00:01:21,154 --> 00:01:23,960
And then finally the posterior is what 
we're usually trying to compute, which is 

20
00:01:23,960 --> 00:01:30,720
the probability of our hypothesis being 
true, given the data collected, okay. 

21
00:01:30,720 --> 00:01:33,890
So, let's think about it this way, 
arrange this into a grid. 

22
00:01:33,890 --> 00:01:38,180
This is like, you know, basically, it's 
actually pretty trivial to derive this 

23
00:01:38,180 --> 00:01:41,200
theorem, if you think carefully about 
what's going on. 

24
00:01:42,410 --> 00:01:48,190
So, a nice way to explain it is so build 
a two by two grid like this. 

25
00:01:48,190 --> 00:01:56,255
Where you have A, and B, and not A and 
not B. 

26
00:01:56,255 --> 00:02:05,190
And, the box here is the probability of A 
and B happening, okay. 

27
00:02:05,190 --> 00:02:11,660
Well, A and B is the probability of A 
given that B has already occurred, 

28
00:02:11,660 --> 00:02:16,550
multiplied by the probability that B 
occurs, right? 

29
00:02:16,550 --> 00:02:18,375
And you should take a moment and convince 
yourself of that. 

30
00:02:18,375 --> 00:02:26,130
Equivalently, is the probability that B 
of B occurring, given that A has already 

31
00:02:26,130 --> 00:02:32,180
occurred, multiplied by the probability 
that A actually occurs, okay. 

32
00:02:32,180 --> 00:02:39,040
So these things are all equivalent. 
Set these two things equal to each other, 

33
00:02:39,040 --> 00:02:44,720
and just, you know, divide. 
So set these two things equal to one 

34
00:02:44,720 --> 00:02:49,758
another, and divide the both sides by the 
probability of B occurring and you have, 

35
00:02:49,758 --> 00:02:52,156
Bayses' rule. 
All right so let's try this out. 

36
00:02:52,156 --> 00:02:58,055
So, there's a question. 
Say that you know that 1% of women at age 

37
00:02:58,055 --> 00:03:03,660
40 who participate in routine screening, 
have Breast Cancer. 

38
00:03:03,660 --> 00:03:07,980
And let's say you know 80% of women who 
actually do have Breast Cancer will get a 

39
00:03:07,980 --> 00:03:15,570
positive, result from the test. 
Further, say you know that 9.6% of women 

40
00:03:15,570 --> 00:03:19,890
who do not have Breast Cancer will also 
get positive results. 

41
00:03:19,890 --> 00:03:23,480
So these, this is the false positive 
rate, okay. 

42
00:03:23,480 --> 00:03:27,502
Now given that you know a woman in this 
age group had a positive test. 

43
00:03:27,502 --> 00:03:30,400
All right, the test came back positive in 
a routine screening. 

44
00:03:30,400 --> 00:03:34,040
What is the probability that she actually 
has breast cancer? 

45
00:03:35,260 --> 00:03:36,820
And you should take a minute to sort of 
work this out. 

46
00:03:39,910 --> 00:03:45,110
So what's sort of remarkable is that, 
this is a fairly straightforward 

47
00:03:45,110 --> 00:03:50,960
application of Bayes' rule. 
But intuitively it's easy to make sort of 

48
00:03:52,270 --> 00:03:56,660
mistakes in, in the reasoning. 
And in fact when you ask, there's been 

49
00:03:56,660 --> 00:04:00,130
some famous studies that are now somewhat 
out of date where they asked doctors this 

50
00:04:00,130 --> 00:04:02,380
question. 
And the doctors came up with wildly wrong 

51
00:04:02,380 --> 00:04:05,560
answers. 
Only 15% of the, doctors they asked were 

52
00:04:05,560 --> 00:04:07,889
able to answer the question correctly. 
Okay, so, here's how to break it down, 

53
00:04:07,889 --> 00:04:10,280
again with our two by two grid. 
We've got 1% of cases that have cancer, 

54
00:04:10,280 --> 00:04:25,030
and we've got 99, which means we have 99% 
of cases that do not have cancer. 

55
00:04:25,030 --> 00:04:27,894
This is in the global population, right? 
And let's say we just know this. 

56
00:04:27,894 --> 00:04:31,980
All right, we know that 1% of people, 
across the board, have cancer, or at 

57
00:04:31,980 --> 00:04:36,740
least in this group that we're studying. 
I shouldn't say across the board, okay. 

58
00:04:36,740 --> 00:04:42,920
Now, we also know that if you have 
cancer, there's an 80% chance that when 

59
00:04:42,920 --> 00:04:46,170
you take the test it'll come back 
positive, okay. 

60
00:04:46,170 --> 00:04:52,088
So the true positive rate is 1% have 
cancer times 80%. 

61
00:04:52,088 --> 00:04:58,850
Now, we're also given that if you do not 
have cancer and you take the test, you'll 

62
00:04:58,850 --> 00:05:01,700
get a positive result of 9, at a rate of 
9.6%. 

63
00:05:01,700 --> 00:05:07,168
And so the false positive rate is the 
rate of not having cancer, 99% times 

64
00:05:07,168 --> 00:05:08,726
9.6%. 
Now, suddenly for the false negative and 

65
00:05:08,726 --> 00:05:11,386
the true negative tests, 1% times 20 
times 1 minus 80 and 99% times 1 minus 

66
00:05:11,386 --> 00:05:24,360
9.6 is 90.4. 
Okay, so we have all the information we 

67
00:05:24,360 --> 00:05:26,840
need. 
So going back to actually answering the 

68
00:05:26,840 --> 00:05:31,460
questions we can, write down, Bayes' 
Rule. 

69
00:05:31,460 --> 00:05:35,830
So, what is the probability of having 
cancer given that we have a positive test 

70
00:05:35,830 --> 00:05:39,120
result? 
Well, that's the probability of getting a 

71
00:05:39,120 --> 00:05:41,490
positive test result, given we have 
cancer. 

72
00:05:43,340 --> 00:05:47,010
Multiplied by the probability of having 
cancer overall, divided by the 

73
00:05:47,010 --> 00:05:56,060
probability of a positive test overall, 
okay. 

74
00:05:56,060 --> 00:05:59,550
So, the only one of these terms, so we 
actually know this terms straight out, 

75
00:05:59,550 --> 00:06:06,670
we're given this one, all right, which is 
80%. 

76
00:06:06,670 --> 00:06:10,330
And we're given this one, the probability 
of cancer overall is, is 1%. 

77
00:06:10,330 --> 00:06:14,110
We don't have this denominator given, so 
we need to figure that out. 

78
00:06:14,110 --> 00:06:20,450
Well, that's the chance of a positive 
test in all other occurrences, the chance 

79
00:06:20,450 --> 00:06:23,020
of a positive test. 
Given that we have cancer and the chance 

80
00:06:23,020 --> 00:06:24,939
of a positive test given that they don't 
have cancer. 

81
00:06:27,430 --> 00:06:29,300
In each case multiplied by the 
probability of that happening. 

82
00:06:29,300 --> 00:06:36,500
Okay, so you can, you can decompose this. 
So that's 0.8 times the 1% probability of 

83
00:06:36,500 --> 00:06:45,770
actually having cancer plus 9.6% times 
99% which gives you this number of 10.3%. 

84
00:06:45,770 --> 00:06:52,148
So that's the overall probability of 
getting a positive test result, all 

85
00:06:52,148 --> 00:06:55,250
right? 
And so now, you plug this in, you end up 

86
00:06:55,250 --> 00:06:59,210
with a number of 7.8% for our, of our 
answer. 

87
00:06:59,210 --> 00:07:03,158
If we have a positive test result, the 
chance of actually having cancer is 7.8%. 

88
00:07:03,158 --> 00:07:07,110
So this is lower than you might come up 
with if you don't think about it 

89
00:07:07,110 --> 00:07:08,910
carefully, right? 
You might think that boy, I got a 

90
00:07:08,910 --> 00:07:11,540
positive test result. 
There's something, there's this 80% 

91
00:07:11,540 --> 00:07:14,170
floating around. 
You know, boy, there's probably a 70 to 

92
00:07:14,170 --> 00:07:21,450
80% change that that, I have cancer. 
But because of the very low percentage of 

93
00:07:21,450 --> 00:07:24,380
having caner in the, in the prior 
probabilities. 

94
00:07:24,380 --> 00:07:26,830
The actual number is still pretty low, 
okay? 

95
00:07:26,830 --> 00:07:28,380
So it's easy to make mistakes with this 
stuff. 

96
00:07:28,380 --> 00:07:31,610
Now, this was a remarkably simple case. 
First of all, we're given all this 

97
00:07:31,610 --> 00:07:33,670
information. 
Second of all, there's only two 

98
00:07:33,670 --> 00:07:36,064
possibilities, these sort of binary 
variables. 

99
00:07:36,064 --> 00:07:38,270
So let's think about something a little 
more complicated. 

100
00:07:39,300 --> 00:07:43,780
Okay, so let's consider a classic 
application of Bayes' rule to a big data 

101
00:07:43,780 --> 00:07:48,820
problem, which is spam filtering. 
Okay, so here our task is to determine 

102
00:07:48,820 --> 00:07:53,130
whether an email message is spam. 
The probability that an email message is 

103
00:07:53,130 --> 00:07:59,416
spam, given the words in the email 
message. 

104
00:07:59,416 --> 00:08:04,720
Okay, and with Bayes' rule, you can 
express that probability as the 

105
00:08:04,720 --> 00:08:07,859
probability that the email message is 
spam overall. 

106
00:08:09,700 --> 00:08:13,750
Multiplied by the probability of seeing 
these particular words in the message, 

107
00:08:13,750 --> 00:08:18,860
given that we already know it's spam. 
And all that divided by the probability 

108
00:08:18,860 --> 00:08:24,440
of seeing these words in the message. 
Now, the interesting one here is this 

109
00:08:24,440 --> 00:08:29,020
numerator. 
And the reason is that, the probability 

110
00:08:29,020 --> 00:08:35,510
of words appearing doesn't involve the 
unknown label of whether it's spam or 

111
00:08:35,510 --> 00:08:38,520
not. 
And so all we're trying to do is get a 

112
00:08:38,520 --> 00:08:42,460
relative frequency of spam or not spam, 
okay? 

113
00:08:42,460 --> 00:08:45,340
And so, just dividing by a constant 
factor of the probability of seeing these 

114
00:08:45,340 --> 00:08:48,730
words doesn't change our decision at all. 
So, we don't care about the actual 

115
00:08:48,730 --> 00:08:51,834
number, we just care about the decision 
of spam or not spam. 

116
00:08:51,834 --> 00:08:57,245
Okay, so fine, so, re-expressing is 
before we get rid of the denominator. 

117
00:08:57,245 --> 00:08:59,675
Re-expressing this, what do we mean by 
words? 

118
00:08:59,675 --> 00:09:04,770
Well, this will, you can write this as 
probability that the email message is 

119
00:09:04,770 --> 00:09:08,580
spam, given that the word viagra appears 
in the message. 

120
00:09:08,580 --> 00:09:10,950
Given that the word rich, appears in the 
message. 

121
00:09:10,950 --> 00:09:13,570
Given that the word, something more 
innocuous, perhaps like friend appears in 

122
00:09:13,570 --> 00:09:16,210
the message. 
So all the words, you know, in the 

123
00:09:16,210 --> 00:09:20,272
English language or, or, all the words 
of, of interest to us in this test, okay. 

124
00:09:20,272 --> 00:09:25,938
So that's, re-expressed that way. 
Now, this numerator can be rewritten in 

125
00:09:25,938 --> 00:09:30,989
the following way. 
Given that it's a conditional 

126
00:09:30,989 --> 00:09:35,920
probability, we can apply, a chain rule, 
repeated, a repeated application of the 

127
00:09:35,920 --> 00:09:39,870
definition of conditional probability to 
obtain this. 

128
00:09:39,870 --> 00:09:44,300
The probability that it's spam multiplied 
by, so let's see, so this expression 

129
00:09:44,300 --> 00:09:47,760
rather, can be expressed as the 
probability seeing the word viagra. 

130
00:09:48,890 --> 00:09:52,610
Given that it's spam, multiplied by the 
probability of seeing all these other 

131
00:09:52,610 --> 00:09:55,550
words. 
given that it's spam and given that it's 

132
00:09:55,550 --> 00:09:58,280
viagra. 
Or given that the email message contains 

133
00:09:58,280 --> 00:10:02,230
viagra. 
And you can keep going. 

134
00:10:02,230 --> 00:10:05,490
This probability times the probability of 
seeing the word rich, given that it's 

135
00:10:05,490 --> 00:10:08,690
spam, and given that is viagra. 
Multiplied by the probability of all the 

136
00:10:08,690 --> 00:10:13,150
other words, given that it is spam, given 
that it contains the word viagra. 

137
00:10:13,150 --> 00:10:16,770
Given that it contains the word rich, and 
so on, okay. 

138
00:10:16,770 --> 00:10:22,300
So this is a long, complicated, 
conditional probability. 

139
00:10:22,300 --> 00:10:26,270
And this is where the Naive Bayes 
assumption comes in. 

140
00:10:26,270 --> 00:10:31,430
So, under the Naive Bayes assumption we 
say that the probabilities of these 

141
00:10:31,430 --> 00:10:34,445
different words appearing in this email 
message are completely independent. 

142
00:10:34,445 --> 00:10:39,700
That it's no more likely for you to see 
the word rich, when you see the word 

143
00:10:39,700 --> 00:10:45,015
wealth than it is, you know, without the 
word wealth there. 

144
00:10:45,015 --> 00:10:49,370
Okay, this isn't true, right? 
Obviously words go, go together. 

145
00:10:49,370 --> 00:10:54,320
There are co-occurrence rates, right? 
But you just ignore that, and just treat 

146
00:10:54,320 --> 00:10:57,760
everything as completely independent, 
which allows you to simplify this 

147
00:10:57,760 --> 00:11:00,590
expression as just a sequence of 
probabilities. 

148
00:11:00,590 --> 00:11:04,050
What is the probability of seeing the 
word viagra, given that it's spam? 

149
00:11:04,050 --> 00:11:07,290
What is the probability of seeing the 
word rich, given that it's spam, and so 

150
00:11:07,290 --> 00:11:09,580
on? 
Now, how do you get these probabilities? 

151
00:11:12,070 --> 00:11:18,528
Well, you have data, right? 
You had a set of documents that had been 

152
00:11:18,528 --> 00:11:24,600
pre-labeled as spam and you can look at 
the number of them that contain the word 

153
00:11:24,600 --> 00:11:29,830
rich. 
And divide by the total number, okay. 

154
00:11:29,830 --> 00:11:37,650
And so now you can calculate the 
probability of, of the two classes, spam 

155
00:11:37,650 --> 00:11:41,770
and not spam. 
And apply a decision procedure to call 

156
00:11:41,770 --> 00:11:44,860
it. 
And in fact, a simple one is just 

157
00:11:44,860 --> 00:11:47,490
whichever one is more likely, whichever 
one has the higher probability. 

158
00:11:47,490 --> 00:11:53,264
It's called the MAP decision rule which 
stands for the Maximum A Posteriori. 

159
00:11:53,264 --> 00:11:53,936
Okay. 

