1
00:00:06,263 --> 00:00:07,991
[MUSIC]. 
Okay, so, so far we've been discussing 

2
00:00:07,991 --> 00:00:11,663
statistical inference from a particular 
perspective, which is the frequentest 

3
00:00:11,663 --> 00:00:17,396
perspective. 
And so frequentists are concerned with 

4
00:00:17,396 --> 00:00:22,620
the Probability of seeing a particular 
data sample, given the null hypothesis. 

5
00:00:22,620 --> 00:00:24,400
Okay, and that's what the p-value gives 
you. 

6
00:00:24,400 --> 00:00:28,310
But there's another perspective to 
statistics called the Bayesian approach, 

7
00:00:28,310 --> 00:00:34,600
which is concerned with the probability 
of Having a certain outcome given the 

8
00:00:34,600 --> 00:00:37,620
data we've already seen. 
Okay, and so I want to talk a little bit 

9
00:00:37,620 --> 00:00:41,715
about Bayesian approaches and give an 
example of their power. 

10
00:00:41,715 --> 00:00:45,470
Okay. 
What are the differences here? 

11
00:00:45,470 --> 00:00:48,520
Well, you can think about the differences 
In terms of what is fixed. 

12
00:00:48,520 --> 00:00:52,720
So the frequency's perspective is that 
data are a repeatable random sample, 

13
00:00:52,720 --> 00:00:54,195
right? 
You're always allowed to go do another 

14
00:00:54,195 --> 00:00:59,050
experiment the same way, and in fact you 
need to at least virtually, to reason 

15
00:00:59,050 --> 00:01:03,390
about, the probabilities that get 
produced by these methods. 

16
00:01:03,390 --> 00:01:05,660
Okay. 
So there is some sort of frequency you 

17
00:01:05,660 --> 00:01:10,171
can reason about, frequency of of 
achieving a certain outcome. 

18
00:01:10,171 --> 00:01:13,570
Okay. 
And the underlying parameters of the 

19
00:01:13,570 --> 00:01:16,660
population remain constant during this 
repeatable process. 

20
00:01:16,660 --> 00:01:21,406
Right, your, your, the population stays 
fixed and you run experiments to 

21
00:01:21,406 --> 00:01:25,430
determine, you know likelihoods. 
Okay, meanwhile, with the Bayesian 

22
00:01:25,430 --> 00:01:31,430
approach, the data observed from a 
realized sample, and the parameters A, of 

23
00:01:31,430 --> 00:01:35,220
the population or unknown, but can be 
described probabilistically, right. 

24
00:01:35,220 --> 00:01:40,340
So there not sort of, fixed values. 
However, the data are fixed, right, you 

25
00:01:40,340 --> 00:01:44,090
don't think about going back and sampling 
more data, you just have the observations 

26
00:01:44,090 --> 00:01:52,110
you have, okay. 
And so, the Bayesian approach is 100% 

27
00:01:52,110 --> 00:01:55,700
concerned with the application of 
Bayesian rule which is this. 

28
00:01:55,700 --> 00:02:00,955
Okay, so this, what Thomas Bayesian did 
was relate these conditional 

29
00:02:00,955 --> 00:02:08,880
probabilities with the prior beliefs. 
Okay, so it allows you to take your- 

30
00:02:08,880 --> 00:02:13,420
Belief about the probabilities of certain 
events happening and update them when 

31
00:02:13,420 --> 00:02:19,950
more data is collected, okay. 
And so here's what it says, it says the 

32
00:02:19,950 --> 00:02:24,630
probability of event A happening, given 
that event B has already happened. 

33
00:02:24,630 --> 00:02:28,610
Is equal to the probability of b 
happening given that a has already 

34
00:02:28,610 --> 00:02:33,510
happened, multiplied by the probability 
of a happening across the board, divided 

35
00:02:33,510 --> 00:02:35,470
by the probability of be happening across 
the board. 

36
00:02:35,470 --> 00:02:38,660
And if that's not clear, that's okay. 
We're, we're going to go into more detail 

37
00:02:38,660 --> 00:02:41,710
here. 
So right now though, recognize that the 

38
00:02:41,710 --> 00:02:46,000
key benefit here is the ability to 
incorporate prior knowledge, which is not 

39
00:02:46,000 --> 00:02:52,720
the case with the frequency approach. 
Now, a key weakness here is that you need 

40
00:02:52,720 --> 00:02:55,770
to incorporate prior knowledge So if you, 
so there's a couple of problems with 

41
00:02:55,770 --> 00:02:58,620
this. 
One is if you don't know anything about 

42
00:02:58,620 --> 00:03:05,300
the population you're modeling then there 
is not much you can do with the prior, 

43
00:03:05,300 --> 00:03:08,510
there's not a good way to model this, the 
distributions. 

44
00:03:08,510 --> 00:03:15,545
Now, there are some techniques to sort of 
derive so called uninformative priors 

45
00:03:15,545 --> 00:03:20,030
that try not to influence things to much 
but- Give you a plugin to be able to 

46
00:03:20,030 --> 00:03:21,690
rule. 
But still, that is an issue. 

47
00:03:21,690 --> 00:03:26,190
fine. 
Perhaps more insideously and the reason 

48
00:03:26,190 --> 00:03:34,440
why Bayes' the Bayesian approach, was 
popular and then it fell out of favor in 

49
00:03:34,440 --> 00:03:36,140
the earlier part of the twentieth 
century. 

50
00:03:36,140 --> 00:03:43,428
Was that, you can use this rule to kind 
of, do anything you want, to confirm or 

51
00:03:43,428 --> 00:03:46,880
deny the affect of any sort of evidence, 
right? 

52
00:03:46,880 --> 00:03:48,495
Just by plugging in your own prior 
belief. 

53
00:03:48,495 --> 00:03:52,325
Okay. 
So, given a fixed set of data, two 

54
00:03:52,325 --> 00:03:56,060
different people can come up with two 
different conclusion about the data 

55
00:03:56,060 --> 00:03:59,440
because they had prior beliefs. 
They modeled the, the, the prior 

56
00:03:59,440 --> 00:04:03,140
distributions differently, okay. 
And this was seen as a major flaw, and 

57
00:04:03,140 --> 00:04:06,404
gave rise to the frequency approach, 
which. 

58
00:04:06,404 --> 00:04:11,990
Looks more or less objectively at the 
data itself. 

59
00:04:11,990 --> 00:04:12,230
Okay. 
Right. 

60
00:04:12,230 --> 00:04:21,150
And so that's what is spelled out here in 
a nice essay by Matthews 1998, is that 

61
00:04:21,150 --> 00:04:23,320
different people could use Bayesian 
Theorem and get different results. 

62
00:04:23,320 --> 00:04:28,819
And so, faced with some, some 
experimental evidence per se ESP true 

63
00:04:28,819 --> 00:04:32,990
believers could uses Bayes's Theorem to 
sh show that the new results confirmed 

64
00:04:32,990 --> 00:04:39,784
it, while skeptics could use it to show 
that it the ESP didn't exist. 

65
00:04:39,784 --> 00:04:43,380
Okay, and both views are possible because 
Bayesian Theorem only shows how to alter 

66
00:04:43,380 --> 00:04:46,660
one's prior level of belief. 
And different people can start out with 

67
00:04:46,660 --> 00:04:48,790
different opinions. 
Okay, so that's the issue. 

68
00:04:48,790 --> 00:04:57,940
However the frequency oriented, the 
frequentist approach So, the frequency 

69
00:04:57,940 --> 00:05:05,050
approach, laid out by Fisher and others, 
were able to achieve what was thought to 

70
00:05:05,050 --> 00:05:07,770
be impossible. 
Which is a way of judging the 

71
00:05:07,770 --> 00:05:12,545
significance of experimental data 
independent of any prior beliefs. 

72
00:05:12,545 --> 00:05:16,260
Okay. 
So he'd found a way that anyone could use 

73
00:05:16,260 --> 00:05:20,060
to show that a result was too impressive, 
too statistically significant to be 

74
00:05:20,060 --> 00:05:23,239
dismissed as a fluke. 
And, you know, all you had to do was 

75
00:05:23,239 --> 00:05:26,873
convert your raw data into this thing 
called a P-value. 

76
00:05:26,873 --> 00:05:29,500
Alright. 
Now, you have to have some sort of 

77
00:05:29,500 --> 00:05:33,605
threshold to measure the significance of 
this P-value, and that was written up as 

78
00:05:33,605 --> 00:05:37,700
0.05 in by Fisher. 
You know, in a, in a, in a original 

79
00:05:37,700 --> 00:05:40,890
paper. 
So what were the insights that led to 

80
00:05:40,890 --> 00:05:45,040
this particular value of 0.05? 
Well, Fisher admitted that there weren't 

81
00:05:45,040 --> 00:05:47,210
any. 
He simply decided 0.05 because it was 

82
00:05:47,210 --> 00:05:50,490
mathematically convenient. 
Okay? 

83
00:05:50,490 --> 00:05:55,170
And you know, in the many people have 
pointed out, say from the 60s on, there's 

84
00:05:55,170 --> 00:05:59,520
been periods of time where there's been a 
significant amount of work showing that 

85
00:05:59,520 --> 00:06:05,745
this approach, with these particular 
thresholds are able to produce a lot of 

86
00:06:05,745 --> 00:06:12,000
incorrect conclusions. 
So James Berger Purdue wrote his entire 

87
00:06:12,000 --> 00:06:15,539
series of papers warning about the quote 
astonishing tendency of Fisher's p-values 

88
00:06:15,539 --> 00:06:20,340
to exaggerate significance. 
Findings that met the 120 standard can 

89
00:06:20,340 --> 00:06:23,659
actually arise when the data provide 
little, very little of no evidence in 

90
00:06:23,659 --> 00:06:26,799
favor of an effect. 
Okay, so that's the problem, so now 

91
00:06:27,980 --> 00:06:31,760
perhaps the, we can say that the pendulum 
is in some sense swinging back towards 

92
00:06:31,760 --> 00:06:35,930
the Bayesian approach. 
One more problem with the Bayesian 

93
00:06:35,930 --> 00:06:39,910
approach that I'll bring up that I don't 
necessarily have a slide on, is that, 

94
00:06:39,910 --> 00:06:43,772
which we'll see in a little bit of detail 
in. 

95
00:06:43,772 --> 00:06:48,930
Perhaps the next segment, is that these 
conditional probabilities and these prior 

96
00:06:48,930 --> 00:06:53,995
probabilities end up forcing you to, to 
model a situation mathematically, and 

97
00:06:53,995 --> 00:06:59,030
that situation may be very difficult to 
model mathematically. 

98
00:06:59,030 --> 00:07:03,130
Okay. 
So what you end up with is these, these 

99
00:07:03,130 --> 00:07:06,580
chains of complicated conditional 
probabilities that need to be integrated 

100
00:07:06,580 --> 00:07:13,600
in order to apply Bayes' rule, and so 
this was computationally intractable, and 

101
00:07:13,600 --> 00:07:18,920
there were a norm, a huge number of 
tricks and simplications and, and ideas 

102
00:07:18,920 --> 00:07:22,290
used to make this more tractable. 
Okay. 

103
00:07:23,460 --> 00:07:27,520
But, thanks to the development of methods 
that allow you sample these complicated 

104
00:07:27,520 --> 00:07:33,940
distributions computationally these, this 
approach is becoming increasingly 

105
00:07:33,940 --> 00:07:35,910
popular. 
Okay. 

106
00:07:35,910 --> 00:07:38,890
And there's a lot of thinkers in this 
space that believe that, you know, for 

107
00:07:38,890 --> 00:07:44,190
the 21st century, a combination of 
frequent and Bayesian approaches is going 

108
00:07:44,190 --> 00:07:49,270
to be dominant. 
Okay, so for these reasons we're going to 

109
00:07:49,270 --> 00:07:54,030
make sure we talk about the Bayesian 
approach at a very introductory level and 

110
00:07:54,030 --> 00:07:58,930
leading up to machine learning algorithm 
called naive bayes. 

