1
00:00:00,001 --> 00:00:02,195
Okay so where are we? 
We're talking about supervised learning 

2
00:00:02,195 --> 00:00:05,103
and we're talking specifically about 
classification problems where we predict 

3
00:00:05,103 --> 00:00:07,330
a class label based on the other 
attributes. 

4
00:00:07,330 --> 00:00:11,570
And so we gave an example of predicting 
whether or not we played some sport based 

5
00:00:11,570 --> 00:00:23,630
on a few different attributes 
representing the weather. 

6
00:00:23,630 --> 00:00:28,020
Or, predicting whether a passenger aboard 
the Titanic did or did not survive, based 

7
00:00:28,020 --> 00:00:33,065
on things, things like their gender, 
their age, the fare they paid, and so on. 

8
00:00:33,065 --> 00:00:35,590
Okay. 
And, so, to get started here, we talked 

9
00:00:35,590 --> 00:00:40,480
about just, you know, manually guessing 
simple rules that might explain the data, 

10
00:00:40,480 --> 00:00:44,310
and, as you might imagine. 
The relationships between the attributes 

11
00:00:44,310 --> 00:00:50,090
and the class label are complex, and so, 
you know, your, your intuition may be 

12
00:00:50,090 --> 00:00:54,430
wrong and so you need some way to 
automate the search for these kinds of 

13
00:00:54,430 --> 00:00:55,180
rules. 
Okay. 

14
00:00:55,180 --> 00:01:01,480
So we've talked about this one-rule 
algorithm that will proceduralizes the 

15
00:01:01,480 --> 00:01:05,870
process of choosing a good rule. 
and that worked fine. 

16
00:01:05,870 --> 00:01:08,840
But, you know, obviously it's pretty 
limited, right, it can only, it can only 

17
00:01:08,840 --> 00:01:13,140
find very simple kinds of relationships, 
and so we also talked about sequential 

18
00:01:13,140 --> 00:01:17,230
cover algorithm that looks for, and 
builds up sets of, rules that may have 

19
00:01:17,230 --> 00:01:20,870
multiple conditions in them. 
Okay, you know, so great, so a set of 

20
00:01:20,870 --> 00:01:26,040
rules they're themselves complex, are 
more descriptive of the data and 

21
00:01:26,040 --> 00:01:28,490
therefore have better strength as a 
classifier. 

22
00:01:28,490 --> 00:01:32,650
But they're also harder to interpret. 
And one of the points that we'll be 

23
00:01:32,650 --> 00:01:35,855
making in this class is that, you know, 
communicability of these models, 

24
00:01:35,855 --> 00:01:40,860
communicability of the results is 
important, okay. 

25
00:01:40,860 --> 00:01:44,850
And so, a single rule is easy, a set of 
complex rules is perhaps less so. 

26
00:01:46,150 --> 00:01:50,380
And for that reason and for other reasons 
we can talk about decision trees. 

27
00:01:50,380 --> 00:01:56,960
And we sort of show that a decision tree, 
I was about to say can be constructed 

28
00:01:56,960 --> 00:02:00,420
from a set of rules. 
Each path from the root is a rule, but 

29
00:02:00,420 --> 00:02:03,120
that's not quite true. 
Constructing a decision tree from a set 

30
00:02:03,120 --> 00:02:06,320
of rules is not necessarily trivial. 
Going the other direction is pretty, is 

31
00:02:06,320 --> 00:02:10,480
straightforward as we as we described. 
Every path in the room is a rule, okay. 

32
00:02:10,480 --> 00:02:13,430
But, but there is a relationship, 
moreover, given a decision tree it's 

33
00:02:13,430 --> 00:02:15,970
pretty easy to interpret I would claim 
right. 

34
00:02:15,970 --> 00:02:21,080
You just sort of take your data item and 
look at the root node and answer the 

35
00:02:21,080 --> 00:02:23,920
question. 
You know, is, is the ginger male or 

36
00:02:23,920 --> 00:02:25,900
female? 
Female, well then go down this branch, 

37
00:02:25,900 --> 00:02:29,410
male, go down this branch so it's very 
easy to understand what's going on. 

38
00:02:29,410 --> 00:02:32,820
it's also easy to understand what the 
most important decisions are because they 

39
00:02:32,820 --> 00:02:36,620
bubble up to the top, Okay? 
And so if someone's asking you well, you 

40
00:02:36,620 --> 00:02:39,725
know, how is your model behaving, why 
does it make one decision over the 

41
00:02:39,725 --> 00:02:41,585
another? 
You can sort of answer those questions, 

42
00:02:41,585 --> 00:02:44,820
Okay? 
And so we talked about using how you make 

43
00:02:44,820 --> 00:02:51,270
the decision of which attributes to 
select at each level and we talked about 

44
00:02:51,270 --> 00:02:54,440
entropy. 
Being a measure of, sort of, the purity, 

45
00:02:54,440 --> 00:02:56,880
right? 
The gain is to try to make an entire 

46
00:02:56,880 --> 00:03:00,276
branch, pure with respect to a class 
label, right? 

47
00:03:00,276 --> 00:03:03,132
Everyone down this branch survived. 
Okay. 

48
00:03:03,132 --> 00:03:08,340
And so entropy gives us a way to do that. 
And we talk about a few extensions for 

49
00:03:08,340 --> 00:03:11,330
numeric attributes, where you could 
split. 

50
00:03:11,330 --> 00:03:16,185
based on, at a particular level as 
opposed to, you know, so if you don't 

51
00:03:16,185 --> 00:03:23,330
have gender male and female but rather 
you have a, a, a number like humidity or 

52
00:03:23,330 --> 00:03:26,580
the [INAUDIBLE] you have to be a little 
more careful. 

53
00:03:26,580 --> 00:03:29,510
You don't want to have a thousand 
different branches coming out of a node, 

54
00:03:29,510 --> 00:03:32,770
one for each unique value. 
Rather you want to bucket them somehow, 

55
00:03:32,770 --> 00:03:35,920
and so what's the way to bucket them? 
Well, one, one idea is just to find a 

56
00:03:35,920 --> 00:03:38,410
split point that splits them to two 
children. 

57
00:03:38,410 --> 00:03:41,130
So everything below this value goes one 
way, and everything above this value goes 

58
00:03:41,130 --> 00:03:43,180
another way, and we talked about how to 
find that threshold. 

59
00:03:43,180 --> 00:03:44,690
Okay. 
So, fine. 

60
00:03:44,690 --> 00:03:48,528
So where we are now is that decision 
trees are potentially prone to over 

61
00:03:48,528 --> 00:03:52,030
fitting and so. 
Let's talk about that. 

62
00:03:53,180 --> 00:03:58,390
This except is from Pedro Domingo's 2012 
CACM paper that we've mentioned before. 

63
00:03:58,390 --> 00:04:01,600
So you know, what if the knowledge and 
data we have are not sufficient to 

64
00:04:01,600 --> 00:04:03,870
completely determine the correct 
classifier. 

65
00:04:03,870 --> 00:04:08,320
So it could be that we've. 
You know, designed a classifier that is 

66
00:04:08,320 --> 00:04:11,260
responding to what, you know, what he 
calls random quirks in the data as 

67
00:04:11,260 --> 00:04:14,960
opposed to sort of uncovering some kind 
of fundamental truth. 

68
00:04:14,960 --> 00:04:17,960
You know, it doesn't have any predictive 
power in the real world, it just sort of 

69
00:04:17,960 --> 00:04:21,880
describes this particular data set. 
And so the problem here is over fitting, 

70
00:04:21,880 --> 00:04:24,030
and you know, he says it's the bugbear of 
machine learning. 

71
00:04:24,030 --> 00:04:27,440
You know a lot of machine learning 
problems really reduced to, how do we 

72
00:04:27,440 --> 00:04:28,640
avoid overfitting? 
Right? 

73
00:04:28,640 --> 00:04:32,870
You could, you can always train a model 
on a particular data set, but, how do you 

74
00:04:32,870 --> 00:04:37,620
avoid, specializing into this data set 
and making sure it has some, some 

75
00:04:37,620 --> 00:04:41,040
predictive power in the real world? 
Okay, and so, you know, the case to look 

76
00:04:41,040 --> 00:04:44,390
out for is, your learner output to 
classify that is 100% accurate on the 

77
00:04:44,390 --> 00:04:48,390
training data, but is only 50% accurate 
on test data, when in fact you could have 

78
00:04:48,390 --> 00:04:52,756
output one that is 75% accurate on both. 
It has overfit. 

79
00:04:52,756 --> 00:04:55,640
Okay. 
And so then I want to call out that this 

80
00:04:55,640 --> 00:04:58,100
is really the, the definition to think of 
when you think about overfitting. 

81
00:04:58,100 --> 00:05:00,708
Low error on training data and high error 
on test data. 

82
00:05:00,708 --> 00:05:05,290
Okay? 
So an image that sometimes is called to 

83
00:05:05,290 --> 00:05:11,460
mind when you're talking about 
overfitting is this one, which is, you 

84
00:05:11,460 --> 00:05:19,220
know, I've, I've fit two polynomials. 
To a small set of data points. 

85
00:05:19,220 --> 00:05:23,240
One with a high degree, I think ten and 
the other with a low degree, in this case 

86
00:05:23,240 --> 00:05:28,770
actually still pretty high as five. 
And you can sort of see that, a couple of 

87
00:05:28,770 --> 00:05:32,380
things. 
One is that the red curve is exactly 

88
00:05:32,380 --> 00:05:35,999
matching the data. 
Okay. 

89
00:05:35,999 --> 00:05:41,500
but you can also see that the sort of, 
what appears to be the trend in the data 

90
00:05:42,910 --> 00:05:45,850
probably is better described by the, the 
green line. 

91
00:05:45,850 --> 00:05:49,499
And, in fact, it was, it was generated by 
a model that sort of added in some random 

92
00:05:49,499 --> 00:05:52,250
noise. 
to a curve, and so the green one probably 

93
00:05:52,250 --> 00:05:57,320
does actually reflect the actual, 
underlying data model, and the underlying 

94
00:05:57,320 --> 00:06:01,837
process which, which generated the data. 
But, you know, the other, the other point 

95
00:06:01,837 --> 00:06:06,570
to look out for here is, it's sensitive 
to change, to perturbations in the data. 

96
00:06:06,570 --> 00:06:11,100
If I move one of these points, just one 
of these points a little bit, how much 

97
00:06:11,100 --> 00:06:14,920
does that curve change. 
And the green one would not change much 

98
00:06:14,920 --> 00:06:19,600
while the red one might, okay. 
But I actually don't think these curve 

99
00:06:19,600 --> 00:06:23,330
fitting model is the best image to call 
up when you think about overfitting 

100
00:06:23,330 --> 00:06:29,341
because it's, it's, in my mind it's a 
little bit specific to polynomial curves 

101
00:06:29,341 --> 00:06:33,850
and the number of degrees of freedom you 
are, you are. 

102
00:06:33,850 --> 00:06:35,170
Working with. 
Okay. 

103
00:06:35,170 --> 00:06:40,380
So it's not always clear, at least to me, 
how to map this image that I call in my 

104
00:06:40,380 --> 00:06:42,750
head of overfitting to machine learning 
problems. 

105
00:06:42,750 --> 00:06:47,875
So I think a more useful one is one 
that's actually associated with the 

106
00:06:47,875 --> 00:06:52,007
Wikipedia page on an article related to 
overfitting. 

107
00:06:52,007 --> 00:06:55,020
that's, released in the Creative Commons, 
is this. 

108
00:06:55,020 --> 00:07:03,180
That over time, your error goes down from 
both your training set and your test set, 

109
00:07:03,180 --> 00:07:05,710
but at some point, it continues to go 
down for your training set. 

110
00:07:05,710 --> 00:07:11,480
I'm pointing at the wrong screen. 
This continues to go down on your 

111
00:07:11,480 --> 00:07:14,670
training set, but then it starts to creep 
up on your test set, and so, at this 

112
00:07:14,670 --> 00:07:18,150
point, when you put a little symbol here 
to indicate, that's where over fitting 

113
00:07:18,150 --> 00:07:19,800
has set in. 
And so, I think, this is the image you 

114
00:07:19,800 --> 00:07:22,461
call to mind when you are thinking about 
overfitting problems.This is this 

115
00:07:22,461 --> 00:07:25,460
diference between the error in your 
training set and the error in your test 

116
00:07:25,460 --> 00:07:27,490
set. 
Now, in some cases it's tough, because 

117
00:07:27,490 --> 00:07:35,420
the test set, you know, maybe. 
you, you know, what is the test set, is 

118
00:07:35,420 --> 00:07:38,350
it something you actually have your hands 
on where you can measure this error, or 

119
00:07:38,350 --> 00:07:40,344
is it more like the predictive power in 
the real world? 

120
00:07:40,344 --> 00:07:46,310
But regardless, any kind of estimate you 
have of, over the error that you're 

121
00:07:46,310 --> 00:07:50,670
actually achieving on data that you 
didn't train on, is, is what to look out 

122
00:07:50,670 --> 00:07:56,051
for, and when these start to diverge, 
that means you were were good, all right? 

123
00:07:56,051 --> 00:07:59,300
So other language you should be familiar 
with around this concept, is you know the 

124
00:07:59,300 --> 00:08:02,520
model able to generalize? 
That's when it's, that's when, if, if 

125
00:08:02,520 --> 00:08:05,290
you're, if it is able to generalize then 
you're not over fit, okay? 

126
00:08:05,290 --> 00:08:08,096
Can it deal with unseen data or does it 
over fit the data it test on. 

127
00:08:08,096 --> 00:08:13,796
So in order to solve for this, you test 
on Hold-out data. 

128
00:08:13,796 --> 00:08:16,140
Okay. 
So one way to do this is, split the data 

129
00:08:16,140 --> 00:08:19,720
to be modeled into training and test set. 
Just in a fixed way. 

130
00:08:19,720 --> 00:08:22,700
Train the model in the training set. 
Evaluate a model in the training set. 

131
00:08:22,700 --> 00:08:26,490
And evaluate the model in the test set. 
And the difference between those two, 

132
00:08:26,490 --> 00:08:30,224
again, is, is, a measure of the ability 
of the model's ability to generalize. 

133
00:08:30,224 --> 00:08:35,696
You know a measure of how over fit it is, 
okay. 

134
00:08:35,696 --> 00:08:44,710
Now, doing this just once, splitting 
between [UNKNOWN] is, not the most 

135
00:08:44,710 --> 00:08:48,620
powerful mechanism to do this. 
And there are a couple a slides, give you 

136
00:08:48,620 --> 00:08:57,910
a better one. 
Alright, so another image to kind of call 

137
00:08:57,910 --> 00:09:01,178
to mind when you're thinking about 
overfitting is and also comes from[ Pedro 

138
00:09:01,178 --> 00:09:06,880
Domingo's paper, which is this is 
sanction between bias and various, 

139
00:09:06,880 --> 00:09:11,430
variance, excuse me. 
So, in the underfit case, right, you, 

140
00:09:11,430 --> 00:09:15,850
you're missing the mark, you're not 
describing the data very well. 

141
00:09:15,850 --> 00:09:19,160
But you might have low variance. 
You know you're, you're, you're clearing 

142
00:09:19,160 --> 00:09:23,130
missing the mark. 
And then, you know high bias and high 

143
00:09:23,130 --> 00:09:26,250
variance well, then you're sort of just 
way off the mark. 

144
00:09:26,250 --> 00:09:28,960
You haven't, you know, you're, your 
learner has not producing anything useful 

145
00:09:28,960 --> 00:09:31,630
at all. 
but in the over-fit case, you might have 

146
00:09:31,630 --> 00:09:35,512
very, very low bias, but you have high 
variance. 

147
00:09:35,512 --> 00:09:40,060
Now, the way to interpret this language 
I'm using, high variance, is variance 

148
00:09:40,060 --> 00:09:43,360
when you, when you add more data, or when 
you perturb the data, or when you 

149
00:09:43,360 --> 00:09:47,720
evaluate on a different data set. 
How my, how, what's your quality look 

150
00:09:47,720 --> 00:09:51,020
like, right, how well do you do? 
Okay. 

151
00:09:51,020 --> 00:10:00,530
And so if you have low bot-, so, if you 
have low bias on your test set, but it's 

152
00:10:00,530 --> 00:10:04,470
sensitive to the data, alright, that's 
where the high variance comes from, then 

153
00:10:04,470 --> 00:10:07,770
you're in an over-fit scenario. 
Okay. 

154
00:10:07,770 --> 00:10:11,530
And in the other case, while you're, 
you're not very sensitive to the data, 

155
00:10:11,530 --> 00:10:14,960
you have low variance, right. 
But you're biased, but you have high 

156
00:10:14,960 --> 00:10:17,320
bias, you're wrong. 
You're getting the wrong answer. 

157
00:10:17,320 --> 00:10:21,870
So this is like, if I just have a model 
that's always predicting, you know, let's 

158
00:10:21,870 --> 00:10:25,180
say I have a model that says, every, on 
the [UNKNOWN] data set that everybody 

159
00:10:25,180 --> 00:10:28,010
dies all the time. 
While this is very insenstive to the 

160
00:10:28,010 --> 00:10:32,630
data, right, it has low variance, right. 
No matter what data set I give it, I 

161
00:10:32,630 --> 00:10:34,020
would either always produce the same 
answer. 

162
00:10:34,020 --> 00:10:37,855
But it has high bias, it's wrong. 
Right, the, the error is, is high, while 

163
00:10:37,855 --> 00:10:41,860
over-fitting is the other way around. 
We've gotta get an exact answer out of my 

164
00:10:41,860 --> 00:10:45,600
training set, where I have very, very low 
bias, but as soon as I add one more data 

165
00:10:45,600 --> 00:10:49,780
point, it it changes drastically. 
Okay. 

166
00:10:49,780 --> 00:10:57,820
So this is another set of terms and kind 
of concepts to think about when you think 

167
00:10:57,820 --> 00:11:04,322
about overfitting. 
Alright. 

