1
00:00:00,000 --> 00:00:04,108
If you run a learning algorithm and it
doesn't do as well as you are hoping,

2
00:00:04,108 --> 00:00:08,486
almost all the time it'll be because you
have either a high finance problem or a

3
00:00:08,486 --> 00:00:12,810
high variance problem. In other words,
either an under fitting problem or an over

4
00:00:12,810 --> 00:00:16,918
fitting problem. And in this case it's
very important to figure out which of

5
00:00:16,918 --> 00:00:21,351
these two problems, is it high or variance
or a bit of both that you actually have

6
00:00:21,351 --> 00:00:25,891
because knowing which of these two things
is happening will give a strong indicator

7
00:00:25,891 --> 00:00:30,270
for whether their useful and how much they
[inaudible] in trying to improve your

8
00:00:30,270 --> 00:00:34,473
algorithm. In this video, I'd like to
delve more deeply into this bias. Various

9
00:00:34,473 --> 00:00:39,390
issues and understand them better. Let's
what take around how to knock the learning

10
00:00:39,390 --> 00:00:44,365
algorithm and evaluate or diagnose whether
we might have a bias problem or variance

11
00:00:44,365 --> 00:00:48,571
problem since this will be critical to
figuring out how to improve the

12
00:00:48,571 --> 00:00:53,547
performance of the learning algorithm that
you may implement. So you've already seen

13
00:00:53,547 --> 00:00:58,048
this figure a few times where if you fit
too simple hypothesis that gives a

14
00:00:58,048 --> 00:01:02,313
straight line that under-fits the data
complex. If you fit a too complex

15
00:01:02,313 --> 00:01:06,756
hypothesis then that might fit the
training set perfectly but over-fit the

16
00:01:06,756 --> 00:01:12,093
data and this may be. Hypothesis of some
intermediate level of complexities of some

17
00:01:12,093 --> 00:01:17,374
two polynomials of not too low and not
high degree that's just right that gives

18
00:01:17,374 --> 00:01:22,524
you the best generalization for these
options. Now that we're arms with the

19
00:01:22,524 --> 00:01:27,541
notion of train, training and validation
the test says we can understand the

20
00:01:27,541 --> 00:01:32,559
concepts that bios and theories a little
bit better. Concretely lets, let our

21
00:01:32,559 --> 00:01:38,104
training error and cross validation error
be defined as in the previous videos just

22
00:01:38,104 --> 00:01:43,312
say the square error the average square
error as mentioned. On the training sets

23
00:01:43,312 --> 00:01:48,375
all has measured on the cross validation
set. Now let's plot the following figure.

24
00:01:48,375 --> 00:01:53,562
On the horizontal axis, I'm going to plot
the degree of polynomial. So as it goes to

25
00:01:53,562 --> 00:01:58,000
the right, I'm going to, I'm going to be
fitting higher and higher order

26
00:01:58,000 --> 00:02:02,687
polynomials. So will the left of this
figure, where maybe D equals one, we're

27
00:02:02,687 --> 00:02:07,500
going to be fitting very simple figures.
Whereas way here on the right of the

28
00:02:07,500 --> 00:02:12,187
horizontal axis, have much larger values
of D. So a much higher degree of

29
00:02:12,187 --> 00:02:17,487
polynomial. And so here that's going to
correspond to fitting. Much more complex

30
00:02:17,487 --> 00:02:22,836
functions to your training set. Let's look
at the training error and the cross

31
00:02:22,836 --> 00:02:27,342
validation error and plot them on this
figure. Let's start with the training

32
00:02:27,342 --> 00:02:32,086
error. As we increase the degree of the
polynomial, we're going to be able to fit

33
00:02:32,086 --> 00:02:36,533
our training set better and better. And
so, if D=1 then it's a relatively high

34
00:02:36,533 --> 00:02:41,157
training error. If we have a very high
degree polynomial our training error is

35
00:02:41,157 --> 00:02:45,960
going to be really low maybe even zero
because we'll fit the training set really

36
00:02:45,960 --> 00:02:50,644
well. And so as we increase the degree of
polynomial, we find typically that the

37
00:02:50,644 --> 00:02:55,776
training error decreases. So, I'm going to
write J. Subscript three of data there.

38
00:02:55,776 --> 00:03:01,565
Because our training error tends to
decrease with the degree of polynomial

39
00:03:01,565 --> 00:03:07,652
that we fit to the data. Next let's look
at the cross validation error. Or for that

40
00:03:07,652 --> 00:03:13,812
matter, if we look at the test set error
we'll get a pretty similar result as if we

41
00:03:13,812 --> 00:03:19,292
were to plot the cross validation error.
So, we know that if D=1. We're fitting a

42
00:03:19,292 --> 00:03:23,758
very simple function. And so we may be
under fitting the training set. And so

43
00:03:23,758 --> 00:03:28,343
we're going to have a very high cross
validation error. If we fit, you know, an

44
00:03:28,343 --> 00:03:33,107
intermediate degree polynomial, there's,
we have a D equals two in our example on

45
00:03:33,107 --> 00:03:38,049
the previous slide, we're going to have a
much lower cross validation error, because

46
00:03:38,049 --> 00:03:43,110
we're just fitting, finding a much better
fit to the data. And conversely, if D were

47
00:03:43,110 --> 00:03:47,874
too high, so if D took on, say, a value of
four, then we're getting over fitting, and

48
00:03:47,874 --> 00:03:52,669
so we ended with a high value for cross
validation error. So, if you were to. Very

49
00:03:52,669 --> 00:03:57,854
[inaudible] and plot the curve. You might
end up with a curve like that. Where,

50
00:03:57,854 --> 00:04:03,737
that's J.C.V. Of staza. Indicating that
you plot J. Tesla's data, you get

51
00:04:03,737 --> 00:04:09,465
something very similar. And so, this sort
of plot also helps us to better understand

52
00:04:09,465 --> 00:04:13,829
the notions of bias and variance.
Concretely, suppose you've applied a

53
00:04:13,829 --> 00:04:18,447
learning algorithm, and it's not
performing as well as you were hoping. So,

54
00:04:18,447 --> 00:04:23,507
so if your cross-validation set error, or
your test set error is high. How can we

55
00:04:23,507 --> 00:04:28,758
figure out if the learning algorithm is
suffering from high bias or if it's suffer

56
00:04:28,758 --> 00:04:34,130
from high variance? So the setting of the
cause validation error being high,

57
00:04:34,130 --> 00:04:39,789
corresponds to either this regime or this
regime. So this regime on the left

58
00:04:39,789 --> 00:04:45,747
corresponds to a high bias problem. That
is, if you're fitting a overly low order

59
00:04:45,747 --> 00:04:51,556
polynomial, such as a D=1. When we really
needed a higher order polynomial to fit

60
00:04:51,556 --> 00:04:57,513
the data. Whereas in contrast, this regime
corresponds to a high variance problem.

61
00:04:57,513 --> 00:05:03,247
That is, if D, the degree of polynomial
was too large for the data set that we

62
00:05:03,247 --> 00:05:08,535
have. And this figure just has a clue for
how to distinguish between these two

63
00:05:08,535 --> 00:05:15,790
cases. Concretely for the high bias case.
That is the case of [inaudible]. What we

64
00:05:15,790 --> 00:05:21,790
find is that both the cross validation
error and the trading error are going to

65
00:05:21,790 --> 00:05:29,770
be high. So if your algorithm is suffering
from a bias problem. The training set

66
00:05:29,770 --> 00:05:38,162
error, will be high. And you might find
that the cross validation error will also

67
00:05:38,162 --> 00:05:45,271
be high. It might be a close. Maybe just
slightly higher than a training error. And

68
00:05:45,271 --> 00:05:52,211
so, if you see this combination that's a
sign your algorithm may be suffering from

69
00:05:52,211 --> 00:05:58,981
high bias. In contrast if your algorithm
is suffering from high variance then if

70
00:05:58,981 --> 00:06:05,921
you look here. We'll notice that J-train
that is the training error is going to be

71
00:06:05,921 --> 00:06:13,626
low. That is your fitting the training set
very well. Where as your, cross validation

72
00:06:13,626 --> 00:06:18,575
error. Assuming that this is say the
squared era. Which we're trying to

73
00:06:18,575 --> 00:06:22,757
minimize [inaudible]. Where as in
contrast, your arrow on the cross

74
00:06:22,757 --> 00:06:28,080
validation set or your cos function in the
cross validation set will be much bigger.

75
00:06:28,080 --> 00:06:32,884
Then your training set error. So, there's
a double greater than sign. That's the

76
00:06:32,884 --> 00:06:37,935
math symbol for much greater than, denoted
by two greater than signs. And so, if you

77
00:06:37,935 --> 00:06:43,047
see this combination of values then that
might give you, that's a clue that your

78
00:06:43,047 --> 00:06:47,482
learning algorithm maybe suffering from
high variance. And might be over

79
00:06:47,482 --> 00:06:51,732
emphasizing. And the key that
distinguishes these two cases is if you

80
00:06:51,732 --> 00:06:56,844
have a high bias problem your training set
error will also be high. Your hypothesis

81
00:06:56,844 --> 00:07:02,080
is just not fitting the training set well.
And if you have a high variance problem.

82
00:07:02,080 --> 00:07:06,852
Your training set error will usually be
low. That is much lower than your cross

83
00:07:06,852 --> 00:07:11,684
allegation error. So hopefully that gives
you a somewhat better understanding of the

84
00:07:11,684 --> 00:07:15,845
two problems of bias and variants. I still
have a lot more to say about bias and

85
00:07:15,845 --> 00:07:19,954
variants in the next few videos. But what
we'll see later is that by diagnosing

86
00:07:19,954 --> 00:07:24,011
whether a learning algorithm may be
suffering from high bias or high variance,

87
00:07:24,011 --> 00:07:28,432
we'll show you even more details of how to
do that in later videos. We'll see that by

88
00:07:28,432 --> 00:07:32,697
figuring out whether a learning algorithm
may be suffering from high bias or high

89
00:07:32,697 --> 00:07:36,806
variance, or a combination of both, that,
that would give us much better guidance

90
00:07:36,806 --> 00:07:41,231
for what might be [inaudible]. Things to
try in order to improve the performance of

91
00:07:41,231 --> 00:07:42,367
a learning algorithm.
