1
00:00:00,000 --> 00:00:04,954
In this video, I'd like to talk about how
to evaluate a hypothesis that has been

2
00:00:04,954 --> 00:00:09,908
[inaudible] by your algorithm. In later
videos, we'll build on this, to talk about

3
00:00:09,908 --> 00:00:14,615
how to prevent the problems of over
fitting and under fitting as well. When we

4
00:00:14,615 --> 00:00:19,136
fit the parameters of our learning
algorithm, we think about choosing the

5
00:00:19,136 --> 00:00:24,338
parameters to minimize the training error.
One might think that getting a really low

6
00:00:24,338 --> 00:00:29,231
value of training error might be a good
thing. But we've already seen that, just

7
00:00:29,231 --> 00:00:34,000
because a hypothesis has low training
error, that doesn't mean it's necessary.

8
00:00:34,000 --> 00:00:39,828
Hypothesis and we already seen the example
of how hypothesis can over fit, and there

9
00:00:39,828 --> 00:00:45,448
fulfill to generalize the new example that
not in the training set. So how do you

10
00:00:45,448 --> 00:00:51,138
tell if hypothesis might be over fitting?
In this simple example we could plot the

11
00:00:51,138 --> 00:00:56,689
hypothesis each of x and just see what is
going on, that in general for problems

12
00:00:56,689 --> 00:01:02,171
with more features than just one feature,
for problems with the large number of

13
00:01:02,171 --> 00:01:07,236
features like this, it becomes hard or
maybe impossible to caught weather

14
00:01:07,236 --> 00:01:12,422
hypothesis. Practices functions looks
like. And so, we need some other way to

15
00:01:12,422 --> 00:01:17,882
evaluate our hypothesis. The standard way
to evaluate a direct hypothesis is as

16
00:01:17,882 --> 00:01:22,995
follows: Suppose we have a data set like
this, here up to show ten training

17
00:01:22,995 --> 00:01:28,800
examples but of course usually we may have
dozens, or hundreds, or maybe thousands of

18
00:01:28,800 --> 00:01:34,260
training examples. In order to make sure
we can evaluate our hypothesis what we

19
00:01:34,260 --> 00:01:39,995
going to do is split the data we have into
two portions. The first portion is going

20
00:01:39,995 --> 00:01:45,787
to be our usual training set. And the
second portion is going to be our test

21
00:01:45,787 --> 00:01:52,154
set. And a pretty typical split of this,
of all the data we have into a training

22
00:01:52,154 --> 00:01:57,634
set and test set, might be around, say, a
70%, 30 percent split. With more of the

23
00:01:57,634 --> 00:02:03,597
data going to the training set, and
relatively less to the test set. And so,

24
00:02:03,597 --> 00:02:09,053
now. If we have some data set. We will
assign only say 70 percent of the data to

25
00:02:09,053 --> 00:02:15,106
be our training set. Where here M is as
usual a number in our training examples

26
00:02:15,106 --> 00:02:21,312
and the remainder of our data might that
be assigned to become our test sets. And

27
00:02:21,312 --> 00:02:27,135
here I'm going to use the notation M
subscript test to denote the number of

28
00:02:27,135 --> 00:02:33,264
test examples. And so, in general this.
Subscript test is going to denote examples

29
00:02:33,264 --> 00:02:39,038
that come from our test set so that X1
subscript test comma Y1 subscript test is

30
00:02:39,038 --> 00:02:44,669
my first test example which I guess in
this example, might be this example over

31
00:02:44,669 --> 00:02:50,300
here. Finally, one last detail. Where as
here I've drawn this as though the first

32
00:02:50,300 --> 00:02:55,147
70 percent goes to the trading set and the
last 30 percent to the test set if there

33
00:02:55,147 --> 00:03:00,636
is any sorted ordering to the data. That
should be better to send a random 70

34
00:03:00,636 --> 00:03:05,619
percent of your data to the twenty-thirds
and a random 30%. Your data to the test

35
00:03:05,619 --> 00:03:10,330
set. So, if your data were already
randomly sorted you could just take the

36
00:03:10,330 --> 00:03:15,235
first 70 percent and last 30%. But if your
data were not randomly ordered it would be

37
00:03:15,235 --> 00:03:20,656
better to randomly shuffle or to randomly
reorder the examples in your training set.

38
00:03:20,656 --> 00:03:25,173
Before you know, sending the first 70
percent to the training set and the last

39
00:03:25,173 --> 00:03:30,581
30 percent to the test set. Here then is a
fairly typical procedure for how you would

40
00:03:30,581 --> 00:03:35,770
train and test a learning algorithm, maybe
linear regression. First you learn the

41
00:03:35,770 --> 00:03:41,023
parameters theta from your training set,
so you minimize the usual training error

42
00:03:41,023 --> 00:03:45,693
objective j of theta, where j of theta
here was defined using that 70 percent of

43
00:03:45,693 --> 00:03:50,947
all the data you have. So that's only the
training data. And then you will compute

44
00:03:50,947 --> 00:03:56,200
the test error, and I'm going to denote
the test error as j subscript test, and so

45
00:03:56,200 --> 00:04:02,269
what you do is you take your parameter.
Data that you've learned from the training

46
00:04:02,269 --> 00:04:08,938
set and plug it in here. And compute your
test set error. Which I'm going to write

47
00:04:08,938 --> 00:04:15,277
as follows. So, this is basically the
average square error as measured on your

48
00:04:15,277 --> 00:04:21,616
test set. This is pretty much what you
would expect. So, run every test example

49
00:04:21,616 --> 00:04:28,119
through your hypothesis with parameter
data and just measure the squared error

50
00:04:28,119 --> 00:04:34,184
the hypothesis has on you M subscript
test, test exam. Apples. And of course

51
00:04:34,184 --> 00:04:39,793
this is the definition of the test set
error if we are using mini regression and

52
00:04:39,793 --> 00:04:45,609
using this squared error metric. How about
if we were doing a classification problem.

53
00:04:45,609 --> 00:04:51,149
And say using logistic regressions then.
In that case the procedure for training

54
00:04:51,149 --> 00:04:56,620
and testing, say logistic regressions is
pretty similar. First, we will learn the

55
00:04:56,620 --> 00:05:01,605
parameters from the training data. That
first 70 percent of the data and then we

56
00:05:01,605 --> 00:05:07,166
will compute the test errors as follows.
It is the same objective function. As we

57
00:05:07,166 --> 00:05:12,918
always use for logistic aggression, and
now it is defined using our M subscript

58
00:05:12,918 --> 00:05:18,451
test, test examples. While this definition
of the tested era J [inaudible] is

59
00:05:18,451 --> 00:05:24,276
perfectly reasonable, sometimes there's an
alternative test set [inaudible] that

60
00:05:24,276 --> 00:05:29,954
might be easier to interpret, and that's
the misclassification error. It's also

61
00:05:29,954 --> 00:05:35,706
called 01 misclassification error, 01
denoting that either get an example right

62
00:05:35,706 --> 00:05:42,061
or you get an example wrong. Here's what I
mean, define the error of a. That is each

63
00:05:42,061 --> 00:05:50,868
of x and give him the label y as equals to
one if my hypothesis opens a value greater

64
00:05:50,868 --> 00:05:59,266
than the amount of five and y is equals to
zero, or, if my hypothesis opens a value

65
00:05:59,266 --> 00:06:07,561
is 0.5 and y is equals to one. Right so
proof of this case is basic respond to, if

66
00:06:07,561 --> 00:06:15,380
you hypothesis mislabel the example
assuming you touch hold it to 0.5. Either

67
00:06:15,380 --> 00:06:22,904
thought it was more likely to be one, but
is was actually zero. Or, your hypothesis

68
00:06:22,904 --> 00:06:29,593
[inaudible] zero, but the label was
actually one. And otherwise, we define

69
00:06:29,593 --> 00:06:36,560
this error function to be zero. If, your
hypothesis [inaudible] example Y

70
00:06:36,560 --> 00:06:44,085
correctly. We could then define the test
error using the [inaudible] error metric

71
00:06:44,085 --> 00:06:51,484
to be one of M tests of sum from I=1 to M
subscript. Rest of the error of each of XI

72
00:06:51,484 --> 00:06:58,456
test, [inaudible]. And so that is just one
way of writing out that this is exactly

73
00:06:58,456 --> 00:07:05,170
the fraction of the examples is my test
set. That my hypothesis has mislabeled,

74
00:07:05,170 --> 00:07:11,712
and so that's the definition of the
[inaudible] using the misclassification

75
00:07:11,712 --> 00:07:18,598
error, or the 01 misclassification error
metric. So that's the standard technique

76
00:07:18,598 --> 00:07:24,150
for evaluating how good a learned
hypothesis is. And the next video will

77
00:07:24,150 --> 00:07:28,943
adapt these ideas helping us do things,
like choose what features like degrees of

78
00:07:28,943 --> 00:07:33,204
polynomials to use in the learning
algorithm or choose deregularization

79
00:07:33,204 --> 00:07:35,157
parameters for the new algorithm.
