1
00:00:00,000 --> 00:00:04,501
We talked earlier about evaluating
question answering. If we have the answer

2
00:00:04,501 --> 00:00:08,637
to a factoid question we can simply
evaluate it by seeing if the factoid

3
00:00:08,637 --> 00:00:11,911
the system returns is the correct
factoid. Now summaries can't be

4
00:00:11,911 --> 00:00:17,947
evaluated that way because we can't have a
single perfect summary for any document.

5
00:00:17,947 --> 00:00:22,651
So we'll introduce a different algorithm
called ROUGE. Rouge stands for

6
00:00:22,651 --> 00:00:28,420
Recall Oriented Understudy for Gisting
Evaluation, proposed by Lin and Hovy. And,

7
00:00:28,420 --> 00:00:34,116
here's the idea; it's is an intrinsic
me-, metric for evaluating summaries, we're

8
00:00:34,116 --> 00:00:39,080
gonna ask is this summary good as a
summary. Not in some other extrinsic

9
00:00:39,080 --> 00:00:44,020
application, but just as a summary. And
it's based on a metric called bleu, or

10
00:00:44,020 --> 00:00:49,210
blue, that defined originally for machine
translation. And it's not as good, rouge

11
00:00:49,210 --> 00:00:54,337
is not as good as using, as using humans
to say, did this summary answer the user's

12
00:00:54,337 --> 00:00:59,339
question? So if we can afford that, we'll
certainly hire users, and have them test

13
00:00:59,339 --> 00:01:04,278
to see if an answer answers a user's
question. But rouge is very convenient for

14
00:01:04,278 --> 00:01:09,671
testing while we're building our system.
And it works as follows. We're given a

15
00:01:09,671 --> 00:01:15,453
document D and let's say we've got our
summarizer and it produces an automatic

16
00:01:15,453 --> 00:01:20,521
summary. And this can be a query focused
summary, so we maybe know about the, the

17
00:01:20,521 --> 00:01:26,451
query that the user asked or even for
generic summarization. Now we have N

18
00:01:26,451 --> 00:01:30,191
humans produce a set of reference
summaries of this document. Again, in

19
00:01:30,191 --> 00:01:34,203
query focused summarization, they look at
the query and write their summaries. In

20
00:01:34,203 --> 00:01:37,714
generic summarization, they just write
their summaries. So we have a set of

21
00:01:37,714 --> 00:01:42,772
summaries, one, two, three, four human
summaries, of our document D. And now the

22
00:01:42,772 --> 00:01:49,918
system produces another summary, call
that x. So we have our automatic summary

23
00:01:49,918 --> 00:01:54,983
and our four human summaries and now we
just ask, what percentage of the bigrams

24
00:01:54,983 --> 00:01:59,513
from these humans summaries occur in X?
Obviously they won't all occur in x. A

25
00:01:59,513 --> 00:02:04,278
good summary will contain a lot of the
bigrams that occur in some of these human

26
00:02:04,278 --> 00:02:08,926
summaries. So by counting the percentage
you get an intuition for what's a good

27
00:02:08,926 --> 00:02:13,515
summary. And there are various versions of
rouge, unigram rouge, bigram rouge,

28
00:02:13,515 --> 00:02:18,221
there's also other versions that, that
talk about length in different ways. We'll

29
00:02:18,221 --> 00:02:22,634
introduce just one rouge two, bigram
rouge, which works pretty well and it's

30
00:02:22,634 --> 00:02:28,297
just asking, out of all the bigrams, in all
the sentences, in all the reference

31
00:02:28,297 --> 00:02:34,214
summaries, take the count of those
bigrams. And notice that out of all those

32
00:02:34,214 --> 00:02:43,093
bigrams look at, again, in each sentence
in this summary, for each bigram ask

33
00:02:43,093 --> 00:02:49,870
what's its minimum count in the, the
summary produced by our system and the

34
00:02:49,870 --> 00:02:57,122
human summary. So it's asking how many
bigrams occurred both in our system and in

35
00:02:57,122 --> 00:03:02,863
the human summary. So that will give us,
out of all bigrams how many occurred both

36
00:03:02,863 --> 00:03:09,910
in our, in the summary and in the human
references. Let's look at an example. Here

37
00:03:09,910 --> 00:03:16,210
I have made up three human summaries and a
system summary, so here's our human three

38
00:03:16,210 --> 00:03:21,688
summaries and our system answer to a
question. Let's compute rouge, so the

39
00:03:21,688 --> 00:03:27,572
numerator we want to know how many bigrams
that occurred in these human summaries, how

40
00:03:27,572 --> 00:03:33,185
many of those bigrams also occurred in our
system answer. We can walk through, well

41
00:03:33,185 --> 00:03:41,207
water spinach. That's in our answer. Here it is
here. And spinach is, and is a. And in this

42
00:03:41,207 --> 00:03:46,781
summary water spinach and spinach is, and
is a, and in this third summary again,

43
00:03:46,781 --> 00:03:52,854
water spinach and spinach is and is a, but
also commonly eaten and leaf vegetable and

44
00:03:52,854 --> 00:03:58,642
of Asia. So if we add all that together,
we have three from this first summary and

45
00:03:58,642 --> 00:04:04,463
three from the second summary and six from
the third. And how many total bigrams are

46
00:04:04,463 --> 00:04:08,961
there in the human summary? Well, you can
count them yourself. There are ten in the

47
00:04:08,961 --> 00:04:12,574
first example up here, nine in here,
nine in here. So we have

48
00:04:12,574 --> 00:04:23,259
(3+3+6)/(10+9+9) or a rouge score of .43. So
we've introduced rouge, an algorithm for

49
00:04:23,259 --> 00:04:29,404
evaluating summaries, whether they're generic
or query focused, by looking at how many

50
00:04:29,404 --> 00:04:35,711
of the bigrams, or N grams in general, in
a human summary occur in our machine

51
00:04:35,711 --> 00:04:40,317
generated summary, and a better summary is
one that overlaps more with the human

52
00:04:40,317 --> 00:04:44,833
summary. Now rouge doesn't work as well as
having humans actually answer the

53
00:04:44,833 --> 00:04:50,069
question, did this answer provide the
information the user asked for. But, that

54
00:04:50,069 --> 00:04:55,014
can be very expensive, and so rouge
can provide a fast-to-run and convenient

55
00:04:55,014 --> 00:04:59,898
intrinsic metric that we can use to test
our systems, and then, at the end we can

56
00:04:59,898 --> 00:05:02,524
use humans to see how well it really
did.
