1
00:00:00,000 --> 00:00:05,417
Let's now see how to generate these
snippets or other summaries from single

2
00:00:05,417 --> 00:00:11,605
documents. Now here's an example of a
snippet, This is again coming from Google.

3
00:00:11,605 --> 00:00:16,231
I've asked a question, was cast metal
movable type invented in Korea? And

4
00:00:16,231 --> 00:00:21,638
Google's giving me three little snippets
with the answer and you can see that it's

5
00:00:21,638 --> 00:00:26,756
bold faced cases where words in my query
occurred in the snippet. And you can see

6
00:00:26,756 --> 00:00:31,342
the use of dot, dot, dots in the
snippets telling you that it's combining

7
00:00:31,342 --> 00:00:36,051
pieces from different pages, different,
excuse me, different places in the page.

8
00:00:36,051 --> 00:00:41,126
So let's see how these kind of snippets and
other kind of summaries based on single

9
00:00:41,126 --> 00:00:45,965
documents are built. You can think of any
summarization algorithm as having three

10
00:00:45,965 --> 00:00:50,182
stages. The first stage is content
selection: extracting the sentences that

11
00:00:50,182 --> 00:00:55,329
we need from the document. So we have some
document as input, and we're gonna need to

12
00:00:55,329 --> 00:01:00,486
extract sentences. So we might segment our
sentences off. Maybe we use sentences with

13
00:01:00,486 --> 00:01:05,658
periods, full sentences. Maybe we use some
kind of moving window. So we've extracted

14
00:01:05,658 --> 00:01:11,247
some kind of little pieces, little
sentences, and now, from this segmented

15
00:01:11,247 --> 00:01:16,003
set of sentences we want to pick the ones
that are important so I have marked those

16
00:01:16,003 --> 00:01:20,559
with little, little black dots here. We
picked some set of extracted sentences and

17
00:01:20,559 --> 00:01:24,956
our next task information ordering is we
are gonna decide what order the sentences

18
00:01:24,956 --> 00:01:29,614
go in. So we have some now ordered set
of important sentences and finally we

19
00:01:29,614 --> 00:01:33,720
might do some modifications to the
sentences; perhaps we're gonna simplify

20
00:01:33,720 --> 00:01:38,050
them or something else. So that's sentence
realization and the result of these four

21
00:01:38,050 --> 00:01:43,018
steps is our summary. Now, the most basic
summarization algorithm, the one that

22
00:01:43,018 --> 00:01:47,883
comes up a lot, really only uses one of
these three steps, the content selection

23
00:01:47,883 --> 00:01:52,317
step. So in the simplest possible
algorithm we don't worry about what

24
00:01:52,317 --> 00:01:56,935
ordering the sentences come in and we
don't modify the sentences at all. We

25
00:01:56,935 --> 00:02:01,738
simply segment our document into sentences
or maybe their windows. We pick the

26
00:02:01,738 --> 00:02:06,625
important ones and we leave them in the
same order they came in. So we're gonna

27
00:02:06,625 --> 00:02:11,791
use what we call document order with the
original sentences. So this is a very simple

28
00:02:11,791 --> 00:02:16,454
baseline for summarization and one that
most web based snippet generation

29
00:02:16,454 --> 00:02:20,595
algorithms certainly use. The most
commonly used algorithm for content

30
00:02:20,595 --> 00:02:25,327
selection dates back to the very earliest
paper in the field, from 1958, it's pretty

31
00:02:25,327 --> 00:02:29,659
exciting that these ideas came out so
early. And the intuition's really very

32
00:02:29,659 --> 00:02:34,220
simple. Choose sentences that have salient
or informative words. Well, what's that

33
00:02:34,220 --> 00:02:38,040
mean? Well, you've seen TFIDF. That's a
way of picking words that are

34
00:02:38,040 --> 00:02:42,692
particularly frequent, and then don't
contain words that occur in all documents. So

35
00:02:42,692 --> 00:02:47,638
that's one way we might define saliency or
informativity. Turns out in summarization

36
00:02:47,638 --> 00:02:52,467
we tend to use another approach. The log
likelihood ratio or, sometimes called

37
00:02:52,467 --> 00:02:56,923
topic signature approach. And, that
differs from TFIDF in two ways. One,

38
00:02:56,923 --> 00:03:01,570
we use a slightly different statistic for
picking, for weighting each of the words.

39
00:03:01,570 --> 00:03:05,737
And second, instead of picking all the
words, we'll choose only the words whose

40
00:03:05,737 --> 00:03:10,578
weight is above some threshold, The very
salient words. Now, log likelihood ratio

41
00:03:10,578 --> 00:03:15,339
gives us a statistic called lambda. I'm
not gonna go into details, but, they're in

42
00:03:15,339 --> 00:03:19,846
some lovely papers. And we're gonna choose
all words for who the value of two log

43
00:03:19,846 --> 00:03:24,806
lambda is greater than in this cutoff of
ten. So that gives us a threshold for

44
00:03:24,806 --> 00:03:29,122
which we can pick words that are
particularly salient by this statistic. So

45
00:03:29,122 --> 00:03:33,236
we're gonna weight every word. And the
weight of word I is gonna be one, if the

46
00:03:33,236 --> 00:03:37,934
word is especially associated with that
document, meaning, occurs more times in

47
00:03:37,934 --> 00:03:41,755
that document than in the background
corpus by some threshold. Otherwise, we're

48
00:03:41,755 --> 00:03:46,167
gonna, we're gonna give it a weight of
zero. And again, for details about how to

49
00:03:46,167 --> 00:03:50,650
compute the log likelihood and intuition
about the statistics, you can see this

50
00:03:50,650 --> 00:03:55,559
lovely Ted Dunning paper, or the Lin and
Hovy paper that proposed using it for

51
00:03:55,559 --> 00:04:01,813
summarization. Now, we want to modify this
algorithm for dealing with query focused

52
00:04:01,813 --> 00:04:06,089
summarization. Again, we're not interested
so much in pure summarization in today's

53
00:04:06,089 --> 00:04:10,904
lecture, but how to use summarization
techniques for question answering. So this

54
00:04:10,904 --> 00:04:15,083
is topic signature based. Topic signature
meaning, pick the words that are

55
00:04:15,083 --> 00:04:19,656
particularly associated with a, with a
document. Content selection, choosing the

56
00:04:19,656 --> 00:04:23,722
sentences where we've got queries,
alright? And so we're gonna modify the

57
00:04:23,722 --> 00:04:27,839
algorithm very slightly. We're gonna
choose words that are informative, either

58
00:04:27,839 --> 00:04:32,359
by log likelihood ratio, or words that
happen to appear in the query. So, we're

59
00:04:32,359 --> 00:04:36,454
going to weight every word in a document.
We're going to give it a weight of one if

60
00:04:36,454 --> 00:04:40,349
it meets the log likelihood threshold,
it passes the threshold of about ten.

61
00:04:40,349 --> 00:04:44,244
We're going to give it a weight of one,
also, if that word happens to appear in

62
00:04:44,244 --> 00:04:48,340
the query or question, and otherwise we're
going to give the word a weight of zero.

63
00:04:48,340 --> 00:04:52,148
And these weights are very simple, one,
one zero you could imagine learning more

64
00:04:52,148 --> 00:04:56,692
complex weights and some research has gone
into coming up with very powerful ways to

65
00:04:56,692 --> 00:05:00,300
learn detailed weights. But one, one zero
works pretty well, it turns out. And now

66
00:05:00,300 --> 00:05:04,214
we're just going to weigh a sentence or
perhaps it's a window, if we don't have

67
00:05:04,214 --> 00:05:07,767
actual sentences. We're going to weigh it
by the weight of the words, so we're just

68
00:05:07,767 --> 00:05:11,576
going to sum over all the words in our
sentence of the weight of the words and

69
00:05:11,576 --> 00:05:16,829
then we're going to take the average. Now
the content selection algorithm we just

70
00:05:16,829 --> 00:05:21,755
described is unsupervised. We didn't
have any labeled training set of which of

71
00:05:21,755 --> 00:05:26,497
summaries to learn weights from or things
like that. So that's an alternative

72
00:05:26,497 --> 00:05:31,731
approach: supervised content selection. So
now if we had a labeled training set where

73
00:05:31,731 --> 00:05:36,843
for each document we had a good summary
and we had an alignment for every sentence

74
00:05:36,843 --> 00:05:41,954
in the summary we knew what sentence it
came from in the document, we had the matching

75
00:05:41,954 --> 00:05:46,274
sentences. Now we could extract all sorts of
features. We could extract the position of

76
00:05:46,274 --> 00:05:49,978
the sentence in the document. First
sentences are very likely to be good

77
00:05:49,978 --> 00:05:54,145
summary sentences. How long it is. We can
have all the features we had before, word

78
00:05:54,145 --> 00:05:58,363
informativeness and things like that. We
can have other kinds of features based on

79
00:05:58,363 --> 00:06:02,839
discourse information that we might have.
And we might associate every sentence with

80
00:06:02,839 --> 00:06:06,852
some vector of features, and now we can
just train a binary classifier. Shall I

81
00:06:06,852 --> 00:06:10,864
put this sentence in summary? Yes or no.
And it might learn weights for all these

82
00:06:10,864 --> 00:06:15,437
features and any other features we can
come up with. And the algorithm, this

83
00:06:15,437 --> 00:06:20,107
sounds good but in practice it turns out
to be very hard to get labeled training

84
00:06:20,107 --> 00:06:24,823
data of this type. When people actually
write abstracts for sentences they're not

85
00:06:24,823 --> 00:06:30,017
always, the authors don't always use exact
words and phrases and certainly don't use

86
00:06:30,017 --> 00:06:34,487
entire sentences that come from the
document, so finding perfectly labeled

87
00:06:34,487 --> 00:06:41,236
abstracts with extracts from the document
is hard. It's hard to do the alignment

88
00:06:41,236 --> 00:06:44,857
because they don't pick entire sentences,
they may be picking words or phrases or

89
00:06:44,857 --> 00:06:48,441
chunks. It's hard to figure out where those
words came from, even when they did pick

90
00:06:48,441 --> 00:06:52,841
them from the document. And it turns out,
surprisingly perhaps, that the performance

91
00:06:52,841 --> 00:06:58,544
is simply not much better than
unsupervised algorithms. So in practice

92
00:06:58,544 --> 00:07:03,616
unsupervised content selection, just using
log likelihood ratio, or other simple

93
00:07:03,616 --> 00:07:10,687
measures of how salient or informative
a word and hence a sentence is, are

94
00:07:10,687 --> 00:07:16,773
the most common method for content
selection. So we've seen how to generate

95
00:07:16,773 --> 00:07:22,012
summaries from a single document and the
baseline algorithm we picked is simply

96
00:07:22,012 --> 00:07:27,188
come up with a simple statistical way to
find a sentence that is very informative

97
00:07:27,188 --> 00:07:32,301
by looking for informative words and we
talked about the log-likelihood ratio as

98
00:07:32,301 --> 00:07:35,079
an important way of finding these
sentences.
