1
00:00:07,970 --> 00:00:11,190
Welcome back. In this video,

2
00:00:11,190 --> 00:00:14,185
we are going to talk about semantic similarity.

3
00:00:14,185 --> 00:00:16,970
Right at the start, I'm going to ask you a question.

4
00:00:16,970 --> 00:00:20,670
Which pair of words are the most similar in the following,

5
00:00:20,670 --> 00:00:22,725
we have deer and elk,

6
00:00:22,725 --> 00:00:27,720
deer and giraffe, deer and horse, and deer and mouse?

7
00:00:27,720 --> 00:00:29,767
What do you think?

8
00:00:29,767 --> 00:00:35,310
Hopefully, you would have given the answer as deer and elk.

9
00:00:35,310 --> 00:00:37,490
But how can you quantify this?

10
00:00:37,490 --> 00:00:41,240
Why do deer and elk appear and sound to be

11
00:00:41,240 --> 00:00:46,284
similar or more similar than other animals here?

12
00:00:46,284 --> 00:00:48,011
To help us with that,

13
00:00:48,011 --> 00:00:51,400
we could use some semantic similarities resources,

14
00:00:51,400 --> 00:00:53,825
and we're going to talk about that in more detail now.

15
00:00:53,825 --> 00:00:58,209
But first, let's see what are the applications of semantic similarity.

16
00:00:58,209 --> 00:01:02,440
Semantic similarity is useful when you're grouping similar words into

17
00:01:02,440 --> 00:01:04,540
semantic concepts into concepts

18
00:01:04,540 --> 00:01:09,100
that have the same meaning - appear to have the same meaning, for example.

19
00:01:09,100 --> 00:01:12,805
Or semantic similarity is very useful

20
00:01:12,805 --> 00:01:16,286
as a building block in natural language understanding tasks.

21
00:01:16,286 --> 00:01:21,111
Tasks such as the textual entailment or paraphrasing.

22
00:01:21,111 --> 00:01:25,455
Paraphrasing is a task where you

23
00:01:25,455 --> 00:01:29,760
rephrase or rewrite some sentence

24
00:01:29,760 --> 00:01:33,090
you get into another sentence that has the same meaning.

25
00:01:33,090 --> 00:01:37,590
Textual entailment, on the other hand, is a little bit more complex.

26
00:01:37,590 --> 00:01:41,350
It says that the smaller sentence or one of

27
00:01:41,350 --> 00:01:48,905
the two sentences derives its meaning or entails its meaning from another piece of text.

28
00:01:48,905 --> 00:01:54,310
So you have a text document or a text passage and a sentence.

29
00:01:54,310 --> 00:01:57,700
And based on the information in the text passage,

30
00:01:57,700 --> 00:02:00,850
you need to say whether the sentence is

31
00:02:00,850 --> 00:02:05,635
correct or it derives its meaning from there or not.

32
00:02:05,635 --> 00:02:09,535
This is a typical task of semantic similarity.

33
00:02:09,535 --> 00:02:14,513
One of the resources useful for semantic similarity is WordNet.

34
00:02:14,513 --> 00:02:21,850
WordNet is a semantic dictionary of words interlinked by semantic relationships.

35
00:02:21,850 --> 00:02:24,490
It is most extensively developed in English,

36
00:02:24,490 --> 00:02:29,050
but there are WordNets available now for quite a few languages.

37
00:02:29,050 --> 00:02:33,345
This WordNet includes a rich linguistic information.

38
00:02:33,345 --> 00:02:36,254
For example, it has the part of speech,

39
00:02:36,254 --> 00:02:38,862
whether something is a noun or an adjective or a verb.

40
00:02:38,862 --> 00:02:43,930
A word senses, different meanings of the same word, synonyms,

41
00:02:43,930 --> 00:02:45,480
other words that mean the same,

42
00:02:45,480 --> 00:02:49,937
hypernyms and hyponyms, that is an is/are relationship.

43
00:02:49,937 --> 00:02:54,405
For example, a deer is a mammal

44
00:02:54,405 --> 00:03:00,605
or meta-name that is a whole and part of relationship and derivationally related forms.

45
00:03:00,605 --> 00:03:04,280
WordNet is also a machine readable and it's freely available,

46
00:03:04,280 --> 00:03:08,790
so it is extensively used in a lot of natural language processing tasks and,

47
00:03:08,790 --> 00:03:12,875
in general, in text mining tasks.

48
00:03:12,875 --> 00:03:16,685
How do you use WordNet for semantic similarity?

49
00:03:16,685 --> 00:03:20,701
WordNet organizes information in a hierarchy, in a tree.

50
00:03:20,701 --> 00:03:27,075
You have a dummy root that is on top of all words of the same part of speech.

51
00:03:27,075 --> 00:03:28,680
So noun has a dummy root.

52
00:03:28,680 --> 00:03:30,290
A verb has a dummy root.

53
00:03:30,290 --> 00:03:34,085
And then, there are many semantic similarity measures

54
00:03:34,085 --> 00:03:36,900
that are using this hierarchy, in some way.

55
00:03:36,900 --> 00:03:42,055
For example, you have different hierarchies for this part of speech.

56
00:03:42,055 --> 00:03:49,834
And let's take an example of our deer that we started with, where deer, elk,

57
00:03:49,834 --> 00:03:55,535
giraffe, horse and so on, these words are grouped together,

58
00:03:55,535 --> 00:03:58,110
in some form, in this hierarchy.

59
00:03:58,110 --> 00:03:59,973
For example, elk, wapiti,

60
00:03:59,973 --> 00:04:03,558
and caribou are all types of deer.

61
00:04:03,558 --> 00:04:06,530
Deer and giraffe are siblings in

62
00:04:06,530 --> 00:04:11,120
this tree hierarchy because they are ruminants, and so on.

63
00:04:11,120 --> 00:04:14,210
And horse's related but not in the same hierarchy.

64
00:04:14,210 --> 00:04:18,038
It's related because horse and deer are ungulates,

65
00:04:18,038 --> 00:04:21,990
but they are not siblings, for example.

66
00:04:21,990 --> 00:04:25,860
So one such measure of using this hierarchy

67
00:04:25,860 --> 00:04:29,806
for defining semantic similarity is path similarity.

68
00:04:29,806 --> 00:04:34,065
You could imagine that you would start with one of these concepts,

69
00:04:34,065 --> 00:04:39,960
and see how many steps you need to take to get to the other.

70
00:04:39,960 --> 00:04:42,870
In other words, you are finding a shortest path

71
00:04:42,870 --> 00:04:46,531
between these two concepts in this hierarchy.

72
00:04:46,531 --> 00:04:48,595
And then, similarity can be just measured

73
00:04:48,595 --> 00:04:52,045
as inversely related to this distance that we computed.

74
00:04:52,045 --> 00:04:56,424
For example, if you have deer and elk,

75
00:04:56,424 --> 00:05:00,210
you would have, the deer and elk, actually are,

76
00:05:00,210 --> 00:05:02,613
have a parent-child relationship in this case,

77
00:05:02,613 --> 00:05:04,805
so the distance is one,

78
00:05:04,805 --> 00:05:10,910
while deer and let's take in another color,

79
00:05:10,910 --> 00:05:14,450
deer and giraffe is the sense of two,

80
00:05:14,450 --> 00:05:17,585
because you need to go up ruminant and down giraffe,

81
00:05:17,585 --> 00:05:21,760
so you have a distance of two.

82
00:05:21,760 --> 00:05:28,030
In general, you can see that when we computed with paths you use this one,

83
00:05:28,030 --> 00:05:30,220
distance of one between deer and elk and say,

84
00:05:30,220 --> 00:05:33,106
it's one over the distance plus one,

85
00:05:33,106 --> 00:05:35,215
so one over two that's .5.

86
00:05:35,215 --> 00:05:40,460
The distance between a deer and giraffe is one over three, so that's,

87
00:05:40,460 --> 00:05:43,705
0.33 and if you just measure the same way,

88
00:05:43,705 --> 00:05:48,370
going from deer to horse you'd say it's one,

89
00:05:48,370 --> 00:05:51,885
two, three, four, five, six.

90
00:05:51,885 --> 00:05:56,950
It's one over seven and that would be 0.14.

91
00:05:56,950 --> 00:06:01,740
The other way to find

92
00:06:01,740 --> 00:06:07,430
similarity between two concepts is using what is called lowest common subsumer.

93
00:06:07,430 --> 00:06:15,930
Lowest common subsumer is that ancestor that is closest to both concepts.

94
00:06:15,930 --> 00:06:19,395
For example, deer and giraffe have

95
00:06:19,395 --> 00:06:25,235
the least common or lowest common subsumer to be ruminants.

96
00:06:25,235 --> 00:06:29,070
You have deer and giraffe and you know that is

97
00:06:29,070 --> 00:06:34,405
the least common subsumer is the one that is an ancestor to both of them,

98
00:06:34,405 --> 00:06:36,505
but the lowest in the hierarchy.

99
00:06:36,505 --> 00:06:42,165
Even though ungulates, and even toward ungulate are both ancestors,

100
00:06:42,165 --> 00:06:45,032
it's the ruminant that is the lowest one in that hierarchy.

101
00:06:45,032 --> 00:06:52,570
With respect to deer and elk,

102
00:06:52,570 --> 00:06:55,930
it's just the deer because deer is a parent for elk

103
00:06:55,930 --> 00:06:59,750
so the one that subsumes both of these would be directive,

104
00:06:59,750 --> 00:07:05,815
and for deer and horse it goes all the way up to ungulate.

105
00:07:05,815 --> 00:07:10,700
Now, you can use this lowest common subsumer notion to find similarity and

106
00:07:10,700 --> 00:07:15,730
that was proposed by Lin and called Lin similarity.

107
00:07:15,730 --> 00:07:19,150
You have similarity measure that is based on the information

108
00:07:19,150 --> 00:07:23,855
contained in the lowest common subsumer of the two concepts.

109
00:07:23,855 --> 00:07:29,680
For example, the formulation for doing that is if you have two concepts u and v,

110
00:07:29,680 --> 00:07:32,215
you take the log of the probability of

111
00:07:32,215 --> 00:07:35,615
this lowest common subsumer and divide it by some of,

112
00:07:35,615 --> 00:07:41,410
log of the probabilities of u and v. And these probabilities are something

113
00:07:41,410 --> 00:07:48,060
that is computed or given by the information content that is learnt over a large corpus.

114
00:07:48,060 --> 00:07:49,800
How do you do all of this in Python?

115
00:07:49,800 --> 00:07:52,970
In Python, especially in NLTK,

116
00:07:52,970 --> 00:07:59,118
you have a lot of semantic similarities already available for use directly.

117
00:07:59,118 --> 00:08:02,625
One, it is very easy to import into Python through NLTK.

118
00:08:02,625 --> 00:08:09,377
You could say import NLTK and from an NLTK corpus import WordNet,

119
00:08:09,377 --> 00:08:16,005
and then you can find appropriate sense of the word that you want to find similarity for.

120
00:08:16,005 --> 00:08:17,392
So for deer you say,

121
00:08:17,392 --> 00:08:22,691
find me the synset of deer which is a noun and give me the first synset,

122
00:08:22,691 --> 00:08:26,205
that's what deer.n.01 means.

123
00:08:26,205 --> 00:08:29,390
It says I want deer in the sense

124
00:08:29,390 --> 00:08:34,770
of given by the noun meaning of it and the first meaning of that.

125
00:08:34,770 --> 00:08:36,575
The same way with elk,

126
00:08:36,575 --> 00:08:41,720
you find the synset that corresponds to elk.n.01 and so on.

127
00:08:41,720 --> 00:08:45,445
Once you have this proper sense of the word,

128
00:08:45,445 --> 00:08:48,014
you could use that to find similarity.

129
00:08:48,014 --> 00:08:53,581
You could say, deer.path _ similarity(elk) or deer.path_similarity(horse).

130
00:08:53,581 --> 00:08:54,835
In this particular case,

131
00:08:54,835 --> 00:08:59,368
you recall that deer and elk were in a parent child relationship,

132
00:08:59,368 --> 00:09:01,200
the similarity was 0.5.

133
00:09:01,200 --> 00:09:06,655
While deer and horse were in two different subtrees,

134
00:09:06,655 --> 00:09:09,560
and the distance was six,

135
00:09:09,560 --> 00:09:11,860
and the similarity was one over seven, that's 1.1428.

136
00:09:11,860 --> 00:09:18,550
Now, if you are using Lin similarity,

137
00:09:18,550 --> 00:09:21,544
you're going to use the information criterion in some way,

138
00:09:21,544 --> 00:09:26,125
and let's say if we use the information criterion that is given by brown clusters.

139
00:09:26,125 --> 00:09:30,480
First, we are going to say from a nltk.corpus import wordnet_ic.

140
00:09:30,480 --> 00:09:36,285
You define brown_ic based on the brown_ic data.

141
00:09:36,285 --> 00:09:39,835
And then say, deer.lin_similarity(elk) using

142
00:09:39,835 --> 00:09:44,080
this brown_ic or the same way with horse with brown_ic,

143
00:09:44,080 --> 00:09:48,872
and you'll see that the similarity there is different.

144
00:09:48,872 --> 00:09:55,320
The Lin similarity is 0.77 for deer and elk and it's 0.86 for deer and horse.

145
00:09:55,320 --> 00:09:57,650
And you'll notice especially here,

146
00:09:57,650 --> 00:10:04,335
that this is not using the distance between two concepts explicitly.

147
00:10:04,335 --> 00:10:05,675
So deer and horse,

148
00:10:05,675 --> 00:10:09,560
that were very far away in the WordNet hierarchy

149
00:10:09,560 --> 00:10:14,520
still get the higher significance and higher similarity between them.

150
00:10:14,520 --> 00:10:15,740
And that is because,

151
00:10:15,740 --> 00:10:21,405
in typical contexts and the information that is contained by these words deer and horse,

152
00:10:21,405 --> 00:10:24,725
you have deer and horse are enough closer in similarity

153
00:10:24,725 --> 00:10:29,680
because they are both basically mammals.

154
00:10:29,680 --> 00:10:34,190
But Elk is a very specific instance of deer and not necessarily,

155
00:10:34,190 --> 00:10:38,090
in the particular Lin similarity doesn't come out as close.

156
00:10:38,090 --> 00:10:43,620
The other different measure

157
00:10:43,620 --> 00:10:48,525
of similarity is using Distributional similarity and Collocations.

158
00:10:48,525 --> 00:10:51,945
Collocations can be defined by this code.

159
00:10:51,945 --> 00:10:54,925
You know a word by the company it keeps.

160
00:10:54,925 --> 00:10:59,520
And that means two words that are frequently appearing in similar concept,

161
00:10:59,520 --> 00:11:02,040
in similar contexts are more likely to

162
00:11:02,040 --> 00:11:05,850
be similar or more likely to be semantically related.

163
00:11:05,850 --> 00:11:09,340
So if you have two words that keep appearing in

164
00:11:09,340 --> 00:11:16,060
very similar contexts or that could replace another word in the similar context,

165
00:11:16,060 --> 00:11:18,190
and still the meaning remains the same,

166
00:11:18,190 --> 00:11:20,890
then they are more likely to be semantically related.

167
00:11:20,890 --> 00:11:24,795
An example is this,

168
00:11:24,795 --> 00:11:29,350
in these four sentences you have something about meeting at a place,

169
00:11:29,350 --> 00:11:32,925
so friends meet at a cafe or Shyam met Ray at a pizzeria

170
00:11:32,925 --> 00:11:38,145
or let's meet up near a coffee shop and so on.

171
00:11:38,145 --> 00:11:43,870
These words, cafe or pizzeria or coffee shop or restaurant are

172
00:11:43,870 --> 00:11:50,545
semantically related because they typically occur around the words meet,

173
00:11:50,545 --> 00:11:53,440
around at, or, near, the.

174
00:11:53,440 --> 00:11:58,615
So there is the determiner right in front of them and there is some notion of location,

175
00:11:58,615 --> 00:12:04,820
and those are the concepts that would form your context around the word.

176
00:12:04,820 --> 00:12:10,940
In general, you would define context based on words before,

177
00:12:10,940 --> 00:12:14,535
after, or within a small window of a target word,

178
00:12:14,535 --> 00:12:16,625
so word what comes before.

179
00:12:16,625 --> 00:12:21,230
For example, for all of these was a cafe and restaurant and so on,

180
00:12:21,230 --> 00:12:23,975
it was 'a' or 'the', alright?

181
00:12:23,975 --> 00:12:27,955
Because it's a noun and you have a determiner right before that.

182
00:12:27,955 --> 00:12:31,460
What comes after or what comes within a small window?

183
00:12:31,460 --> 00:12:34,595
Let's say, of size three and you will remember that

184
00:12:34,595 --> 00:12:38,300
all of those examples had some form of meet.

185
00:12:38,300 --> 00:12:43,935
Met, meet, meeting and so on in that small window of three to five words.

186
00:12:43,935 --> 00:12:47,650
You could also use parts of speech as context,

187
00:12:47,650 --> 00:12:49,575
so part of speech of words before,

188
00:12:49,575 --> 00:12:51,175
after, within a small window.

189
00:12:51,175 --> 00:12:55,780
You could say that this particular target word occurs right

190
00:12:55,780 --> 00:13:01,510
after a determiner or occurs after location morality, two words and so on.

191
00:13:01,510 --> 00:13:05,180
You could have some specific semantic relation to

192
00:13:05,180 --> 00:13:10,745
the target word or you could have words that come from the same sentence,

193
00:13:10,745 --> 00:13:14,820
in same document, and you can define that document as any length you want.

194
00:13:14,820 --> 00:13:19,035
Let's say, a passage in a document,

195
00:13:19,035 --> 00:13:22,000
a paragraph that would constitute your context.

196
00:13:22,000 --> 00:13:28,610
Once you have defined this context you can compute the strength of

197
00:13:28,610 --> 00:13:31,280
association between words based on

198
00:13:31,280 --> 00:13:36,430
how frequently these words co-worker or how frequently they collocate.

199
00:13:36,430 --> 00:13:38,900
That's why it's called Collocations.

200
00:13:38,900 --> 00:13:42,635
For example, if you have two words that keep coming next to each other,

201
00:13:42,635 --> 00:13:47,150
then you would want to say that they are very highly related to each other.

202
00:13:47,150 --> 00:13:50,495
On the other side, if they don't occur together,

203
00:13:50,495 --> 00:13:54,295
then they are not necessarily very similar.

204
00:13:54,295 --> 00:13:58,205
It's also important to see how frequent individual words are.

205
00:13:58,205 --> 00:14:01,630
For example, the word 'the' is so frequent that it

206
00:14:01,630 --> 00:14:06,100
would occur with every other word, fairly often.

207
00:14:06,100 --> 00:14:09,665
The similarity score would be very high

208
00:14:09,665 --> 00:14:14,320
with 'the' just because 'the' itself happens to be very frequent.

209
00:14:14,320 --> 00:14:17,600
There is a way in which you can normalize

210
00:14:17,600 --> 00:14:20,890
such that this very frequent word does not kind of,

211
00:14:20,890 --> 00:14:24,865
super ride all the other similarity measures you find.

212
00:14:24,865 --> 00:14:28,690
And one way to do it would be using Pointwise Mutual Information.

213
00:14:28,690 --> 00:14:32,140
Pointwise Mutual Information is defined as

214
00:14:32,140 --> 00:14:36,695
the log of this ratio of seeing two things together.

215
00:14:36,695 --> 00:14:39,230
Seeing the word and the context together,

216
00:14:39,230 --> 00:14:44,300
divided by the probability of these occurring independently.

217
00:14:44,300 --> 00:14:49,410
What is the chance that you would see the world in the overall corpus?

218
00:14:49,410 --> 00:14:53,890
What is the chance that you can see the context word in the overall corpus and

219
00:14:53,890 --> 00:14:58,570
what is the chance that they are actually occurring together?

220
00:14:58,570 --> 00:14:59,795
This Pointwise Mutual Information is

221
00:14:59,795 --> 00:15:04,700
also something that you can directly call in the NLTK.

222
00:15:04,700 --> 00:15:08,140
You can use NLTK Collocations and Association measures.

223
00:15:08,140 --> 00:15:09,640
For example, you can say,

224
00:15:09,640 --> 00:15:15,250
input NLTK and import the collocations from NLTK.

225
00:15:15,250 --> 00:15:19,285
You can define bigrams as NLTK collocations bigrams,

226
00:15:19,285 --> 00:15:26,260
bigram association measures, and then you can learn that based on a corpus.

227
00:15:26,260 --> 00:15:28,015
In this case, given as text here,

228
00:15:28,015 --> 00:15:30,625
so text corpus and then,

229
00:15:30,625 --> 00:15:36,315
using the PMI measure you can say,

230
00:15:36,315 --> 00:15:46,115
I'm going to get the top 10 pairs using the PMI measure from bigram_measures.

231
00:15:46,115 --> 00:15:51,290
You can use a Use Finder for other useful tasks such as frequency filtering.

232
00:15:51,290 --> 00:15:54,850
So suppose you want all bigram measures that are,

233
00:15:54,850 --> 00:15:59,665
there you have supposed 10 or more occurrences of words only then can you keep them,

234
00:15:59,665 --> 00:16:04,150
then you could do something like finder.apply_ freq_filter (10).

235
00:16:04,150 --> 00:16:12,810
That would then restrict any pair that does not occur at least 10 times in your corpus.

236
00:16:12,810 --> 00:16:17,655
So, the big take home messages from this discussion we have had so far,

237
00:16:17,655 --> 00:16:21,640
is that finding similarity between words and text is non-trivial.

238
00:16:21,640 --> 00:16:23,670
But there are a lot of resources available,

239
00:16:23,670 --> 00:16:26,715
such as WordNet that could be very useful for

240
00:16:26,715 --> 00:16:31,120
semantic relationship between words and semantic similarity between words.

241
00:16:31,120 --> 00:16:36,750
There are many similar functions that are available in WordNet and NLTK provides

242
00:16:36,750 --> 00:16:39,060
a useful mechanism to actually access

243
00:16:39,060 --> 00:16:43,200
the similarity functions and is available for many such tasks,

244
00:16:43,200 --> 00:16:46,290
to find similarity between words or text and so on.

245
00:16:46,290 --> 00:16:48,760
In fact, you could start from what similarity and

246
00:16:48,760 --> 00:16:51,780
then compute text similarity between two sentences.

247
00:16:51,780 --> 00:16:54,090
And in general,

248
00:16:54,090 --> 00:16:58,715
the similarity functions are very useful for natural language understanding tasks.

249
00:16:58,715 --> 00:17:00,245
In the next video,

250
00:17:00,245 --> 00:17:05,000
we are going to go into more detail about topic modeling. See you there.