1
00:00:08,998 --> 00:00:12,780
Let's talk in more detail about
generative models and LDA.

2
00:00:14,020 --> 00:00:19,620
The generative models for text
basically starts with this magic chest.

3
00:00:21,020 --> 00:00:24,740
Suppose you have this chest and
words come out of this chest magically.

4
00:00:26,320 --> 00:00:30,623
And you pick words from this chest
to then create your document.

5
00:00:33,165 --> 00:00:37,710
So when you start pulling out words,
you start seeing words like this.

6
00:00:37,710 --> 00:00:44,130
Harry, Potter, Is, and then you have other
words like movie and the, and so on.

7
00:00:45,490 --> 00:00:48,630
Already, by just looking
at the first two words,

8
00:00:48,630 --> 00:00:53,360
you know that this chest kind of gives
out words about harry potter, so

9
00:00:53,360 --> 00:00:57,770
this is some distribution that favors
words coming from harry potter.

10
00:00:59,020 --> 00:01:03,959
And then you can these words that come
out to create this document to generate

11
00:01:03,959 --> 00:01:05,049
this document.

12
00:01:05,049 --> 00:01:09,791
And this document will be something like
the movie harry potter is based on books

13
00:01:09,791 --> 00:01:11,020
from j.k rowling.

14
00:01:11,020 --> 00:01:16,980
And now you see that in the generation
process, you have a model that gives out

15
00:01:16,980 --> 00:01:22,350
words, and then you use those words coming
from that model to generate the document.

16
00:01:24,580 --> 00:01:26,760
But then you could go the other way.

17
00:01:26,760 --> 00:01:29,400
You could start from the document and

18
00:01:29,400 --> 00:01:35,840
see how many times the word the occurs,
or harry occurs, or potter occurs.

19
00:01:35,840 --> 00:01:40,591
And then create a distribution of words,
that is you create

20
00:01:40,591 --> 00:01:45,723
a probability distribution of how
likely it is to see the words,

21
00:01:45,723 --> 00:01:50,968
harry, in this document or
the word, movie, in this document.

22
00:01:50,968 --> 00:01:55,825
And you'll notice that when
you generate this model,

23
00:01:55,825 --> 00:02:01,429
when you infer this model,
the word, the, is most frequent.

24
00:02:01,429 --> 00:02:03,140
The probability is 0.1.

25
00:02:03,140 --> 00:02:06,120
That means one in ten words is the.

26
00:02:07,280 --> 00:02:09,890
And then you have is, and
then harry, and potter, and so on.

27
00:02:10,930 --> 00:02:15,210
So notice that because the documents
were about harry potter,

28
00:02:15,210 --> 00:02:20,140
the model favors the words
harry potter there.

29
00:02:20,140 --> 00:02:22,630
It's very unlikely that
you would see harry and

30
00:02:22,630 --> 00:02:26,600
potter being this frequent
in any other topic model or

31
00:02:26,600 --> 00:02:31,123
in any other corpus of documents.

32
00:02:31,123 --> 00:02:34,580
So here you had a very
simple generative process.

33
00:02:34,580 --> 00:02:37,210
You had one topic model and

34
00:02:37,210 --> 00:02:40,990
you pulled out words from that topic
model to create your document.

35
00:02:40,990 --> 00:02:42,130
That was a generation story.

36
00:02:43,850 --> 00:02:47,570
However the generation story can
be very complex in most cases.

37
00:02:48,680 --> 00:02:52,474
Instead, suppose you have,
instead of one topic,

38
00:02:52,474 --> 00:02:56,440
you have four topic models, four chests.

39
00:02:56,440 --> 00:03:01,760
And you have this magic hat that pulls
out words from these chests at random or

40
00:03:01,760 --> 00:03:05,980
it has its own policy of choosing
one chest over the other.

41
00:03:05,980 --> 00:03:09,360
And then you have these words that come
and then you still create this document.

42
00:03:10,610 --> 00:03:13,880
Now your model is more complex
because instead of learning,

43
00:03:15,550 --> 00:03:20,700
the generation is almost like where you
decide which chest the word comes out of.

44
00:03:20,700 --> 00:03:22,010
And once you have made that choice,

45
00:03:22,010 --> 00:03:26,030
then you have a different distribution
of that word coming from that chest.

46
00:03:27,450 --> 00:03:31,743
But you still create the same document,
but then, when you are using these

47
00:03:31,743 --> 00:03:35,443
documents to infer your models,
you need to infer four models.

48
00:03:35,443 --> 00:03:40,284
And you need to somehow infer
what was the combination of

49
00:03:40,284 --> 00:03:45,331
words coming from these four chests,
these four topics.

50
00:03:45,331 --> 00:03:51,336
So you not only have to somehow figure out
what were the individual topic models,

51
00:03:51,336 --> 00:03:53,920
individual word distributions.

52
00:03:53,920 --> 00:03:57,720
But also, this mixture model of how you

53
00:03:57,720 --> 00:04:02,070
use these four topic models and
combine them to create one document.

54
00:04:03,450 --> 00:04:06,080
So this is typically
called the mixture model,

55
00:04:06,080 --> 00:04:10,830
the first one that we saw in
the previous slide was unique model,

56
00:04:10,830 --> 00:04:14,910
where you have one topic distribution and
you get words from there.

57
00:04:14,910 --> 00:04:16,840
Whereas here, it's a mixture of topics.

58
00:04:16,840 --> 00:04:21,600
So you have the same document,
generated by four different topics.

59
00:04:21,600 --> 00:04:25,600
Some of them represented with a higher
proportion and others that are not.

60
00:04:25,600 --> 00:04:31,138
It should remind you of the example we
started this topic model discussion from.

61
00:04:31,138 --> 00:04:35,911
That was on the bare necessities in
science article, and you saw that there

62
00:04:35,911 --> 00:04:40,820
was a topic model for computation and
another topic model for genetics.

63
00:04:40,820 --> 00:04:42,412
And a third topic model for

64
00:04:42,412 --> 00:04:46,480
anatomy that was not represented
as well in the document and so on.

65
00:04:47,690 --> 00:04:51,518
So this is kind of similar model here.

66
00:04:51,518 --> 00:04:56,138
LDA is another such generative model.

67
00:04:56,138 --> 00:05:01,039
And the generative model for a document
d is you choose the length of document,

68
00:05:01,039 --> 00:05:05,806
you first decide what is the length of
the document that you are generating.

69
00:05:05,806 --> 00:05:09,870
Then you choose a mixture of topics for
that document.

70
00:05:11,180 --> 00:05:15,405
And then you use that topic's multinomial
distribution, that is the word

71
00:05:15,405 --> 00:05:20,620
distribution, to output the words to
fill up that quota, that topic's quota.

72
00:05:20,620 --> 00:05:22,380
Suppose you decide that for

73
00:05:22,380 --> 00:05:27,580
a particular document, 40% of the words
come from topic A, then you use that

74
00:05:27,580 --> 00:05:31,929
topic A's multinomial distribution
to output the 40% of the words.

75
00:05:33,690 --> 00:05:38,619
This is a very simplistic
explanation of LDA.

76
00:05:38,619 --> 00:05:44,040
You could, some of you might have
seen more complex mathematical

77
00:05:44,040 --> 00:05:48,900
notations, called plate notations,
to define topic models.

78
00:05:48,900 --> 00:05:51,930
And that is something that we'll leave for
a future study.

79
00:05:51,930 --> 00:05:54,490
I can point to some of
it in the reading list.

80
00:05:54,490 --> 00:05:58,260
But for now, it is enough to
kind of understand that LDA

81
00:05:58,260 --> 00:06:02,790
is also a generative model and
it creates its documents based on

82
00:06:03,960 --> 00:06:08,680
some notion of length of the document,
mixture of topics in that document and

83
00:06:08,680 --> 00:06:11,420
then, individual topics
multinomial distributions.

84
00:06:13,580 --> 00:06:16,250
In practice the questions you need to ask

85
00:06:16,250 --> 00:06:20,950
though when you create a model such
as LDA is how many topics you want.

86
00:06:22,190 --> 00:06:24,570
There is no good answer for it.

87
00:06:24,570 --> 00:06:27,780
Finding or even guessing that
number is actually very hard.

88
00:06:27,780 --> 00:06:32,460
So you have to somehow say that okay I
believe that there might be five topics.

89
00:06:32,460 --> 00:06:35,990
Or I would prefer learning
five distinct topics

90
00:06:35,990 --> 00:06:39,580
than 25 topics that are very
similar to each other.

91
00:06:39,580 --> 00:06:41,050
So you make a choice,

92
00:06:41,050 --> 00:06:45,740
just based on a guess of how
distinct these topics could be.

93
00:06:45,740 --> 00:06:49,990
But if you are in a domain where you
know these topics a little bit well.

94
00:06:49,990 --> 00:06:52,440
So for example,
you have all medical documents.

95
00:06:52,440 --> 00:06:56,960
And you know that these medical
documents come from radiology and

96
00:06:56,960 --> 00:07:01,450
pathology and urology, and there are these
streams, then you might say, okay,

97
00:07:01,450 --> 00:07:06,198
I'm interested in these seven streams
of medicine, and those are my topics.

98
00:07:06,198 --> 00:07:10,530
So that there at least you have some
sense of how many topics there should be.

99
00:07:11,920 --> 00:07:15,170
The other big problem is
interpreting the topics.

100
00:07:15,170 --> 00:07:18,400
So you would get topics, but
topics are just word distributions.

101
00:07:18,400 --> 00:07:21,320
They just tell you which
words are more frequent or

102
00:07:21,320 --> 00:07:26,620
more probable coming from particular
topic and which ones are not as probable.

103
00:07:28,070 --> 00:07:32,900
But making sense of that or
generating a coherent label for

104
00:07:32,900 --> 00:07:35,100
the topic is a subjective decision.

105
00:07:35,100 --> 00:07:39,780
There have been some work that have looked
into generating names for these topics.

106
00:07:39,780 --> 00:07:46,380
But most likely whenever you see a name in
a topic model, it just comes out manually.

107
00:07:46,380 --> 00:07:51,150
When people just look at these
words like genetics and genes and

108
00:07:51,150 --> 00:07:54,481
so on, and
say that this is a genetic topic.

109
00:07:54,481 --> 00:07:59,044
Or if they say computation, and
model, and data and information and

110
00:07:59,044 --> 00:08:04,408
there's something to do with computation
or computer science or informatics.

111
00:08:04,408 --> 00:08:08,482
So those names are fairly subjective.

112
00:08:08,482 --> 00:08:12,680
But actual topics that you
learn from LDA is basically

113
00:08:12,680 --> 00:08:15,949
a solution of an optimization function.

114
00:08:15,949 --> 00:08:20,398
So it is more kind of
deterministic in that sense.

115
00:08:20,398 --> 00:08:24,930
To summarize, topic modeling is a great
tool for exploratory text analysis

116
00:08:24,930 --> 00:08:28,755
that helps you kind of answer the question
about what these documents are about.

117
00:08:28,755 --> 00:08:30,310
What is this corpus about?

118
00:08:30,310 --> 00:08:32,910
And you could think about
it as a corpus of tweets,

119
00:08:32,910 --> 00:08:35,410
a corpus of reviews, of news articles.

120
00:08:35,410 --> 00:08:37,690
So you might get a big dump of tweets and

121
00:08:37,690 --> 00:08:40,500
say, what are people
talking about in tweets?

122
00:08:40,500 --> 00:08:44,170
What are the different themes
that come from tweets?

123
00:08:45,520 --> 00:08:50,500
And there are many tools available to
do it fairly effortlessly in Python.

124
00:08:50,500 --> 00:08:52,544
So let's take an example of how to do it.

125
00:08:52,544 --> 00:08:55,290
There are many packages available.

126
00:08:55,290 --> 00:08:58,070
Some of them are gensim, lda and so on.

127
00:08:58,070 --> 00:09:04,080
We are going to go and talk about
gensim more in the next few slides.

128
00:09:04,080 --> 00:09:09,027
But before you use any of these packages,
you need to pre-process text.

129
00:09:09,027 --> 00:09:13,490
And I would encourage you to kind of
recall what we talked about early on

130
00:09:13,490 --> 00:09:18,122
in this course, right in the first
module about pre-processing text.

131
00:09:18,122 --> 00:09:20,428
That you need to tokenize text and
normalize it,

132
00:09:20,428 --> 00:09:22,430
that means make them all lowercase.

133
00:09:22,430 --> 00:09:26,630
Decide whether you should make them
lowercase or not, you remove stop words.

134
00:09:26,630 --> 00:09:31,380
Stop words are common words
that occur frequently

135
00:09:31,380 --> 00:09:35,860
in a particular domain and
is not meaningful in that domain.

136
00:09:35,860 --> 00:09:40,640
So for example, in general English,
the word the and is and so

137
00:09:40,640 --> 00:09:42,980
on might be words that you want to remove.

138
00:09:44,180 --> 00:09:52,520
While if you are in the area of medical
documents, let's say, so clinical notes,

139
00:09:52,520 --> 00:09:56,970
you would always see the word patient and
always see the word doctor and so on.

140
00:09:56,970 --> 00:10:00,150
And they may not be as
important as the other words,

141
00:10:00,150 --> 00:10:04,460
like what is the medication and
what is the disease.

142
00:10:04,460 --> 00:10:09,319
Then you may want to say patient and
doctor are stop words for that context.

143
00:10:09,319 --> 00:10:12,220
The other pre-processing
step would be stemming.

144
00:10:12,220 --> 00:10:15,734
That means you would need to remove
the derivation in related forms,

145
00:10:15,734 --> 00:10:19,780
somehow normalize the derivation
in related forms to the same word.

146
00:10:19,780 --> 00:10:24,052
Meet, meeting, met,
all should be called meet, let's say.

147
00:10:24,052 --> 00:10:27,496
And then once you have done
the pre-processing steps,

148
00:10:27,496 --> 00:10:32,260
you convert this tokenized document
into a document term matrix.

149
00:10:32,260 --> 00:10:37,110
So going from which document
has what words to what words

150
00:10:37,110 --> 00:10:39,110
are occurring in which documents.

151
00:10:39,110 --> 00:10:43,988
Getting that document term matrix
would be the important first

152
00:10:43,988 --> 00:10:47,587
step in finding out, and
in working with LDA.

153
00:10:47,587 --> 00:10:51,851
And then once you have done that, once
you have build this document term matrix,

154
00:10:51,851 --> 00:10:53,800
you build the LDA models on top of it.

155
00:10:55,330 --> 00:11:00,960
So once you have built the mapping
between the terms and documents, then

156
00:11:00,960 --> 00:11:06,060
suppose you have a set of pre-processed
text documents in this variable doc_set.

157
00:11:06,060 --> 00:11:09,690
Then you could use gensim
to learn LDA this way.

158
00:11:09,690 --> 00:11:14,460
You could import gensim and specifically
you import the corpora and the models.

159
00:11:14,460 --> 00:11:20,170
First you create a dictionary, dictionary
is mapping between IDs and words.

160
00:11:21,350 --> 00:11:25,410
Then you create corpus, and
corpus you create going through this,

161
00:11:25,410 --> 00:11:29,550
all the documents in the doc_set, and
creating a document to bag of words model.

162
00:11:29,550 --> 00:11:34,904
This is the step that creates
the document term matrix.

163
00:11:34,904 --> 00:11:39,494
Once you have that, then you input
that in the LdaModel call, so

164
00:11:39,494 --> 00:11:42,470
that you use a gensim.models LdaModel,

165
00:11:42,470 --> 00:11:46,906
where you also specify the number
of topics you want to learn.

166
00:11:46,906 --> 00:11:50,396
So in this case, we said number of
topics is going to be four, and

167
00:11:50,396 --> 00:11:53,490
you also specify this mapping,
the id2word mapping.

168
00:11:53,490 --> 00:11:56,210
That's a dictionary that
is learned two steps ahead.

169
00:11:57,570 --> 00:11:59,000
Once you have learned that,
then that's it, and

170
00:11:59,000 --> 00:12:02,280
you can say how many passes
it should go through.

171
00:12:02,280 --> 00:12:06,990
And there are other parameters that
I would encourage you to read upon.

172
00:12:06,990 --> 00:12:09,257
But once you have defined this ldamodel,

173
00:12:09,257 --> 00:12:12,520
you can then use the ldamodel
to print the topics.

174
00:12:12,520 --> 00:12:15,100
So, in this particular case,
we learnt four topics.

175
00:12:15,100 --> 00:12:18,140
And you can say, give me the top
five words of these four topics and

176
00:12:18,140 --> 00:12:19,740
then it will bring that one out for you.

177
00:12:21,450 --> 00:12:24,930
ldamodel model can also be used to
find topic distributions of documents.

178
00:12:24,930 --> 00:12:30,520
So when you have a new document and you
apply the ldamodel on it, so you infer it.

179
00:12:30,520 --> 00:12:32,090
You can say,

180
00:12:32,090 --> 00:12:36,030
what was the topic distribution, across
these four topics, for that new document.

181
00:12:37,970 --> 00:12:42,089
So the take home concepts here are,
the topic modeling is an exploratory tool,

182
00:12:42,089 --> 00:12:44,191
that is frequently used in text mining.

183
00:12:44,191 --> 00:12:48,564
LDA or Linear Dirichlet Allocation
is a generative model,

184
00:12:48,564 --> 00:12:53,550
that is used extensively,
in modeling large text corpora.

185
00:12:53,550 --> 00:12:58,060
There are other topic models available,
PLSA being another one.

186
00:12:58,060 --> 00:13:02,060
In addition to being an exploratory tool,
LDA can also be

187
00:13:02,060 --> 00:13:06,590
used as a feature selection technique for
text classification and other tasks.

188
00:13:06,590 --> 00:13:11,320
So for example,
if you want to remove all features

189
00:13:11,320 --> 00:13:15,520
that are coming from words that are very
fairly common in your corpus or you want

190
00:13:15,520 --> 00:13:20,720
to focus your features to only those
that are coming from specific topics.

191
00:13:20,720 --> 00:13:23,750
Then you would want to
first train an LDA model.

192
00:13:23,750 --> 00:13:28,720
And then, based on just the words that

193
00:13:28,720 --> 00:13:32,250
come from specific topics of interest, you
might actually generate features that way.

194
00:13:33,340 --> 00:13:38,760
So in general, LDA is a very powerful
tool and a text clustering tool

195
00:13:38,760 --> 00:13:43,490
that is fairly commonly used as the first
step to understand what a corpus is about.

196
00:13:44,620 --> 00:13:47,980
Hope you learned how we could
use topic modeling, and

197
00:13:47,980 --> 00:13:50,940
this gives you a brief
introduction to topic models.

198
00:13:50,940 --> 00:13:55,020
There is many, many things you can
go into more detail about, but

199
00:13:55,020 --> 00:13:57,850
for this course,
I think we'll leave it there.