1
00:00:09,316 --> 00:00:10,660
Welcome back.

2
00:00:10,660 --> 00:00:13,470
In this video,
we are going to talk about topic modeling.

3
00:00:15,330 --> 00:00:20,182
Let's take an example of this
article from science, and

4
00:00:20,182 --> 00:00:25,570
this is about Seeking Life's Bare
Necessities, Bare Genetic Necessities.

5
00:00:26,810 --> 00:00:28,120
And as you look through this,

6
00:00:28,120 --> 00:00:31,190
you'll notice that some
words have been highlighted.

7
00:00:31,190 --> 00:00:33,160
So you have words such as genes and

8
00:00:33,160 --> 00:00:37,700
genomes that are highlighted in yellow,
words such as computer, and predictions,

9
00:00:37,700 --> 00:00:42,320
and computer analysis, and
computation are in blue.

10
00:00:42,320 --> 00:00:49,395
And then you have organism,
or survive, or life in pink.

11
00:00:49,395 --> 00:00:54,960
This demonstrates that any article
you see is more likely to be

12
00:00:54,960 --> 00:01:00,179
formed of different topics or
sub-units that

13
00:01:00,179 --> 00:01:04,870
intermingle very seamlessly
in weaving out an article.

14
00:01:06,400 --> 00:01:10,904
This is the basis of one of the leading
research work that has happened in

15
00:01:10,904 --> 00:01:13,232
text manning on topic modeling, and

16
00:01:13,232 --> 00:01:17,160
this one particularly is from
Latent Dirichlet Allocation.

17
00:01:18,810 --> 00:01:20,357
Well, you have three topics.

18
00:01:20,357 --> 00:01:25,438
You have genetics that's in yellow,
or computation that's in blue,

19
00:01:25,438 --> 00:01:29,686
and life-related, life science,
let's say, in pink.

20
00:01:32,206 --> 00:01:36,890
This shows that documents
are typically a mixture of topics.

21
00:01:36,890 --> 00:01:41,430
So you have topics coming from genetics,
computation, or even anatomy.

22
00:01:42,490 --> 00:01:46,866
And each of these topics
are basically words that

23
00:01:46,866 --> 00:01:50,715
are more probable coming from that topic.

24
00:01:50,715 --> 00:01:56,245
When you're talking about genes and DNA
and so on, you are mostly in the genetics

25
00:01:56,245 --> 00:02:02,282
realm, while if you're talking about brain
and neuron and nerve, you are in anatomy.

26
00:02:02,282 --> 00:02:06,265
If you're talking about computers and
numbers and data and so on,

27
00:02:06,265 --> 00:02:08,552
you're most likely in computation.

28
00:02:08,552 --> 00:02:14,313
So when a new document comes in,
in this case this article on seeking

29
00:02:14,313 --> 00:02:20,903
life's bare genetic necessities,
it comes with it of topic distribution.

30
00:02:20,903 --> 00:02:22,785
And so for that particular article,

31
00:02:22,785 --> 00:02:26,460
there is some sort of topic
distribution over these topics.

32
00:02:26,460 --> 00:02:28,950
Assume there are only
four topics in the world.

33
00:02:28,950 --> 00:02:30,763
Genetics, computation,
life sciences, anatomy.

34
00:02:30,763 --> 00:02:32,118
Obviously, that's not true.

35
00:02:32,118 --> 00:02:37,416
But let's take in this sense that these
are the only four topics you have,

36
00:02:37,416 --> 00:02:40,580
and this particular
article is generated by

37
00:02:40,580 --> 00:02:44,680
these four topics in some
combination of words.

38
00:02:44,680 --> 00:02:49,700
Where anatomy, the green one,
is absent, and computation,

39
00:02:49,700 --> 00:02:51,800
for example, is the most probable.

40
00:02:52,810 --> 00:02:55,920
But then you have genetics also
including a percentage and

41
00:02:55,920 --> 00:02:57,160
a little bit of life sciences.

42
00:02:59,660 --> 00:03:01,700
So what is a topic modeling?

43
00:03:01,700 --> 00:03:07,320
Topic modeling is a coarse-level analysis
of what is in a text collection.

44
00:03:07,320 --> 00:03:10,660
When you have a large corpus, and
you want to make sense of what

45
00:03:10,660 --> 00:03:14,620
this collection is about,
you would probably use topic modeling.

46
00:03:14,620 --> 00:03:15,910
Because you would say,

47
00:03:15,910 --> 00:03:20,130
let's figure out what kind of
documents you have in this collection.

48
00:03:20,130 --> 00:03:22,260
Are they all about sports?

49
00:03:22,260 --> 00:03:25,310
Are they all about business?

50
00:03:26,370 --> 00:03:28,180
Or are they all about computers?

51
00:03:28,180 --> 00:03:35,200
And if they are all about computers,
then are they all about architecture?

52
00:03:35,200 --> 00:03:38,800
Or are they all about algorithms?

53
00:03:38,800 --> 00:03:41,520
Which are all different
subunits within a larger unit.

54
00:03:43,940 --> 00:03:48,670
A topic is a subject of
theme of a discourse, and

55
00:03:48,670 --> 00:03:51,430
topics are represented
by a word distribution.

56
00:03:51,430 --> 00:03:54,590
And that means that you
have some probability

57
00:03:54,590 --> 00:03:57,669
of a word appearing in that topic.

58
00:03:59,560 --> 00:04:02,737
And different words have different
probabilities in that topic.

59
00:04:02,737 --> 00:04:07,224
So for example,
if you see a basketball, or

60
00:04:07,224 --> 00:04:10,496
a player, or a fee, or a score,

61
00:04:10,496 --> 00:04:15,368
you are more likely to be
in the topic of sports.

62
00:04:15,368 --> 00:04:20,095
And if you are in the topic of sports,
then words such as player and team and

63
00:04:20,095 --> 00:04:22,240
score are more likely to appear.

64
00:04:23,700 --> 00:04:29,140
Team may also appear in social science
studies but maybe not as frequently,

65
00:04:29,140 --> 00:04:33,890
or it's not as probable to have that in,
let's say, in life.

66
00:04:35,500 --> 00:04:37,730
Though it is likely there as well, right?

67
00:04:37,730 --> 00:04:41,290
So for a particular word,
you have different distribution or

68
00:04:41,290 --> 00:04:44,650
probable occurring from a topic, and

69
00:04:44,650 --> 00:04:49,570
topics are basically this probability
of distribution over all words.

70
00:04:49,570 --> 00:04:52,130
A document is assumed to
be a mixture of topics.

71
00:04:53,470 --> 00:04:55,890
So for example,
you will have a lot of topics like this.

72
00:04:55,890 --> 00:04:59,020
So you have humans,
genomes, DNA, and so on.

73
00:04:59,020 --> 00:05:01,870
That is probably about the genetics topic.

74
00:05:01,870 --> 00:05:06,490
You have another topic that is evolution,
and species, and organism, and life, and

75
00:05:06,490 --> 00:05:07,825
biology.

76
00:05:07,825 --> 00:05:11,120
You have a third topic on disease,
and host, and bacteria and so

77
00:05:11,120 --> 00:05:13,330
on, on new strains.

78
00:05:13,330 --> 00:05:16,050
And another one in computer
modeled information data.

79
00:05:16,050 --> 00:05:21,780
And you can see that these topics or
these word distributions where,

80
00:05:21,780 --> 00:05:25,030
for example,
a topic is what is there in a column, and

81
00:05:25,030 --> 00:05:30,520
they are salted maybe weekly by
how probable these words are.

82
00:05:30,520 --> 00:05:36,820
So computer or model is the most probable
word in this topic, the fourth topic.

83
00:05:36,820 --> 00:05:40,740
So when you're doing topic modeling,
what's known, what's given to you?

84
00:05:40,740 --> 00:05:44,940
What you're given is a text collection or
a corpus, and

85
00:05:44,940 --> 00:05:48,420
you are somehow given
the number of topics.

86
00:05:49,446 --> 00:05:55,690
Let's say we are interested in 20 topics,
and we are somehow group these words and

87
00:05:55,690 --> 00:05:59,110
find these topics and
find 20 of them from a large collection.

88
00:06:01,140 --> 00:06:03,370
What's not known are the actual topics.

89
00:06:04,530 --> 00:06:09,560
You are not given that you are interested
in these specific 20 topics.

90
00:06:09,560 --> 00:06:13,741
You say, well, what you've given us
that you want me to find 20 topics.

91
00:06:13,741 --> 00:06:16,381
But you could find any 20 topics, and

92
00:06:16,381 --> 00:06:19,511
you want to find a topic
that is more coherent.

93
00:06:19,511 --> 00:06:20,650
So that's part of the problem.

94
00:06:22,010 --> 00:06:25,210
And you're also not given the topic
distribution for each document, so

95
00:06:25,210 --> 00:06:29,550
you're not given that this particular
document is all about sports.

96
00:06:29,550 --> 00:06:34,145
This particular document is 50% sports and
50% genetics.

97
00:06:36,413 --> 00:06:39,191
So that distribution is not known either.

98
00:06:39,191 --> 00:06:43,115
Essentially, topic modeling
is a text clustering problem.

99
00:06:43,115 --> 00:06:46,083
However, in this particular case,
the documents and

100
00:06:46,083 --> 00:06:48,220
words are clustered simultaneously.

101
00:06:49,400 --> 00:06:51,958
You need to figure out
what words come together.

102
00:06:51,958 --> 00:06:56,800
What were they similar to each other or
semantically related to each other?

103
00:06:56,800 --> 00:07:00,920
Recall one of the previous videos where we
talked about semantic similarity of words.

104
00:07:02,220 --> 00:07:06,380
And then you also need to figure
out what documents come together.

105
00:07:06,380 --> 00:07:11,500
What documents are of the same topic or
mostly about the same topic?

106
00:07:12,500 --> 00:07:17,770
And how does these words get
derived based on these documents?

107
00:07:17,770 --> 00:07:22,500
So how do you build this topic modeling
to understand what is a distribution of

108
00:07:22,500 --> 00:07:26,990
words in a particular document and what
is a probability of a word in a topic.

109
00:07:28,720 --> 00:07:32,530
Different topic modeling approaches are
available, and there have been new models

110
00:07:32,530 --> 00:07:37,022
that are defined very regularly
in computer science literature.

111
00:07:37,022 --> 00:07:38,660
The most common ones and

112
00:07:38,660 --> 00:07:43,275
the ones that started this field are
Probabilistic Latent Semantic Analysis,

113
00:07:43,275 --> 00:07:47,040
PLSA, that was first proposed in 1999.

114
00:07:47,040 --> 00:07:53,870
And then Latent Dirichlet Allocation,
that's LDA, that was proposed in 2003.

115
00:07:53,870 --> 00:07:58,730
LDA is by far one of the most
popular topic models, and

116
00:07:58,730 --> 00:08:01,150
we're going to talk about it in
more detail in the next video.