1
00:00:07,850 --> 00:00:15,005
When we talked about the Naïve Bayes model and the theory and the formulation behind it,

2
00:00:15,005 --> 00:00:19,110
we didn't really focus on the features and what the features represented.

3
00:00:19,110 --> 00:00:25,855
There are two ways in which Naïve Bayes features could be learned.

4
00:00:25,855 --> 00:00:30,410
There are the two classic variants of Naïve Bayes for text.

5
00:00:30,410 --> 00:00:36,690
You have the multinomial Naïve Bayes model and the other one would be a Bernoulli model,

6
00:00:36,690 --> 00:00:39,265
and we will talk about it soon.

7
00:00:39,265 --> 00:00:43,500
The multinomial Naïve Bayes model is one in

8
00:00:43,500 --> 00:00:48,730
which you assume that the data follows a multinomial distribution.

9
00:00:48,730 --> 00:00:50,425
So what does that mean?

10
00:00:50,425 --> 00:00:56,040
It means that when you have the set of features that define a particular data instance,

11
00:00:56,040 --> 00:01:01,410
we're assuming that these each come independent of each other and

12
00:01:01,410 --> 00:01:08,055
can also have multiple occurrences or multiple instances of each feature.

13
00:01:08,055 --> 00:01:13,225
So, counts become important in this multinomial distribution model.

14
00:01:13,225 --> 00:01:15,785
So you have each feature value,

15
00:01:15,785 --> 00:01:19,950
a some sort of a count or a weighted count.

16
00:01:19,950 --> 00:01:25,970
Example would be word occurrence counts or TF-IDF weighting and so on.

17
00:01:25,970 --> 00:01:29,750
So, suppose you have a piece of text, a document,

18
00:01:29,750 --> 00:01:34,850
and you are finding out what are all the words that were used in this model.

19
00:01:34,850 --> 00:01:37,675
That would be called a bag-of-words model.

20
00:01:37,675 --> 00:01:40,520
And if you just use the words,

21
00:01:40,520 --> 00:01:42,080
whether they were present or not,

22
00:01:42,080 --> 00:01:46,450
then that is a Bernoulli distribution for each feature.

23
00:01:46,450 --> 00:01:52,395
So, it becomes a multivariate Bernoulli when you're talking about it for all the words.

24
00:01:52,395 --> 00:01:56,000
But if you say that the number of

25
00:01:56,000 --> 00:01:59,960
times a particular word occurs is important, so for example,

26
00:01:59,960 --> 00:02:03,665
if the statement is to be or not to be,

27
00:02:03,665 --> 00:02:07,945
and you want to somehow say that the word to occur twice,

28
00:02:07,945 --> 00:02:10,400
the word be occur twice,

29
00:02:10,400 --> 00:02:16,052
the word or occur just once and so on,

30
00:02:16,052 --> 00:02:22,025
you want to somehow keep track of what was the frequency of each of these words.

31
00:02:22,025 --> 00:02:27,410
And then, if you want to give more importance to more rare words,

32
00:02:27,410 --> 00:02:31,705
then you would add on something called a term frequency,

33
00:02:31,705 --> 00:02:34,175
inverse document frequency weighting.

34
00:02:34,175 --> 00:02:37,070
So you don't, not only give importance to the frequency,

35
00:02:37,070 --> 00:02:41,625
but say how common is this word in the entire collection,

36
00:02:41,625 --> 00:02:45,650
and that's what the idea of weighting comes from.

37
00:02:45,650 --> 00:02:48,375
So for example, the word THE is very common,

38
00:02:48,375 --> 00:02:50,840
it occurs on almost every sentence,

39
00:02:50,840 --> 00:02:52,421
it occurs in every document,

40
00:02:52,421 --> 00:02:54,405
so it is not very informative.

41
00:02:54,405 --> 00:02:57,660
But if it is the word, like,

42
00:02:57,660 --> 00:03:03,100
SIGNIFICANT, it is significant because it's not gonna be occurring in every document.

43
00:03:03,100 --> 00:03:07,380
So, you want to give a higher importance to a document that

44
00:03:07,380 --> 00:03:12,140
has this word significant as compared to the word the,

45
00:03:12,140 --> 00:03:14,390
and that kind of variation in

46
00:03:14,390 --> 00:03:19,250
weighting is possible when you're doing a multinomial Naïve Bayes model.

47
00:03:19,250 --> 00:03:22,755
The second model is the Bernoulli Naïve Bayes model.

48
00:03:22,755 --> 00:03:29,185
Here, the assumption is that the data follows a multivariate Bernoulli distribution,

49
00:03:29,185 --> 00:03:32,560
where each feature is a binary feature, that is,

50
00:03:32,560 --> 00:03:34,687
the word is present or not present,

51
00:03:34,687 --> 00:03:39,700
and it's only that information about just the word being present that is

52
00:03:39,700 --> 00:03:46,925
significant and modeled and it does not matter how many times that word was present.

53
00:03:46,925 --> 00:03:50,455
In fact, it also does not matter whether the word is

54
00:03:50,455 --> 00:03:54,900
significant or not in the sense that is the word THE,

55
00:03:54,900 --> 00:03:57,115
which is fairly common in everything,

56
00:03:57,115 --> 00:03:59,574
or is the word something like SIGNIFICANT,

57
00:03:59,574 --> 00:04:02,445
which is less common in all documents.

58
00:04:02,445 --> 00:04:07,330
So when you have just the binary features, I mean,

59
00:04:07,330 --> 00:04:09,415
just a binary model for every feature,

60
00:04:09,415 --> 00:04:11,110
then the entire data,

61
00:04:11,110 --> 00:04:14,970
the set of features follows what is called a multivariate Bernoulli model.

62
00:04:14,970 --> 00:04:20,690
So these are the two standard classic variants in Naïve Bayes,

63
00:04:20,690 --> 00:04:26,170
and you'll see that most of the approaches and most of

64
00:04:26,170 --> 00:04:32,120
the tools that you have for Naïve Bayes modeling give you that option,

65
00:04:32,120 --> 00:04:35,325
give you the option of multinomial Naïve Bayes or Bernoulli Naïve Bayes.

66
00:04:35,325 --> 00:04:40,465
It's fairly common in text documents to use the multinomial Naïve Bayes,

67
00:04:40,465 --> 00:04:44,942
but there are instances where you would want to go the Bernoulli route,

68
00:04:44,942 --> 00:04:47,920
especially if you want to somehow say that the frequency is

69
00:04:47,920 --> 00:04:49,180
immaterial and it's just

70
00:04:49,180 --> 00:04:52,780
whether the presence or absence of a word that is more important.