1
00:00:00,000 --> 00:00:04,507
The Naive Bayes algorithm is one of the most
important algorithms for text

2
00:00:04,507 --> 00:00:10,735
classification. The intuition of the Naive
Bayes algorithm is really quite simple.

3
00:00:10,735 --> 00:00:15,512
It's based on Bayes Rule, which we'll see
in a second, and it relies on a very

4
00:00:15,512 --> 00:00:20,290
simple representation of the document
called the bag of words representation.

5
00:00:20,290 --> 00:00:25,191
Let's see the intuition of the bag of
words representation. Imagine I have some

6
00:00:25,191 --> 00:00:29,969
document that says, I love this movie,
it's sweet but with satirical humor. And

7
00:00:29,969 --> 00:00:36,060
so, and our job. Is to, is to build this
function gamma, which takes the document

8
00:00:36,060 --> 00:00:44,278
and returns a class. The class could be
positive. Or the class could be negative

9
00:00:44,278 --> 00:00:51,065
in case of, of sentiment analysis. Which is it
a positive or negative? In order to solve

10
00:00:51,065 --> 00:00:55,681
this task, one thing we might do is look
at individual words in the document, like

11
00:00:55,681 --> 00:01:00,410
love or satirical or great. We might look
at all of the words. In some kinds of text

12
00:01:00,410 --> 00:01:05,026
classification we're gonna look at all the
word, we're gonna look at every single

13
00:01:05,026 --> 00:01:09,584
word. In other cases we'll look at just
some subset. If we were to look at a subset,

14
00:01:09,584 --> 00:01:14,484
we might imagine that the document looks
something like this. It just looks like it

15
00:01:14,484 --> 00:01:19,271
has the word love, and the word satirical,
and the word great and all the other words

16
00:01:19,271 --> 00:01:24,213
have disappeared. Whether we use a subset
of words or all of the words in the

17
00:01:24,213 --> 00:01:28,463
document. The bag of words
representation loses all the information

18
00:01:28,463 --> 00:01:32,888
about the order of the words in the
document and all we represent about the

19
00:01:32,888 --> 00:01:38,538
document is the set of words that occurred
and their accounts. So for example for the

20
00:01:38,538 --> 00:01:43,379
previous document, we might represent the
document as just a vector of words: great,

21
00:01:43,379 --> 00:01:47,690
love, recommend, laugh, happy. And for
each one account: great occurred twice,

22
00:01:47,690 --> 00:01:52,236
love occurred twice, recommend occurred
once. And again, we can keep all of the

23
00:01:52,236 --> 00:01:57,196
words in the document, and we'll often do
that or we can just keep some of the words

24
00:01:57,196 --> 00:02:01,624
in the document if we have an idea that
some of the words are particularly

25
00:02:01,624 --> 00:02:06,584
indicative cues. So the idea of the bag of
words' representation is that we're gonna

26
00:02:06,584 --> 00:02:11,344
represent our document just by a list of
words. And there counts, and throw away

27
00:02:11,344 --> 00:02:16,244
everything else about the document. Which
order the words occurred in, what font

28
00:02:16,244 --> 00:02:20,710
they were in, anything else, and our
function will, our function gamma, our

29
00:02:20,710 --> 00:02:25,361
classifier, will take that representation,
and assign us a class positive or

30
00:02:25,361 --> 00:02:31,762
negative. And this applies, I've shown it
to you for the two class problem of

31
00:02:31,762 --> 00:02:36,527
sentiment analysis, positive or negative
sentiment. But this applies for all sorts

32
00:02:36,527 --> 00:02:40,999
of document classification tasks. So I
might have some document I need to

33
00:02:40,999 --> 00:02:45,706
classify into a different computer science
topic because I'm building an online

34
00:02:45,706 --> 00:02:50,590
library of computer science papers. Or I'm
giving advice on computer science topics.

35
00:02:50,590 --> 00:02:55,568
So I have some text, some, some document
here with words like parser or language or

36
00:02:55,568 --> 00:03:00,486
label or translation and I wanna know
which aspect of computer science it should

37
00:03:00,486 --> 00:03:05,161
go in so I can file my paper automatically
and a good text classifier should

38
00:03:05,161 --> 00:03:09,775
automatically figure out that that's a,
that's a natural language processing

39
00:03:09,775 --> 00:03:18,560
paper. So that's the intuition of the
naive base classifier.
