1
00:00:09,137 --> 00:00:14,040
Today, we'll be going through an example
of using scikit-learn to perform sentiment

2
00:00:14,040 --> 00:00:15,710
analysis on Amazon Reviews.

3
00:00:17,190 --> 00:00:21,515
The data set we'll be working with today
is the Amazon Reviews on Unlocked_Mobile

4
00:00:21,515 --> 00:00:22,450
phones dataset.

5
00:00:24,135 --> 00:00:28,926
Looking at the head of the dataframe, we
can see we have the Product Name, Brand,

6
00:00:28,926 --> 00:00:34,570
Price, Rating, Review text and the number
of people who found the review helpful.

7
00:00:36,400 --> 00:00:40,039
For our purposes, we'll be focusing
on the Rating and Reviews columns.

8
00:00:41,870 --> 00:00:44,660
Let's start by cleaning
up the dataframe a bit.

9
00:00:44,660 --> 00:00:47,040
First, we'll drop any
rows with missing values.

10
00:00:48,410 --> 00:00:52,500
Next, let's remove any ratings = 3,
we'll assume these are neutral.

11
00:00:54,310 --> 00:00:58,245
Finally, we'll create a new column that
will serve as our target for our model,

12
00:00:58,245 --> 00:01:03,685
where any reviews that were rated
more than 3 will be encoded as a 1,

13
00:01:03,685 --> 00:01:05,422
indicating it was positively rated.

14
00:01:05,422 --> 00:01:11,000
Otherwise, it'll be encoded as a 0,
indicating it was not positively rated.

15
00:01:13,460 --> 00:01:16,260
Looking at the mean of
the positively rated column,

16
00:01:16,260 --> 00:01:18,459
we can see that we have
imbalanced classes.

17
00:01:20,460 --> 00:01:24,350
Now, let's put our data into training and
test sets using the reviews and

18
00:01:24,350 --> 00:01:25,690
positively rated columns.

19
00:01:28,579 --> 00:01:30,300
Looking at X_train,

20
00:01:30,300 --> 00:01:35,836
we can see we have a series of
over 231,000 reviews or documents.

21
00:01:35,836 --> 00:01:42,690
We'll need to convert these into a numeric
representation that scikit-learn can use.

22
00:01:42,690 --> 00:01:47,438
The bag-of-words approach is simple and
commonly used way to represent text for

23
00:01:47,438 --> 00:01:50,603
use in machine learning,
which ignores structure and

24
00:01:50,603 --> 00:01:53,660
only counts how often each word occurs.

25
00:01:53,660 --> 00:01:57,140
CountVectorizer allows us to
use the bag-of-words approach

26
00:01:57,140 --> 00:02:01,399
by converting a collection of text
documents into a matrix of token counts.

27
00:02:02,810 --> 00:02:07,430
First, we instantiate the CountVectorizer
and fit it to our training data.

28
00:02:09,270 --> 00:02:14,400
Fitting the CountVectorizer consists of
the tokenization of the trained data and

29
00:02:14,400 --> 00:02:15,740
building of the vocabulary.

30
00:02:18,420 --> 00:02:23,330
Fitting the CountVectorizer tokenizes
each document by finding all sequences of

31
00:02:23,330 --> 00:02:28,910
characters of at least two letters or
numbers separated by word boundaries.

32
00:02:28,910 --> 00:02:33,680
Converts everything to lowercase and
builds a vocabulary using these tokens.

33
00:02:36,100 --> 00:02:39,497
We can get the vocabulary by using
the get_feature_names method.

34
00:02:40,710 --> 00:02:44,860
This vocabulary is built on any tokens
that occurred in the training data.

35
00:02:47,010 --> 00:02:48,460
Looking at every 2,000th feature,

36
00:02:48,460 --> 00:02:53,200
we can get a small sense of
what the vocabulary looks like.

37
00:02:53,200 --> 00:02:55,110
We can see it looks pretty messy,

38
00:02:55,110 --> 00:02:57,730
including words with numbers
as well as misspellings.

39
00:02:59,879 --> 00:03:02,444
By checking the length
of get_feature_names,

40
00:03:02,444 --> 00:03:05,809
we can see that we're working
with over 53,000 features.

41
00:03:09,028 --> 00:03:13,810
Next, we use the transform method to
transform the documents in X_train to

42
00:03:13,810 --> 00:03:19,132
a document term matrix, giving us
the bag-of-word representation of X_train.

43
00:03:21,390 --> 00:03:26,360
This representation is stored in a SciPy
sparse matrix, where each row corresponds

44
00:03:26,360 --> 00:03:30,280
to a document and each column a word
from our training vocabulary.

45
00:03:31,780 --> 00:03:34,230
The entries in this matrix

46
00:03:34,230 --> 00:03:36,910
are the number of times each
word appears in each document.

47
00:03:38,260 --> 00:03:40,850
Because the number of words
in the vocabulary is so

48
00:03:40,850 --> 00:03:44,580
much larger than the number of words
that might appear in a single review,

49
00:03:44,580 --> 00:03:46,280
most entries of this matrix are zero.

50
00:03:48,789 --> 00:03:53,928
Now let's use this feature matrix X_
train_ vectorized to train our model.

51
00:03:53,928 --> 00:03:58,450
We'll use LogisticRegression, which works
well for high dimensional sparse data.

52
00:04:01,300 --> 00:04:03,971
Next, we'll make predictions
using X_test and

53
00:04:03,971 --> 00:04:06,148
compute the area under the curve score.

54
00:04:06,148 --> 00:04:10,420
We'll transform X_test
using our vectorizer

55
00:04:10,420 --> 00:04:11,720
that was fitted to the training data.

56
00:04:12,900 --> 00:04:18,589
Note that any words in X_test that didn't
appear in X_train will just be ignored.

57
00:04:18,589 --> 00:04:25,835
Looking at our AUC score,
we see we achieve a score of about 0.927.

58
00:04:25,835 --> 00:04:28,070
Let's take a look at
the coefficients from our model.

59
00:04:29,260 --> 00:04:31,470
Sorting them and
looking at the ten smallest and

60
00:04:31,470 --> 00:04:36,920
ten largest coefficients, we can see
the model has connected words like worst,

61
00:04:36,920 --> 00:04:40,430
worthless and junk to negative reviews.

62
00:04:40,430 --> 00:04:44,348
And words like excellent, loves,
and amazing to positive reviews.

63
00:04:46,861 --> 00:04:49,376
Next, let's look at a different approach,

64
00:04:49,376 --> 00:04:52,365
which allows us to rescale
features called tf–idf.

65
00:04:54,060 --> 00:04:57,700
Tf–idf, or
Term frequency-inverse document frequency,

66
00:04:57,700 --> 00:05:01,900
allows us to weight terms based on
how important they are to a document.

67
00:05:03,520 --> 00:05:07,600
High weight is given to terms that appear
often in a particular document, but

68
00:05:07,600 --> 00:05:10,650
don't appear often in the corpus.

69
00:05:10,650 --> 00:05:15,830
Features with low tf–idf are either
commonly used across all documents or

70
00:05:15,830 --> 00:05:18,440
rarely used and
only occur in long documents.

71
00:05:19,440 --> 00:05:24,360
Features with high tf–idf are frequently
used within specific documents, but

72
00:05:24,360 --> 00:05:26,119
rarely used across all documents.

73
00:05:28,480 --> 00:05:31,100
Similar to how we used CountVectorizer,

74
00:05:31,100 --> 00:05:35,380
we'll instantiate the tf–idf vectorizer
and fit it to our training data.

75
00:05:37,350 --> 00:05:41,810
Because tf–idf vectorizer goes through
the same initial process of tokenizing

76
00:05:41,810 --> 00:05:45,579
the document, we can expect it to
return the same number of features.

77
00:05:46,800 --> 00:05:50,670
However, let's take a look at a few tricks
for reducing the number of features

78
00:05:50,670 --> 00:05:53,720
that might help improve our model's
performance or reduce a refitting.

79
00:05:56,126 --> 00:06:00,690
CountVectorizor and
tf–idf Vectorizor both take an argument,

80
00:06:00,690 --> 00:06:05,530
mindf, which allows us to specify
a minimum number of documents

81
00:06:05,530 --> 00:06:08,890
in which a token needs to appear
to become part of the vocabulary.

82
00:06:10,430 --> 00:06:14,844
This helps us remove some words
that might appear in only a few and

83
00:06:14,844 --> 00:06:17,550
are unlikely to be useful predictors.

84
00:06:17,550 --> 00:06:22,526
For example, here we'll pass in min_df
= 5, which will remove any words

85
00:06:22,526 --> 00:06:26,661
from our vocabulary that appear
in fewer than five documents.

86
00:06:29,101 --> 00:06:33,754
Looking at the length,
we can see we've reduced the number of

87
00:06:33,754 --> 00:06:39,047
features by over 35,000 to
just under 18,000 features.

88
00:06:39,047 --> 00:06:43,718
Next, when we transform our training data,
fit our model,

89
00:06:43,718 --> 00:06:49,397
make predictions on the transform
test data, and compute the AUC score,

90
00:06:49,397 --> 00:06:53,628
we can see we, again,
get an AUC of about 0.927.

91
00:06:53,628 --> 00:06:56,392
No improvement in AUC score, but

92
00:06:56,392 --> 00:07:01,340
we were able to get the same
score using far fewer features.

93
00:07:01,340 --> 00:07:04,985
Let's take a look at which features
have the smallest and largest tf–idf.

94
00:07:06,920 --> 00:07:10,490
List of features with the smallest
tf–idf either commonly

95
00:07:10,490 --> 00:07:14,540
appeared across all reviews or
only appeared rarely in very long reviews.

96
00:07:16,430 --> 00:07:20,050
List of features with the largest
tf–idf contains words which appeared

97
00:07:20,050 --> 00:07:24,169
frequently in a review, but did not
appear commonly across all reviews.

98
00:07:27,205 --> 00:07:30,733
Looking at the smallest and
largest coefficients from our new model,

99
00:07:30,733 --> 00:07:34,080
we can again see which words our
model has connected to negative and

100
00:07:34,080 --> 00:07:35,120
positive reviews.

101
00:07:37,880 --> 00:07:43,230
One problem with our previous bag-of-words
approach is word order is disregarded.

102
00:07:43,230 --> 00:07:48,350
So, not an issue, phone is working
is seen the same as an issue,

103
00:07:48,350 --> 00:07:49,260
phone is not working.

104
00:07:50,280 --> 00:07:53,660
Our current model sees both of
these reviews as negative reviews.

105
00:07:56,020 --> 00:08:00,460
One way we can add some context is
by adding sequences of word features

106
00:08:00,460 --> 00:08:02,530
known as n-grams.

107
00:08:02,530 --> 00:08:07,046
For example, bigrams,
which count pairs of adjacent words,

108
00:08:07,046 --> 00:08:11,492
could give us features such as
is working versus not working.

109
00:08:11,492 --> 00:08:14,663
And trigrams,
which give us triplets of adjacent words,

110
00:08:14,663 --> 00:08:17,070
could give us features
such as not an issue.

111
00:08:19,480 --> 00:08:21,610
To create these n-gram features,

112
00:08:21,610 --> 00:08:25,110
we'll pass in a tuple to
the parameter ngram_range,

113
00:08:25,110 --> 00:08:29,499
where the values correspond to the minimum
length and maximum lengths of sequences.

114
00:08:30,760 --> 00:08:35,820
For example, if I pass in the tuple,
1, 2, CountVectorizer

115
00:08:35,820 --> 00:08:39,670
will create features using the individual
words, as well as the bigrams.

116
00:08:41,890 --> 00:08:45,960
Let's see what kind of AUC score we can
achieve by adding bigrams to our model.

117
00:08:47,270 --> 00:08:51,530
Keep in mind that, although n-grams
can be powerful in capturing meaning,

118
00:08:51,530 --> 00:08:54,840
longer sequences can cause an explosion
of the number of features.

119
00:08:55,900 --> 00:08:57,460
Just by adding bigrams,

120
00:08:57,460 --> 00:09:02,190
the number of features we have
has increased to almost 200,000.

121
00:09:02,190 --> 00:09:07,030
And after training our logistic
regression model and our new features,

122
00:09:07,030 --> 00:09:14,960
looks like by adding bigrams, we were
able to improve our AUC score to 0.967.

123
00:09:14,960 --> 00:09:18,590
If we take a look at what features our
model connected with negative reviews,

124
00:09:18,590 --> 00:09:23,110
we can see that we now have bigrams
such as no good and not happy,

125
00:09:24,190 --> 00:09:28,130
while for positive reviews we
have not bad and no problems.

126
00:09:30,130 --> 00:09:34,830
If we again try to predict not an issue,
phone is working, and an issue,

127
00:09:34,830 --> 00:09:39,810
phone is not working, we can see that our
newest model now correctly identifies them

128
00:09:39,810 --> 00:09:42,050
as positive and
negative reviews respectively.

129
00:09:44,370 --> 00:09:47,710
The vectorizers we saw in this
tutorial are very flexible and

130
00:09:47,710 --> 00:09:51,880
also support tasks such as removing
stop words or limitization.

131
00:09:51,880 --> 00:09:54,920
So be sure to check the documentation for
more info.

132
00:09:54,920 --> 00:09:56,270
As always, thanks for watching.