1
00:00:08,855 --> 00:00:13,560
In this video we are going to answer
the first of the three questions

2
00:00:13,560 --> 00:00:18,580
on what to think about in
a supervised learning scenario.

3
00:00:18,580 --> 00:00:20,870
And that is feature identification.

4
00:00:22,030 --> 00:00:23,941
How do you identify features from text?

5
00:00:26,172 --> 00:00:29,311
So why is text data so unique?

6
00:00:29,311 --> 00:00:32,726
The text data presents very
unique set of challenges

7
00:00:32,726 --> 00:00:36,695
when you're looking at in
a supervised learning scenario.

8
00:00:36,695 --> 00:00:41,090
All the information that you need or that
you have in these cases is all in text.

9
00:00:42,110 --> 00:00:44,410
But text is a very weird concept.

10
00:00:44,410 --> 00:00:46,570
There are different ways you
can parse a text document.

11
00:00:48,070 --> 00:00:52,490
The features can be pulled from
text in different granularities.

12
00:00:54,310 --> 00:00:58,510
So let's take an example of the type
of textual features that you will get.

13
00:01:00,330 --> 00:01:04,180
The basic constructs in
text is the set of words.

14
00:01:05,960 --> 00:01:10,182
So, these are by far,
the most common class of features and

15
00:01:10,182 --> 00:01:13,600
add a significant number of
features when you add to it.

16
00:01:13,600 --> 00:01:17,680
So, for example, in English language,
there are about 40,000 unique words.

17
00:01:17,680 --> 00:01:22,778
So just looking at that in common English,
you would have 40,000 features.

18
00:01:22,778 --> 00:01:27,400
If you're looking at social media or
some other genres of data,

19
00:01:27,400 --> 00:01:32,790
you might get very many more
number of features because you

20
00:01:32,790 --> 00:01:37,820
could have unique word spellings and
so on and all of those would be words.

21
00:01:39,200 --> 00:01:43,540
When you get these many features, one of
the questions you're to start answering

22
00:01:43,540 --> 00:01:46,120
is, how do you handle
commonly-occurring words?

23
00:01:47,315 --> 00:01:50,240
In some cases, they are called stop words.

24
00:01:50,240 --> 00:01:54,440
Words like the,
that occurs fairly commonly,

25
00:01:54,440 --> 00:01:58,360
and that's the most frequent
word in the English language.

26
00:01:59,450 --> 00:02:02,820
And every document, every sentence
would probably have the word the.

27
00:02:05,360 --> 00:02:10,120
But more generally the word
the is not as important for

28
00:02:10,120 --> 00:02:13,410
a classification task than any other word.

29
00:02:13,410 --> 00:02:21,240
Let's say if you're talking about politics
and you have parliament as a word.

30
00:02:21,240 --> 00:02:26,090
Parliament as a word is more
important than the word

31
00:02:26,090 --> 00:02:31,070
the to determine whether the document
belongs to the politics class.

32
00:02:33,700 --> 00:02:37,020
The next step is normalization.

33
00:02:37,020 --> 00:02:42,550
Do you make all the words lower case so
that parliament with a capital P and

34
00:02:42,550 --> 00:02:47,060
a small p are treated the same,
or add the same feature?

35
00:02:48,310 --> 00:02:51,204
Or should we leave it as-is?

36
00:02:51,204 --> 00:02:55,575
US, capitals, would be the United States.

37
00:02:55,575 --> 00:02:59,785
Whereas if you make it lowercase, it would
be indistinguishable from the word us.

38
00:03:00,870 --> 00:03:05,770
So, in some cases, you want to leave it as
is, in some cases you want it lowercase,

39
00:03:05,770 --> 00:03:06,850
so how do you make that choice?

40
00:03:08,690 --> 00:03:11,692
There are also issues about stemming and
lemmatization.

41
00:03:11,692 --> 00:03:17,720
So for example, you don't want
plurals to be different features.

42
00:03:17,720 --> 00:03:21,740
So you would want to lemmatize them or
to stand them.

43
00:03:21,740 --> 00:03:24,060
Then, all of this is still about words.

44
00:03:24,060 --> 00:03:25,370
Let's go beyond words.

45
00:03:27,050 --> 00:03:31,280
You can actually identify features or
characteristics of words.

46
00:03:31,280 --> 00:03:32,670
For example, capitalization.

47
00:03:34,420 --> 00:03:38,570
The word, as I said, US is capitalized,

48
00:03:39,770 --> 00:03:44,310
White House, with the W capitalized and
the H capitalized,

49
00:03:44,310 --> 00:03:50,260
it's very different than white house,
a white house, right?

50
00:03:50,260 --> 00:03:56,020
So, capitalization is important feature to
identify certain words and their meanings.

51
00:03:57,773 --> 00:04:01,840
You could use parts of speech
of words in a sentence.

52
00:04:01,840 --> 00:04:03,120
And those could be a feature.

53
00:04:03,120 --> 00:04:07,770
So for
example if it's important to say that this

54
00:04:07,770 --> 00:04:12,690
particular word had
a determinant in front.

55
00:04:14,500 --> 00:04:18,289
Then that word is, so
then that becomes an important feature.

56
00:04:18,289 --> 00:04:21,990
An example would be the weather,
whether example.

57
00:04:21,990 --> 00:04:26,860
Recall that we talked about a spelling
correction problem where you want to

58
00:04:26,860 --> 00:04:34,400
determine whether the correct spelling
is whether as in W-H-E-T-H-E-R or

59
00:04:34,400 --> 00:04:37,950
weather as in W-E-A-T-H-E-R.

60
00:04:37,950 --> 00:04:43,990
If you see a determiner like,
the, in front of that word,

61
00:04:43,990 --> 00:04:50,060
it most likely would be
the weather as in W-E-A-T-H-E-R.

62
00:04:50,060 --> 00:04:53,570
So that particular part of
speech before this word

63
00:04:53,570 --> 00:04:55,010
becomes a very important feature.

64
00:04:58,000 --> 00:05:01,320
You also might want to know
the grammatical structure or

65
00:05:01,320 --> 00:05:04,930
parse the sentence and
get the sentence parse structure

66
00:05:04,930 --> 00:05:08,340
to see what is the verb associated
with a particular noun.

67
00:05:08,340 --> 00:05:12,380
How far is it from the associated noun and
so on.

68
00:05:14,750 --> 00:05:18,640
And then, you may want to group
words of similar meaning to have

69
00:05:18,640 --> 00:05:20,860
one feature to represent a set of words.

70
00:05:22,140 --> 00:05:26,260
An example would be buy,
purchase, and so on.

71
00:05:26,260 --> 00:05:31,026
These are all synonyms, and you don't
want to have two different features,

72
00:05:31,026 --> 00:05:35,938
one for buy, one for purchase, you may
want to group them together because they

73
00:05:35,938 --> 00:05:38,820
mean the same,
they have the same semantics.

74
00:05:40,280 --> 00:05:44,401
But it could also be other groups,
like titles or

75
00:05:44,401 --> 00:05:49,236
honorifics, like Mr, Ms,
Dr, Professor, and so on.

76
00:05:49,236 --> 00:05:54,100
Or the set of numbers, or digits, because
you don't want to specifically have

77
00:05:54,100 --> 00:05:58,686
a feature for zero and other features for
one, and so on, all the numbers.

78
00:05:58,686 --> 00:06:02,470
So you might want to say, if it's
a number anywhere between 0 to 10,000,

79
00:06:02,470 --> 00:06:05,140
I'm just going to call it something,
a number, right, so

80
00:06:05,140 --> 00:06:08,000
suddenly you've reduced
10,000 features into one.

81
00:06:09,720 --> 00:06:10,740
Or the same with dates.

82
00:06:10,740 --> 00:06:14,624
If you are able to recognize dates
using irregular expressions for

83
00:06:14,624 --> 00:06:19,125
example and you do a very good job of it,
then you may want say, you know what,

84
00:06:19,125 --> 00:06:23,701
maybe all dates would be identified and
called one feature as a date, because I

85
00:06:23,701 --> 00:06:28,171
don't want to learn something that is for
every individual date possible.

86
00:06:28,171 --> 00:06:29,860
Then it's an infinite list.

87
00:06:31,110 --> 00:06:35,930
Other type of features would be
depending on the classification tasks.

88
00:06:35,930 --> 00:06:40,450
So for example, you may have features
that come from inside the words or

89
00:06:40,450 --> 00:06:43,130
have features with word sequences.

90
00:06:43,130 --> 00:06:46,030
An example would be bigrams or trigrams.

91
00:06:46,030 --> 00:06:47,920
The White House example comes to mind.

92
00:06:47,920 --> 00:06:50,370
Where you want to say, White House,

93
00:06:50,370 --> 00:06:55,860
as a two word construct as a bigram
is conceptually one thing.

94
00:06:56,870 --> 00:07:01,640
As compared to white and house as two
different features, two different things.

95
00:07:03,580 --> 00:07:08,710
You may also want to have character
sub-sequences such as ing or ion.

96
00:07:08,710 --> 00:07:12,438
And just by looking at
it you know that ing is

97
00:07:12,438 --> 00:07:17,130
basically saying it is a word
in its continuous form, right?

98
00:07:17,130 --> 00:07:24,649
So, just looking at ing in a word,
we'd be able to call it as a verb.

99
00:07:24,649 --> 00:07:31,410
And ion more likely at the end of
a word would be a noun of some form.

100
00:07:31,410 --> 00:07:35,800
So just these character subsequences
can help you identify some classes if

101
00:07:35,800 --> 00:07:39,020
that is important for this particular
classification task that you have.

102
00:07:41,620 --> 00:07:42,460
So how would you do it?

103
00:07:44,600 --> 00:07:47,780
We have talked about some
of these features and

104
00:07:47,780 --> 00:07:51,510
I would suggest you recall
the lectures from previous week

105
00:07:51,510 --> 00:07:55,710
that was about natural language
processing and basic NLP tasks.

106
00:07:55,710 --> 00:07:58,810
And we have addressed some
of these things there.

107
00:07:58,810 --> 00:08:02,858
You just want to identify them now and
make it available as features for

108
00:08:02,858 --> 00:08:04,570
your classification task.

109
00:08:04,570 --> 00:08:06,530
We'll see more examples of features soon.