1
00:00:00,480 --> 00:00:02,110
To a different module.

2
00:00:02,110 --> 00:00:05,900
You might have seen my Big Data
Architecture Fundamental module here

3
00:00:05,900 --> 00:00:07,120
earlier in the class.

4
00:00:07,120 --> 00:00:09,960
So welcome to a completely different
module, three part series here.

5
00:00:09,960 --> 00:00:14,180
We're going to talk about Content
Detection and Analysis for Big Data okay.

6
00:00:15,280 --> 00:00:18,580
So we're going to cover
a number of different topics.

7
00:00:18,580 --> 00:00:21,310
We're going to talk about sort of the
landscape of all of the different content

8
00:00:21,310 --> 00:00:23,810
types that are out there when
you're building big data systems.

9
00:00:23,810 --> 00:00:27,350
Like search engines or
just data analysis systems, what types of

10
00:00:27,350 --> 00:00:31,550
file formats are out there, what their
meaning is, what their importance is.

11
00:00:31,550 --> 00:00:35,030
Why it's important to be able to
automatically and rapidly detect them.

12
00:00:35,030 --> 00:00:38,590
We'll talk about some of the challenges
in doing so, in detecting them,

13
00:00:38,590 --> 00:00:42,670
in extracting text and information from
them, and extracting their language.

14
00:00:42,670 --> 00:00:46,860
Different approaches to doing that and
sort of the final the part of this sort of

15
00:00:46,860 --> 00:00:50,680
three part model here will cover
a specific technology called Apache Tika.

16
00:00:50,680 --> 00:00:53,090
Which is really good to
have in your tool belt for

17
00:00:53,090 --> 00:00:55,310
dealing with content detection and
analysis.

18
00:00:55,310 --> 00:00:59,023
So, this particular first module will
cover the information landscape and

19
00:00:59,023 --> 00:01:02,027
the importance all the different
content types out there and

20
00:01:02,027 --> 00:01:05,741
the importance of automatic appro,
approaches for detecting those types, and

21
00:01:05,741 --> 00:01:10,100
for eventually extracting texts,
metadata, and information from them.

22
00:01:10,100 --> 00:01:11,830
So, just some quick notes again.

23
00:01:11,830 --> 00:01:13,960
This talk is optimized for
breadth, not depth.

24
00:01:13,960 --> 00:01:16,330
We're going to cover a lot
of things in this talk,

25
00:01:16,330 --> 00:01:19,690
so we're going to, I'm going to talk,
I'm going to try not to talk too fast.

26
00:01:19,690 --> 00:01:20,990
I might talk fast.

27
00:01:20,990 --> 00:01:24,120
But just understand that if there's
something that you didn't catch along

28
00:01:24,120 --> 00:01:27,590
the way, I have a website in which I
teach a full version of this class.

29
00:01:27,590 --> 00:01:29,690
It's at the University
of Southern California.

30
00:01:29,690 --> 00:01:33,450
The link is there, I teach a class on
search engines and information retrieval,

31
00:01:33,450 --> 00:01:36,520
so I encourage you to check that out,
there are lecture notes and slides.

32
00:01:36,520 --> 00:01:40,010
So, any of the things you don't feel like
you get enough information about just

33
00:01:40,010 --> 00:01:42,830
kind of go to that site, and
you can get more information there.

34
00:01:42,830 --> 00:01:44,860
And then also,
feel free to ask me questions, again,

35
00:01:44,860 --> 00:01:48,590
like if you saw my other lecture here, and
my other module on Big Data Fundamentals.

36
00:01:48,590 --> 00:01:52,020
Feel free to reach out to me via my
e-mail, or via Twitter, or, you know,

37
00:01:52,020 --> 00:01:53,970
however, to reach out to
me if you have a question.

38
00:01:53,970 --> 00:01:55,770
I'd be happy to try and ans, answer it.

39
00:01:55,770 --> 00:01:57,550
And again,
I want to welcome you to JPL and

40
00:01:57,550 --> 00:02:01,460
Caltech's sort of virtual summer
school on Big Data Analytics.

41
00:02:01,460 --> 00:02:03,030
So let's get right to it.

42
00:02:04,250 --> 00:02:05,410
You're dealing with content.

43
00:02:05,410 --> 00:02:07,750
You're searching for
it in a Big Data System.

44
00:02:07,750 --> 00:02:08,620
You're processing it.

45
00:02:08,620 --> 00:02:11,230
And there's a lot of different
file types that are out there.

46
00:02:11,230 --> 00:02:15,600
There's, that you might have to deal with
even in a big data scientific data system.

47
00:02:15,600 --> 00:02:19,860
You might have to parse information or
get information out of PowerPoint slides.

48
00:02:19,860 --> 00:02:21,920
You know, definitely if you're building
a search engine you're going to have to

49
00:02:21,920 --> 00:02:23,660
get it out of web documents.

50
00:02:23,660 --> 00:02:28,060
Heck nowadays, you gotta get it
out of PDF documents binary files,

51
00:02:28,060 --> 00:02:31,250
code, there's all kinds of
ways to search for code now.

52
00:02:31,250 --> 00:02:35,200
Okay, charts, graphs,
all these different types of things.

53
00:02:35,200 --> 00:02:37,807
By some estimates,
there's between 16,000 and

54
00:02:37,807 --> 00:02:41,960
51,000 different types of content
that's out there on the internet, okay?

55
00:02:41,960 --> 00:02:43,335
So this is sort of a more,

56
00:02:43,335 --> 00:02:49,090
kind of course grained estimate
from a website called filext.com.

57
00:02:49,090 --> 00:02:51,150
And that website takes
a look at web logs and

58
00:02:51,150 --> 00:02:55,340
search logs, and so forth, for various
web companies like Google and so forth.

59
00:02:55,340 --> 00:02:57,360
And what people and
URLs people are going to.

60
00:02:57,360 --> 00:03:01,447
Looks at the end of the URL and tries to
determine simply based on extension and

61
00:03:01,447 --> 00:03:03,540
URL extension if it is a new content type.

62
00:03:03,540 --> 00:03:07,350
So if you create a dot caltec
big data file that's going to

63
00:03:07,350 --> 00:03:12,800
be even if it is a text file as a new
type by this sort screen estimate.

64
00:03:12,800 --> 00:03:15,630
There are other estimates,
we'll talk about them sort of later,

65
00:03:15,630 --> 00:03:18,640
in this module, that are a little
bit more finer grain, and

66
00:03:18,640 --> 00:03:21,530
probably more accurate that sets
the amount of content types about 1200.

67
00:03:21,530 --> 00:03:23,550
Richly curated types out there.

68
00:03:23,550 --> 00:03:26,020
So what do you want to do with content?

69
00:03:26,020 --> 00:03:29,640
Well, once you know that there are all
these types, the common thing you want to

70
00:03:29,640 --> 00:03:32,260
do in big data systems a lot
of times is to parse and

71
00:03:32,260 --> 00:03:34,220
to extract information from it.

72
00:03:34,220 --> 00:03:38,350
To extract their text, texts out
of these different file types and

73
00:03:38,350 --> 00:03:40,610
content types, and
maybe their structure, for

74
00:03:40,610 --> 00:03:44,890
example, this text is important
it's emphasized or its heading.

75
00:03:44,890 --> 00:03:49,060
You know, this text is a matrix actually,
it actually corresponds to numbers and

76
00:03:49,060 --> 00:03:51,360
this role in the matrix and things.

77
00:03:51,360 --> 00:03:54,690
You want to index or maybe capture or
do something with the meta data.

78
00:03:54,690 --> 00:03:56,670
Meta data is data about data.

79
00:03:56,670 --> 00:03:59,230
The way I like to describe it
is if your data is a book,

80
00:03:59,230 --> 00:04:01,744
like a PDF file, a digital book.

81
00:04:01,744 --> 00:04:04,470
Meta data might be properties
about that book, it's author.

82
00:04:04,470 --> 00:04:06,130
If you're dealing with Tika.

83
00:04:06,130 --> 00:04:08,180
in Action it might be Chris Mattman.

84
00:04:08,180 --> 00:04:12,260
It's title, might be Tika, in Action,
it might have a meta data property ISBN,

85
00:04:12,260 --> 00:04:17,580
which is its, international
standardized book number okay for that.

86
00:04:17,580 --> 00:04:19,760
So meta data are properties about data.

87
00:04:19,760 --> 00:04:22,700
And you definitely want to extract
that from the content somehow.

88
00:04:22,700 --> 00:04:25,990
But then again, there are all these
different types so how do we do that.

89
00:04:25,990 --> 00:04:29,420
You might also want to identify
what language that content is in.

90
00:04:29,420 --> 00:04:33,550
Especially these large scale big data
systems that you guys are dealing with.

91
00:04:33,550 --> 00:04:36,590
You're dealing with,
basically a lot of the problem, you know,

92
00:04:36,590 --> 00:04:39,010
today is that everything
just isn't in English.

93
00:04:39,010 --> 00:04:39,650
Right?
You know,

94
00:04:39,650 --> 00:04:41,930
everything is in a number
of different languages.

95
00:04:43,100 --> 00:04:45,730
Defense projects are really
interested in Arabic.

96
00:04:45,730 --> 00:04:49,360
Okay, especially as we deal with
sort of international terrorism and

97
00:04:49,360 --> 00:04:50,320
things like that.

98
00:04:50,320 --> 00:04:55,290
We might, have financial interests in
decoding or understanding languages from

99
00:04:55,290 --> 00:04:59,820
Chinese documents or it might, you know,
need to take a French document from

100
00:04:59,820 --> 00:05:03,320
NASA which represents some
spacecraft design and extract.

101
00:05:03,320 --> 00:05:05,920
It's information and
I understand it in English.

102
00:05:05,920 --> 00:05:09,240
So we might need to, we, it's really
important to identify the language

103
00:05:09,240 --> 00:05:10,680
that text, the metadata, come with.

104
00:05:10,680 --> 00:05:14,520
And then potentially act on it,
like translate it or things like that.

105
00:05:14,520 --> 00:05:16,710
So, content types are really important.

106
00:05:16,710 --> 00:05:18,660
Not to mention the fact that there are so
many of them and

107
00:05:18,660 --> 00:05:20,640
there's all these things we want to
do with them but, you know,

108
00:05:20,640 --> 00:05:22,120
they're actually really important.

109
00:05:22,120 --> 00:05:23,090
Take a look at Google.

110
00:05:23,090 --> 00:05:26,010
I mean this is sort of common place now,
but, you know, maybe two,

111
00:05:26,010 --> 00:05:30,520
two years ago with the first appearance in
Google of this little kind of circle here.

112
00:05:30,520 --> 00:05:32,990
You see this little red
circle on the slide,

113
00:05:32,990 --> 00:05:35,910
of a little content type
by your search result.

114
00:05:35,910 --> 00:05:37,480
When you get back a PDF document, and

115
00:05:37,480 --> 00:05:39,130
your search result's going to
tell you it's a PDF, right?

116
00:05:39,130 --> 00:05:41,950
That's actually really important.

117
00:05:41,950 --> 00:05:43,860
When search engines like Google and

118
00:05:43,860 --> 00:05:48,890
Bing started identifying what the content
type are in search results, okay.

119
00:05:48,890 --> 00:05:52,230
Beyond that we need to identify what
the content types are because all

120
00:05:52,230 --> 00:05:55,930
these down stream applications act on,
what the content type is.

121
00:05:55,930 --> 00:06:00,183
For example, content type detection is so
important that in your browser,

122
00:06:00,183 --> 00:06:05,470
your browser the wa, the way that it knows
what to do, like Firefox or Chrome or.

123
00:06:05,470 --> 00:06:09,520
You know, Internet Explorer, whatever your
browser is, there are mappings inside of

124
00:06:09,520 --> 00:06:13,660
your browser knowing that when
you click on a video file for

125
00:06:13,660 --> 00:06:17,500
example that there's a set of applications
or handlers to deal with that video file.

126
00:06:17,500 --> 00:06:20,680
Like load Quicktime if it's a movie file,
right.

127
00:06:20,680 --> 00:06:22,830
Pull it into Excel if it's an Excel file.

128
00:06:22,830 --> 00:06:27,100
Well to be able to tell that it's
an Excel file we need some mapping.

129
00:06:27,100 --> 00:06:30,960
Of the content types, and some ability
to detect the content types, and

130
00:06:30,960 --> 00:06:33,660
then to map them to a specific
application to deal with that.

131
00:06:33,660 --> 00:06:35,020
Okay?
So, it's so important,

132
00:06:35,020 --> 00:06:36,710
it's actually appearing
in the software that,

133
00:06:36,710 --> 00:06:40,050
you know, tens of millions of
people are using nowadays.

134
00:06:40,050 --> 00:06:41,220
Content detection is important.

135
00:06:41,220 --> 00:06:45,000
It's also detect important in
the context of search engines.

136
00:06:45,000 --> 00:06:48,380
There are multiple places in
a search engine that really rely on

137
00:06:48,380 --> 00:06:49,520
content detection.

138
00:06:49,520 --> 00:06:52,410
This is sort of the conical search engine

139
00:06:52,410 --> 00:06:55,270
architecture from
the Apache nutch project.

140
00:06:55,270 --> 00:06:58,080
Nutch is sort of in an open
source web search engine.

141
00:06:58,080 --> 00:07:01,770
It implements the the architecture and
anatomy, if you will,

142
00:07:01,770 --> 00:07:05,380
of a large scale web hypertextural
search engine is defined by Britain and

143
00:07:05,380 --> 00:07:10,060
Page of their conical paper and computer
networks and ISDN systems on Google.

144
00:07:10,060 --> 00:07:13,120
Right?
So, search engines have things like, like

145
00:07:13,120 --> 00:07:17,050
protocol frameworks to download content
over different protocols like FTP, HTTP.

146
00:07:17,050 --> 00:07:23,040
They have things like parsing frameworks,
which when they get content or download

147
00:07:23,040 --> 00:07:27,790
them over a particular protocol they've
gotta parse the text on the meta data out.

148
00:07:27,790 --> 00:07:30,210
They have indexing
frameworks which decide.

149
00:07:30,210 --> 00:07:33,240
How to in, take the text and
how to take the meta data and

150
00:07:33,240 --> 00:07:36,640
how to put it in a search index,
to make it available later, for search.

151
00:07:36,640 --> 00:07:40,350
They have ranking sort of elements and
components to it,

152
00:07:40,350 --> 00:07:44,320
that allow it rank the results that
come back from a particular query.

153
00:07:44,320 --> 00:07:47,500
And so they have URL filtering and
filtering frameworks to decided,

154
00:07:47,500 --> 00:07:51,070
which URLs to go fetch, and
which to kind of throw out and discard.

155
00:07:51,070 --> 00:07:53,390
So there are a number of places in
the search engine architecture where

156
00:07:53,390 --> 00:07:55,450
content detection and
analysis is important.

157
00:07:55,450 --> 00:07:57,110
First, start with filtering.

158
00:07:57,110 --> 00:07:58,190
URL filtering.

159
00:07:58,190 --> 00:08:02,070
We may only want to build a vertical
search engine that goes after movies.

160
00:08:02,070 --> 00:08:04,770
Because maybe we're building a search
engine for Netflix and we don't

161
00:08:04,770 --> 00:08:09,270
necessarily care in our search engine to
allow you to search for URLs like HTML.

162
00:08:09,270 --> 00:08:11,140
URLs we just want present
movie files to you.

163
00:08:11,140 --> 00:08:13,210
So it's important to know and to filter.

164
00:08:13,210 --> 00:08:16,080
URLs that are only for
movies in that case.

165
00:08:16,080 --> 00:08:17,360
Take Parson for example.

166
00:08:17,360 --> 00:08:21,630
We may only want to parse
things like author or

167
00:08:21,630 --> 00:08:25,800
number of pages or you know headings or
things like that out of

168
00:08:25,800 --> 00:08:29,350
things that are documents,
like PDF documents or Word documents.

169
00:08:29,350 --> 00:08:30,810
Author, number of pages and

170
00:08:30,810 --> 00:08:33,450
things like that have no applicability for
example a JPEG image.

171
00:08:33,450 --> 00:08:34,060
Image.

172
00:08:34,060 --> 00:08:37,150
Okay, so it's really important
to know how to parse content

173
00:08:37,150 --> 00:08:39,440
in a search engine framework,
by its content type.

174
00:08:39,440 --> 00:08:42,330
And thus, we have to know the content
type to be able to detect that,

175
00:08:42,330 --> 00:08:44,870
so being able to detect
it is really important.

176
00:08:44,870 --> 00:08:46,690
Take indexing, in a search engine.

177
00:08:46,690 --> 00:08:47,830
Indexing is really important.

178
00:08:47,830 --> 00:08:52,210
We know if it's, a Microsoft Office file,
that it has a property in

179
00:08:52,210 --> 00:08:55,210
it called number of pages,
that we may want to search for later.

180
00:08:55,210 --> 00:08:58,370
Well, we only maybe want to store that or

181
00:08:58,370 --> 00:09:02,040
store that particular metadata field
if it's a Microsoft Office file.

182
00:09:02,040 --> 00:09:04,810
We need ability to detect
that sort of as well and

183
00:09:04,810 --> 00:09:08,490
that's sort of where content detection and
analysis comes in for that.

184
00:09:08,490 --> 00:09:12,090
So that's sort of the end of
the first part of of this mod,

185
00:09:12,090 --> 00:09:14,850
module on content detection and analysis.

186
00:09:14,850 --> 00:09:17,680
In the next part,
we're going to talk about some of the,

187
00:09:17,680 --> 00:09:21,220
sort of more information about
things like mime types and

188
00:09:21,220 --> 00:09:25,120
mime hierarchies, more information
about parsing and why that's important,

189
00:09:25,120 --> 00:09:28,468
more information about ways and
methodologies for doing that.

190
00:09:28,468 --> 00:09:31,010
And so we'll head right into
that next module or right now.

191
00:09:31,010 --> 00:09:31,510
Thanks.

