1
00:00:00,006 --> 00:00:01,820
It's Chris Mattmann, I'm back.

2
00:00:01,820 --> 00:00:05,500
It's time for part two of content
detection and analysis for big data.

3
00:00:05,500 --> 00:00:09,540
Welcome JPL Caltech,
virtual summer school big data analytics,

4
00:00:09,540 --> 00:00:10,270
really appreciate that.

5
00:00:10,270 --> 00:00:13,890
In the first part of this
module we covered sort of

6
00:00:13,890 --> 00:00:16,590
the importance of file types,
how there are so

7
00:00:16,590 --> 00:00:20,950
many by some estimates by 18 to 51,000
different types that are out there.

8
00:00:20,950 --> 00:00:23,600
And growing nowadays, so
the information landscape.

9
00:00:23,600 --> 00:00:27,190
And then why it's important to do sort
of content type detection, and what,

10
00:00:27,190 --> 00:00:31,090
why you want to do it, to parse text,
to parse metadata, to identify language.

11
00:00:31,090 --> 00:00:34,260
In all of these different contents that
you find in your big data systems.

12
00:00:34,260 --> 00:00:38,410
So, in this lecture,
we're going to cover a little bit more on,

13
00:00:38,410 --> 00:00:42,680
on why it's important to content detection
by looking at the mime hierarchy and

14
00:00:42,680 --> 00:00:44,200
the mime database, in a more sort of

15
00:00:45,480 --> 00:00:48,530
kind of fine grained estimate of how
many content types that are out there.

16
00:00:48,530 --> 00:00:52,290
Then, we'll look at some of the challenges
with doing things like extracting text and

17
00:00:52,290 --> 00:00:54,790
metadata from content types.

18
00:00:54,790 --> 00:00:56,370
We'll study some of those challenges and

19
00:00:56,370 --> 00:01:00,370
think about some of the approaches
to sort of mitigate that.

20
00:01:00,370 --> 00:01:03,750
So, so, I've been sort of flippantly
saying content types, and

21
00:01:03,750 --> 00:01:05,260
so forth, and file types.

22
00:01:05,260 --> 00:01:09,460
And kind of using them interchangeably
here in the first portion of this module.

23
00:01:09,460 --> 00:01:12,970
We'll kind of get a little bit
more formal here in part two, and

24
00:01:12,970 --> 00:01:18,224
we'll define file types by MIME,
MIME types.

25
00:01:18,224 --> 00:01:24,710
Which are multi sort of multi-media
internet message exchange,

26
00:01:24,710 --> 00:01:30,890
or MIME, which was a a if you call
it sort of a standard or an RFC.

27
00:01:30,890 --> 00:01:34,120
They came out of the people that were
working in the internet mail community.

28
00:01:34,120 --> 00:01:39,270
And they were interested in when you
send mail, like SMTP mail being able

29
00:01:39,270 --> 00:01:42,160
to classify and identify the different
attachments that you have.

30
00:01:42,160 --> 00:01:45,660
So, eventually this RFC
was taken over by IANA.

31
00:01:45,660 --> 00:01:49,670
The Internet Assigned Numbers Authority,
if you've ever registered a DNS site

32
00:01:49,670 --> 00:01:53,460
through now one of the many providers,
like GoDaddy or Google or whatever.

33
00:01:53,460 --> 00:01:57,250
What's happening in the background is that
eventually they're registering your site

34
00:01:57,250 --> 00:01:58,820
with the domain registrar.

35
00:01:58,820 --> 00:02:01,770
And that's sort of governed by this
international body called IANA,

36
00:02:01,770 --> 00:02:03,240
the Internet Assigned Numbers Authority.

37
00:02:03,240 --> 00:02:07,810
Well, one thing they also did at
IANA was they defined a hierarchy or

38
00:02:07,810 --> 00:02:11,170
taxonomy of MIME types, okay?

39
00:02:11,170 --> 00:02:15,110
They broke it down into
initially a seven-layer, or

40
00:02:15,110 --> 00:02:20,960
seven sort of category classification for
audio, images, multipart messages,

41
00:02:20,960 --> 00:02:25,510
video, text and application and things
like that that you see here in your,

42
00:02:25,510 --> 00:02:29,220
in your diagram in,
in terms of a mail message.

43
00:02:29,220 --> 00:02:33,400
And then after they sort of broke
it down like that as part of

44
00:02:33,400 --> 00:02:35,430
this they find this sort of taxonomy and

45
00:02:35,430 --> 00:02:40,440
ability to have sort of sub lease
on these sort of top level types.

46
00:02:40,440 --> 00:02:45,610
And eventually a tree that included
around 1200, which are the canonical sort

47
00:02:45,610 --> 00:02:49,260
of vial, and kinds of text that you see
out there on the internet nowadays.

48
00:02:49,260 --> 00:02:51,835
So along with this hierarchy
that they developed, and

49
00:02:51,835 --> 00:02:56,050
1,200 is probably a more accurate estimate
than simply looking at the URLs like,

50
00:02:56,050 --> 00:02:59,188
and, and the extension or
the end of the URLs like file.txt does.

51
00:02:59,188 --> 00:03:01,760
1200 is, is a lot more accurate.

52
00:03:01,760 --> 00:03:04,020
These are richly curated types.

53
00:03:04,020 --> 00:03:07,850
They are the primary and subtype,
so there's a parental sort of

54
00:03:07,850 --> 00:03:11,090
classification and
a hierarchical classification for that.

55
00:03:11,090 --> 00:03:17,170
The IANA's MIME type's registry also
includes different ways of detecting,uh,

56
00:03:17,170 --> 00:03:19,550
these MIME types or
file types, or content types.

57
00:03:19,550 --> 00:03:24,410
They include information about sort
of what we call Glob extensions or

58
00:03:24,410 --> 00:03:29,800
the extension pattern of the,
the file like startup.txt or startup.pdf.

59
00:03:31,010 --> 00:03:33,100
What you might see in terms of a URL?

60
00:03:33,100 --> 00:03:36,600
Like what file EXT uses to
determine if it's a PDF file or

61
00:03:36,600 --> 00:03:38,410
a text file, or whatever.

62
00:03:38,410 --> 00:03:40,410
You might see it look at the end of so

63
00:03:40,410 --> 00:03:44,580
these are also defined
in the IANA registry.

64
00:03:44,580 --> 00:03:49,220
Magic bytes which basically correspond
to digital file signatures.

65
00:03:49,220 --> 00:03:52,540
So, so all files have some sort of
digital fingerprint to them, and

66
00:03:52,540 --> 00:03:56,920
they're usually in the form of, of things
like what character set the file is in.

67
00:03:56,920 --> 00:04:01,910
Most US most Western countries
don't use something called UTF8

68
00:04:01,910 --> 00:04:03,680
where as the rest of the world does.

69
00:04:04,680 --> 00:04:06,590
Which is particular encoding format.

70
00:04:06,590 --> 00:04:10,570
So it just, immediately we can tell
a lot about, for example, a files type,

71
00:04:10,570 --> 00:04:13,340
its language or whatever just based
on the encoding that its using.

72
00:04:13,340 --> 00:04:16,890
So that's a,
a element of sort of magic byte.

73
00:04:16,890 --> 00:04:21,570
But even more so
files typically start at offsets with

74
00:04:21,570 --> 00:04:25,130
a particular byte sequence that
are authored at a particular type.

75
00:04:25,130 --> 00:04:29,960
So, for example, PDF, all PDF 7 files
start with the characters bang or

76
00:04:29,960 --> 00:04:34,520
exclamation mark PDF 7 you know,
at offset zero.

77
00:04:34,520 --> 00:04:36,670
Okay?
And these types of magic bytes or

78
00:04:36,670 --> 00:04:37,670
digital finger prints,

79
00:04:37,670 --> 00:04:41,380
our file signatures are also
defined in this IANA MIME registry.

80
00:04:41,380 --> 00:04:42,800
Okay, for that.

81
00:04:42,800 --> 00:04:48,080
So you can also use combinations or
different combinations, and

82
00:04:48,080 --> 00:04:51,700
combine them heuristically to be able
to accurately detect a file's type.

83
00:04:51,700 --> 00:04:54,170
Or you can say, look at the URL.

84
00:04:54,170 --> 00:04:56,860
If you can't discern from that,
see if there's a glob pattern, and

85
00:04:56,860 --> 00:05:01,570
at the end of the day try Magic bytes
on maybe the first 1024 bytes of

86
00:05:01,570 --> 00:05:04,360
a particular file that you get,
or something like that.

87
00:05:04,360 --> 00:05:05,890
So, so classifying and

88
00:05:05,890 --> 00:05:09,670
being able automatically use this
information to identify what

89
00:05:09,670 --> 00:05:13,680
the MIME type of a file is, allows you to
target your interaction with that file.

90
00:05:13,680 --> 00:05:17,740
Like I said, how to parse it,
what applications can read or write it, or

91
00:05:17,740 --> 00:05:20,285
you know, a number of the different things
that you need to do downstream with that.

92
00:05:20,285 --> 00:05:24,370
So basically the goals for

93
00:05:24,370 --> 00:05:29,040
sort of exploiting this information
are to be able to sort of

94
00:05:29,040 --> 00:05:32,130
accurately have a means for
representing this MIME type registry.

95
00:05:32,130 --> 00:05:36,330
And we're going to talk later about a
technology called Apache Tika that sort of

96
00:05:36,330 --> 00:05:41,450
fully realizes and implements this
MIME type registry, okay, for that.

97
00:05:41,450 --> 00:05:45,070
And then allows you to use that
information to detect files,

98
00:05:45,070 --> 00:05:48,680
to extract text and
metadata from these files, and

99
00:05:48,680 --> 00:05:51,910
to basically exploit that information and
content detection and analysis, okay?

100
00:05:51,910 --> 00:05:54,970
So, Tika, is something that we're
going to talk about later, and

101
00:05:54,970 --> 00:05:58,780
it's something that sort of is,
a softer technology that

102
00:05:58,780 --> 00:06:01,210
allows you to deal with the number
of these issues that you face.

103
00:06:03,240 --> 00:06:07,200
One thing that you face if you're trying
to deal with content detection and

104
00:06:07,200 --> 00:06:09,730
analysis, deal with all of
these different file types,

105
00:06:09,730 --> 00:06:11,770
whether you are using the IANA registry or

106
00:06:11,770 --> 00:06:15,928
you're just trying other simplistic
techniques like filext.com does is.

107
00:06:15,928 --> 00:06:20,040
You're faced with the fact
that many times to,

108
00:06:20,040 --> 00:06:23,950
once you've targeted and figured out
what type of file it is, to extract and

109
00:06:23,950 --> 00:06:27,450
get information out is typically
dependent on the type of file that it is.

110
00:06:27,450 --> 00:06:31,970
For example, to read an Office file,
like a Word, a,

111
00:06:31,970 --> 00:06:36,050
or PowerPoint, or an Excel file,
you need Microsoft Office.

112
00:06:36,050 --> 00:06:38,380
Right?
So you deal with a Photoshop file or

113
00:06:38,380 --> 00:06:42,250
a particular type of image, or a JPG or
a GIF, you might need Photoshop.

114
00:06:42,250 --> 00:06:45,470
Or you might need some other library
to read image files that can,

115
00:06:45,470 --> 00:06:51,260
like, like Gimp, if for example, you're
a Linux person or, so on and so forth.

116
00:06:51,260 --> 00:06:55,430
You might need Adobe Reader or
Acrobat Reader to read PDF files.

117
00:06:55,430 --> 00:06:57,390
So on and so forth, our preview on Mac.

118
00:06:57,390 --> 00:07:02,550
So, there are many, many, many,
many custom file types that accompany with

119
00:07:02,550 --> 00:07:06,900
them custom applications and tools
that read and write those file types.

120
00:07:06,900 --> 00:07:07,790
So, that's something that we

121
00:07:07,790 --> 00:07:10,680
want to exploit when we deal
with content detection analysis.

122
00:07:10,680 --> 00:07:12,970
Like, we want to exploit the fact
that there are these readers and

123
00:07:12,970 --> 00:07:14,200
writers out there.

124
00:07:14,200 --> 00:07:17,690
So most of the custom applications for
these file types, which also in

125
00:07:17,690 --> 00:07:22,150
turn create the proliferation of content
types and file types that are out there,

126
00:07:22,150 --> 00:07:25,750
typically include readers and writers
that let you get information from it.

127
00:07:25,750 --> 00:07:30,600
For example, Photoshop needs to have
a reader for JPG files in order to

128
00:07:30,600 --> 00:07:34,250
be able to present you information
about a JPG file, to visualize it,

129
00:07:34,250 --> 00:07:37,480
to allow you to look at the meta data
properties, and so on and so forth.

130
00:07:37,480 --> 00:07:39,340
Same goes for things like Office.

131
00:07:39,340 --> 00:07:41,430
So, those are commercial and
proprietary examples.

132
00:07:41,430 --> 00:07:44,940
It's always nice if you can find reader,
in this case, reader libraries which we're

133
00:07:44,940 --> 00:07:48,180
concerned with associated with file
types that are, that are open source.

134
00:07:48,180 --> 00:07:51,460
And as it turns out, there are many
open source libraries out there that

135
00:07:51,460 --> 00:07:54,990
accompany the software products to
read these particular file types.

136
00:07:54,990 --> 00:07:56,850
For example, Microsoft Office.

137
00:07:56,850 --> 00:08:00,890
There's a software called
Apache POI which reads and

138
00:08:00,890 --> 00:08:03,170
writes difference Office file formats and
types.

139
00:08:03,170 --> 00:08:05,230
Excel files, document files,
and things like that.

140
00:08:05,230 --> 00:08:09,830
So you can take POI and basically use
it as a parser to extract text and

141
00:08:09,830 --> 00:08:11,140
metadata from these various files.

142
00:08:11,140 --> 00:08:15,220
There's something called Apache PDFBox,
and FontBox,

143
00:08:15,220 --> 00:08:18,100
which reads PDF files,
and extracts metadata and

144
00:08:18,100 --> 00:08:22,810
information and in some cases text and
other information From these file types.

145
00:08:22,810 --> 00:08:25,870
So it's good to be able to take these
sort of reader and writer libraries and

146
00:08:25,870 --> 00:08:30,350
then exploit them, and then use them,
right, to pull out the information.

147
00:08:30,350 --> 00:08:35,060
So that's good, but what we find when we
do that, is that not all of these sort of

148
00:08:35,060 --> 00:08:38,779
parsing libraries for text and metadata
all deal with text in the same way.

149
00:08:39,820 --> 00:08:40,970
Or metadata in this same way.

150
00:08:40,970 --> 00:08:44,240
So more better at extracting text
from certain types of PDF files,

151
00:08:44,240 --> 00:08:46,270
or HTML files, or whatever.

152
00:08:46,270 --> 00:08:47,390
Some aren't as good.

153
00:08:47,390 --> 00:08:48,450
Some don't get all the text.

154
00:08:48,450 --> 00:08:49,950
Some miss information.

155
00:08:49,950 --> 00:08:54,060
Some incorrectly extract text, or
they extract it in weird character sets.

156
00:08:54,060 --> 00:08:55,570
Some are faster than others.

157
00:08:55,570 --> 00:08:58,650
Some take a long time to run,
and, you know, they're,

158
00:08:58,650 --> 00:09:00,260
they're things you want to avoid.

159
00:09:00,260 --> 00:09:02,310
Some of are more or
less reliable than others.

160
00:09:02,310 --> 00:09:05,715
Some parsing libraries, are libraries,
such that, every time or

161
00:09:05,715 --> 00:09:09,350
fifth time you call them, they cause
a big memory crash on your computer.

162
00:09:09,350 --> 00:09:10,840
If you have an out of memory thing.

163
00:09:10,840 --> 00:09:14,160
Some don't, and they represent the content
types efficiently or inefficiently.

164
00:09:14,160 --> 00:09:17,560
So, dealing with these issues
is one of the challenges in

165
00:09:17,560 --> 00:09:20,250
dealing with the integration of
these existing parsing libraries,

166
00:09:20,250 --> 00:09:22,920
to get information out of text and
metadata files.

167
00:09:22,920 --> 00:09:26,060
To be able to handle those 1200 files
in the IANA registry, you need to

168
00:09:26,060 --> 00:09:29,780
mitigate all of these different
parsers to parse those 1200 files and

169
00:09:29,780 --> 00:09:31,150
bring them together in a particular way.

170
00:09:33,250 --> 00:09:36,610
Thinking about metadata,
metadata has its own sort of issues.

171
00:09:36,610 --> 00:09:39,910
You guys are likely, I, I believe here
in the summer school are going to be

172
00:09:39,910 --> 00:09:43,610
covering metadata as a topic amongst
the other lectures and things like that.

173
00:09:43,610 --> 00:09:46,710
But just in the context relevant to
our own lecture, the, I was again,

174
00:09:46,710 --> 00:09:49,340
thinking of metadata as data about data.

175
00:09:49,340 --> 00:09:52,530
It's really important to understand
the different content types correspond to

176
00:09:52,530 --> 00:09:54,110
different metadata models.

177
00:09:54,110 --> 00:09:56,530
For example,
Word has its own metadata model,

178
00:09:56,530 --> 00:10:01,350
about 192 different metadata elements,
like, like again, author, number of pages.

179
00:10:02,402 --> 00:10:05,910
What slide you're on, if it's
a PowerPoint file, and things like that.

180
00:10:05,910 --> 00:10:11,600
So Microsoft Office actually defines this
sort of canonical metadata representation.

181
00:10:11,600 --> 00:10:14,160
EXIF is a metadata model
that's about images.

182
00:10:14,160 --> 00:10:17,570
It defines things like the number of
frames that are actually in your image,

183
00:10:17,570 --> 00:10:20,940
number of bits and
pixels in which your image was taken.

184
00:10:20,940 --> 00:10:24,060
In some cases, EXIF used to define
the geographic latitude and

185
00:10:24,060 --> 00:10:27,180
longitude by where and which your image
was taken before people started to get

186
00:10:27,180 --> 00:10:30,300
freaked out when Facebook could tell where
you are when they took pictures of you.

187
00:10:30,300 --> 00:10:32,160
And people were uploading it to Flickr and

188
00:10:32,160 --> 00:10:34,860
finding out that people could suddenly
understand where people were.

189
00:10:34,860 --> 00:10:39,230
So, EXIF has these model X and
P as a metadata model published by

190
00:10:39,230 --> 00:10:44,140
Adobe to represents for the Photoshop and
the other Adobe family of file formats.

191
00:10:44,140 --> 00:10:48,270
So, the metadata model typically
also corresponds to the content type

192
00:10:48,270 --> 00:10:49,560
that you're dealing with, okay.

193
00:10:49,560 --> 00:10:51,240
And there's lots of standards and

194
00:10:51,240 --> 00:10:54,430
models out there, and, and typically
they correspond with content type.

195
00:10:54,430 --> 00:10:59,340
And we need ways of extracting not just
the models and their attributes and so

196
00:10:59,340 --> 00:11:02,180
forth, but their values,
understanding what units.

197
00:11:02,180 --> 00:11:05,910
They are like understanding that for
example Word in,

198
00:11:05,910 --> 00:11:10,130
in, in, or in Microsoft office metadata,
number of pages is an integer, okay?

199
00:11:10,130 --> 00:11:12,090
And not a strain and things like that.

200
00:11:12,090 --> 00:11:12,610
Okay?

201
00:11:12,610 --> 00:11:13,860
And then it may have a value range.

202
00:11:17,350 --> 00:11:21,330
So thinking about metadata in the context
of actual example that maybe relevant to

203
00:11:21,330 --> 00:11:23,450
big data, you might think
about cancer research, okay.

204
00:11:23,450 --> 00:11:28,760
This is a, a, actual image, a slide
image related to sort of looking at,

205
00:11:28,760 --> 00:11:34,540
at cancer cells originally taken from
a bright a white light bronchoscopy.

206
00:11:34,540 --> 00:11:39,170
Thinking about lung cancer, this is
a cell's image taken related to that.

207
00:11:39,170 --> 00:11:42,330
And this is it's associated metadata,
okay?

208
00:11:42,330 --> 00:11:46,710
And so this metadata is going to have
things here represented in RDF format.

209
00:11:46,710 --> 00:11:49,070
And extracted that way,
is going to have things like attributes.

210
00:11:49,070 --> 00:11:51,500
Again, these are the properties
of metadata.

211
00:11:51,500 --> 00:11:55,490
This might be number of pages or, or
things like that and a particular value.

212
00:11:55,490 --> 00:11:57,960
And the metadata is also
going to have relationships,

213
00:11:57,960 --> 00:12:00,600
which are relationships between
the different attributes.

214
00:12:00,600 --> 00:12:01,230
Okay?

215
00:12:01,230 --> 00:12:05,580
Here, in this particular cancer research
example, this is a biomarker if you will.

216
00:12:05,580 --> 00:12:09,640
There's a biomarker that's related to
this particular cancerous image, and

217
00:12:09,640 --> 00:12:13,760
that biomarker's recorded as is
properties about this particular image,

218
00:12:13,760 --> 00:12:18,010
who has access to it and
relationships related to that.

219
00:12:18,010 --> 00:12:19,260
Okay?

220
00:12:19,260 --> 00:12:21,720
So these are all important things
that a content detection and

221
00:12:21,720 --> 00:12:26,090
analysis framework, in that the entire
realm of content detection and analysis.

222
00:12:26,090 --> 00:12:28,480
These are what you want to be thinking
about when you're thinking about how to

223
00:12:28,480 --> 00:12:30,470
extract and capture metadata from there.

224
00:12:30,470 --> 00:12:33,410
It's also important to
understand language, okay?

225
00:12:33,410 --> 00:12:36,650
So it's hard, you know, when you're
parsing text and metadata out of different

226
00:12:36,650 --> 00:12:39,530
file types, you really need to understand
the language that they're in, right?

227
00:12:39,530 --> 00:12:40,930
So you have a French document, you know,

228
00:12:40,930 --> 00:12:46,110
j'aime la classe de CS 572 that I use at
my, in my search engines class at USC.

229
00:12:46,110 --> 00:12:46,700
Right?

230
00:12:46,700 --> 00:12:50,060
And maybe the publisher in terms
of the metadata of this is

231
00:12:50,060 --> 00:12:54,420
L'University de Californie
en Etas-Unis de Sud, right?

232
00:12:54,420 --> 00:12:57,170
And the English equivalent is
I love the CS 572 class which

233
00:12:57,170 --> 00:12:59,710
is the text that's
present in this document.

234
00:12:59,710 --> 00:13:02,920
The metadata's publisher is
the University of Southern California.

235
00:13:02,920 --> 00:13:03,660
Okay?

236
00:13:03,660 --> 00:13:06,630
So how do you compare the extracted
text and the metadata from these two

237
00:13:06,630 --> 00:13:09,990
different documents without understanding
that one's in French and one's in English?

238
00:13:09,990 --> 00:13:13,210
Okay, but they're effectively
equivalent doc, documents, all right.

239
00:13:13,210 --> 00:13:15,520
So it's really important
from a content detection and

240
00:13:15,520 --> 00:13:19,420
analysis perspective to have means,
and hopefully automated means of

241
00:13:19,420 --> 00:13:22,030
making these type of language
identifications and detection.

242
00:13:22,030 --> 00:13:24,220
So that we can act and
understand that, in fact,

243
00:13:24,220 --> 00:13:27,010
these are the same content that we're
analyzing from a big data perspective.

244
00:13:28,790 --> 00:13:30,940
There are different methods for
language identification.

245
00:13:30,940 --> 00:13:33,890
They basically break down
to computational methods.

246
00:13:33,890 --> 00:13:39,270
And non-computational approaches, the very
common computational approach is N-grams.

247
00:13:39,270 --> 00:13:43,530
Which is basically looking at N
size sub-sequences of words, or

248
00:13:43,530 --> 00:13:47,950
grams or character sequences in,
in snippets of text.

249
00:13:47,950 --> 00:13:51,936
And using those character, sequence or
snippets, those N-sized or

250
00:13:51,936 --> 00:13:56,630
those n-word sized snippets,
to basically determine whether or

251
00:13:56,630 --> 00:14:00,570
not these snippets of text actually
belong or correspond to a language.

252
00:14:00,570 --> 00:14:05,190
As it turns out there are only so
many three grams, if you will, in,

253
00:14:05,190 --> 00:14:09,920
in the context of English and in the
context of in French, or things like that.

254
00:14:09,920 --> 00:14:13,830
Statistically, where very rapidly
by examining a snippet of text and

255
00:14:13,830 --> 00:14:17,520
comparing it against these sort of
trained and built N-grams models.

256
00:14:17,520 --> 00:14:20,090
You can very rapidly determine whether or
not this is a French or

257
00:14:20,090 --> 00:14:22,810
an English document,
based on those sub sequences of words.

258
00:14:22,810 --> 00:14:28,360
It's a computational technique based
on what, what language and what text.

259
00:14:28,360 --> 00:14:32,370
And how much sort of sample data, and
how many how much text you have to

260
00:14:32,370 --> 00:14:36,840
see before you can sort of statistically
in a good way, detect the language.

261
00:14:36,840 --> 00:14:41,180
But it's a very sort of automated
technique which tends to lend itself to

262
00:14:41,180 --> 00:14:43,670
a lot of people wanting to use
these types of approaches,

263
00:14:43,670 --> 00:14:45,315
trading accuracy usually for that.

264
00:14:45,315 --> 00:14:50,180
Non-computational approaches for language
identification and detection are tagging.

265
00:14:50,180 --> 00:14:52,902
Either having a human or
some type of automated process.

266
00:14:52,902 --> 00:14:53,450
All right, you know,

267
00:14:53,450 --> 00:14:57,440
you may tag n-grams on content after
you've run it through a particular model.

268
00:14:57,440 --> 00:15:01,780
Or you may have humans simply go in and
classify text or documents or

269
00:15:01,780 --> 00:15:06,300
content as particular language, and then
use an act based on those classifications,

270
00:15:06,300 --> 00:15:08,290
either try to learn a classifier for it.

271
00:15:08,290 --> 00:15:11,780
Or simply just you know,
have humans maintain a large repository of

272
00:15:11,780 --> 00:15:13,440
these taggings for
languages and notification.

273
00:15:13,440 --> 00:15:16,310
Which of course is not really
computationally effective on large

274
00:15:16,310 --> 00:15:17,790
amounts of data, and so on and so forth.

275
00:15:18,840 --> 00:15:21,140
Identifying the language of
content is really important,

276
00:15:21,140 --> 00:15:22,490
because once you identified the language,

277
00:15:22,490 --> 00:15:25,890
you might be able to do something called
machine translation if you have a model.

278
00:15:25,890 --> 00:15:26,880
So once you detected the language,

279
00:15:26,880 --> 00:15:29,450
you can automatically translate
from a source language, say,

280
00:15:29,450 --> 00:15:32,980
English to French, or from French
to English, and so on and so forth.

281
00:15:32,980 --> 00:15:37,040
And there's an entire field of
statistical machine translation that we,

282
00:15:37,040 --> 00:15:38,520
I'm not going to cover
in this lecture and,

283
00:15:38,520 --> 00:15:40,720
and I don't think we're,
we're covering here in the summer school.

284
00:15:40,720 --> 00:15:44,430
But I encourage you guys to take a look
into this field, because it's really sort

285
00:15:44,430 --> 00:15:48,760
of emerging especially within the realm
of content detection and analysis.

286
00:15:48,760 --> 00:15:50,030
Right?
And there are many APIs and

287
00:15:50,030 --> 00:15:51,640
tool kits to take a look at, for

288
00:15:51,640 --> 00:15:54,780
example, Google translate,
Bing translate, Lingo 24.

289
00:15:54,780 --> 00:15:58,100
These are all API based
machine translation services.

290
00:15:58,100 --> 00:16:01,292
You give it text,
it gives you back text in one language and

291
00:16:01,292 --> 00:16:05,444
you tell it what language you want it
to translate to, and it will do that.

292
00:16:05,444 --> 00:16:08,033
Or it can translate from a language
to another language and, and so

293
00:16:08,033 --> 00:16:08,690
on and so forth.

294
00:16:08,690 --> 00:16:12,040
And then there are toolkits which you
can download free and open source today.

295
00:16:12,040 --> 00:16:12,770
Many of them,

296
00:16:12,770 --> 00:16:17,480
very popular within the MT community,
are things like Moses and Joshua Decoder.

297
00:16:17,480 --> 00:16:20,870
And you can take these and you can use
them to train a model, based on data that

298
00:16:20,870 --> 00:16:24,360
you give it, to then perform machine
translation statistically after that.

299
00:16:26,120 --> 00:16:29,910
So, there's lots of challenges
when you deal of course with

300
00:16:29,910 --> 00:16:34,000
language identification, text and
metadata extraction, machine translation.

301
00:16:34,000 --> 00:16:37,190
I'll just cover a couple of these
challenges here and point them out.

302
00:16:37,190 --> 00:16:39,450
Scalability is a really big challenge.

303
00:16:39,450 --> 00:16:42,140
Well, first,
the ability to uniformly extract and

304
00:16:42,140 --> 00:16:45,730
present metadata is difficult because
there are so many metadata models.

305
00:16:45,730 --> 00:16:48,380
There aren't as many metadata
models as there are file types.

306
00:16:48,380 --> 00:16:51,480
But there are near,
almost the amount of metadata models.

307
00:16:51,480 --> 00:16:54,880
You know, almost each file type a lot
of times presents metadata models.

308
00:16:54,880 --> 00:16:57,180
They're not a lot of use between them.

309
00:16:57,180 --> 00:16:59,950
Scale is really important when
you're doing content detection and

310
00:16:59,950 --> 00:17:02,460
analysis on large numbers of documents.

311
00:17:02,460 --> 00:17:05,550
Being able to do things
automatically is really important,

312
00:17:05,550 --> 00:17:07,410
especially as we sort
of scale out on that.

313
00:17:07,410 --> 00:17:11,630
Having humans in the loop is really
prohibitive in this environment.

314
00:17:11,630 --> 00:17:15,410
We need approaches for automatically
doing these types of activities.

315
00:17:15,410 --> 00:17:18,090
Integrating third-party parsing
libraries is really difficult, for

316
00:17:18,090 --> 00:17:20,460
the reasons,
in the aforementioned reasons that I have.

317
00:17:20,460 --> 00:17:23,160
Like I said you know,
some perform in different ways.

318
00:17:23,160 --> 00:17:27,060
But also, a number of these parsing
libraries have intrinsic dependencies

319
00:17:27,060 --> 00:17:27,690
on one another.

320
00:17:27,690 --> 00:17:31,200
Like some depend on all sorts of other
software that when you start to use these

321
00:17:31,200 --> 00:17:35,370
parsing libraries, you basically end up
having a sort of big, bloated you know,

322
00:17:35,370 --> 00:17:37,622
software that is suddenly
hard to download and

323
00:17:37,622 --> 00:17:39,740
very hard to install on other machines.

324
00:17:39,740 --> 00:17:43,770
And then, again, there, there's not,
there's a lack of sort of uniform ways, or

325
00:17:43,770 --> 00:17:47,350
extraction interfaces, or bringing
these libraries together, to all bring,

326
00:17:47,350 --> 00:17:51,030
take out text and metadata and language
and so fort from different content.

327
00:17:52,530 --> 00:17:56,320
Another one of the challenges is that
this is a graph that shows one of,

328
00:17:56,320 --> 00:18:00,110
probably benefacto to look at for
the content detection analysis Tika.

329
00:18:00,110 --> 00:18:03,260
And it shows the quality of it's
ability to do charset detection and

330
00:18:03,260 --> 00:18:05,590
language detection, and
it's actually really difficult.

331
00:18:05,590 --> 00:18:08,850
This is on the Y or
on the X axis for this is.

332
00:18:08,850 --> 00:18:12,180
The number of pages that it's
looking at within a sample based on

333
00:18:12,180 --> 00:18:15,036
this crawl part of the public
terabyte dataset project.

334
00:18:15,036 --> 00:18:20,790
And on the on the y axis
is the percentage correct

335
00:18:20,790 --> 00:18:26,570
of those number of documents that Tika was
looking at from this particular web crawl.

336
00:18:26,570 --> 00:18:28,660
The percentage graph how
often it got corrected,

337
00:18:28,660 --> 00:18:33,670
what the associated character set or,
or language was correctly.

338
00:18:33,670 --> 00:18:38,158
And, and you can tell that the,
the character sets that appear most

339
00:18:38,158 --> 00:18:45,004
commonly in sorts of downloads or sample
of web pages, like, for example, ISO85591.

340
00:18:45,004 --> 00:18:49,220
There is a very commonly occurring
character set type on this graph.

341
00:18:49,220 --> 00:18:51,120
Tika actually detects very poorly.

342
00:18:52,190 --> 00:18:55,760
If you look at the percentage correct
that it got there on the y y axis.

343
00:18:55,760 --> 00:18:58,860
And, you know, the ones that
it detects very correctly are,

344
00:18:58,860 --> 00:19:00,780
are typically the ones that
don't occur very much.

345
00:19:00,780 --> 00:19:03,540
And so language and
character detection is hard.

346
00:19:03,540 --> 00:19:06,060
We're looking at this, you know,
in content detection, and

347
00:19:06,060 --> 00:19:07,040
analysis, libraries.

348
00:19:07,040 --> 00:19:10,180
We're looking at sort of non
computational approaches.

349
00:19:10,180 --> 00:19:13,600
We're looking at more accurate curated
approaches, and things like that.

350
00:19:13,600 --> 00:19:17,560
And that's something that we're going to
have to work on and get better at.

351
00:19:17,560 --> 00:19:21,350
main, maintaining a MIME database is
really difficult, especially as more and

352
00:19:21,350 --> 00:19:25,580
more content types are being added
in identifying and things like that.

353
00:19:25,580 --> 00:19:28,520
So, just ensure that data
base MIME information can be

354
00:19:28,520 --> 00:19:30,550
kept up to date is really difficult.

355
00:19:30,550 --> 00:19:35,190
Making sure that basically various sort of

356
00:19:35,190 --> 00:19:37,650
down stream applications
beyond simply search engines.

357
00:19:37,650 --> 00:19:39,810
But things like, you know,
web browsers and

358
00:19:39,810 --> 00:19:42,800
web servers and so forth,
are actively using and, and

359
00:19:42,800 --> 00:19:46,650
leveraging this MIME information in the
content detection and analysis approaches.

360
00:19:46,650 --> 00:19:48,450
It's really difficult because there's,

361
00:19:48,450 --> 00:19:51,290
like everybody needs to
understand content nowadays.

362
00:19:51,290 --> 00:19:54,590
And they don't always know about the right
ways to perform this type of dete,

363
00:19:54,590 --> 00:19:55,690
detection and analysis.

364
00:19:55,690 --> 00:19:57,170
And they don't always know
about the libraries and

365
00:19:57,170 --> 00:19:58,730
toolkits that are out there to do it.

366
00:19:58,730 --> 00:20:00,070
And so a lot of times,
they roll their own.

367
00:20:00,070 --> 00:20:02,560
And so
that's a real big challenge nowadays.

368
00:20:02,560 --> 00:20:06,980
It's just dealing with the fact that,
that's going on right now.

369
00:20:06,980 --> 00:20:12,020
So, just real quick, to wrap up on this
second part of the content detection and

370
00:20:12,020 --> 00:20:13,896
analysis module here.

371
00:20:13,896 --> 00:20:18,190
We covered sort of different
MIME detection, what MIME is.

372
00:20:18,190 --> 00:20:20,570
We covered parsing, and
integrating parsing libraries.

373
00:20:20,570 --> 00:20:24,180
We covered language identification,
machine translation.

374
00:20:24,180 --> 00:20:26,080
Common metadata models and formats, and

375
00:20:26,080 --> 00:20:29,160
then we talked about kind of
the challenges in each of these areas.

376
00:20:29,160 --> 00:20:32,456
And in the final portion of this module,
we're going to cover specific

377
00:20:32,456 --> 00:20:36,760
technology called Apache Tika, which fully
realizes and implements a number of,

378
00:20:36,760 --> 00:20:39,700
of the types of tools that you'll need
to do content detection and analysis.

379
00:20:39,700 --> 00:20:40,430
So thanks, and

380
00:20:40,430 --> 00:20:43,620
we'll cover that in the next the final
portion here of this module.

