1
00:00:01,410 --> 00:00:02,132
Hey guys.

2
00:00:02,132 --> 00:00:03,120
It's Chris Mattmann.

3
00:00:03,120 --> 00:00:04,790
Here we are for the wrap-up,

4
00:00:04,790 --> 00:00:10,510
the penultimate, the third portion of
the Content Detection and Analysis for

5
00:00:10,510 --> 00:00:14,510
Big Data module at the JPL-Caltech Virtual
Summer School on Big Data Analytics.

6
00:00:14,510 --> 00:00:18,170
So thank you, my time is shortly about
to end with you guys, so we'll try and

7
00:00:18,170 --> 00:00:19,150
wrap it up and make it fun.

8
00:00:20,280 --> 00:00:21,120
Let's get into it.

9
00:00:21,120 --> 00:00:23,240
We're talking about content detection and
analysis.

10
00:00:23,240 --> 00:00:25,540
I talked to you guys about
the information landscape,

11
00:00:25,540 --> 00:00:27,694
all the different file
formats that are out there.

12
00:00:27,694 --> 00:00:31,181
Like I said by some
estimates 18,000 to 51,000.

13
00:00:31,181 --> 00:00:34,785
We talked about why it's important to be
able to detect content types, because we

14
00:00:34,785 --> 00:00:38,480
want to parse text and metadata, and
language information out of them.

15
00:00:38,480 --> 00:00:40,800
And it's really important because
there's all these different uses and

16
00:00:40,800 --> 00:00:45,550
search engines, and your browser, and
big data systems for translating things

17
00:00:45,550 --> 00:00:49,090
from one language to another or
from one content type to another.

18
00:00:49,090 --> 00:00:51,390
We talked about some of
the challenges in doing all this.

19
00:00:51,390 --> 00:00:55,160
Ranging from the fact that integrating
third party parsing libraries are hard.

20
00:00:55,160 --> 00:00:57,980
The different file types typically
have different software that

21
00:00:57,980 --> 00:00:59,760
extracts information from them.

22
00:00:59,760 --> 00:01:04,680
We have the fact that detecting language
and software is, or and content, is hard.

23
00:01:04,680 --> 00:01:06,150
Metadata is typically hard.

24
00:01:06,150 --> 00:01:08,290
There are many metadata
models that are out there.

25
00:01:08,290 --> 00:01:10,540
So we talked about a lot of the issues,

26
00:01:10,540 --> 00:01:12,950
the kind of fundamentals of
content detection and analysis.

27
00:01:12,950 --> 00:01:16,650
And in this lecture, we're going to talk
about specific, a specific approach and

28
00:01:16,650 --> 00:01:19,080
a specific technology called Apache Tika.

29
00:01:19,080 --> 00:01:22,170
Which has a number of the sort of
tools that, will give you a number of

30
00:01:22,170 --> 00:01:25,890
your tools in your tool belt to deal with
content detection and analysis systems.

31
00:01:25,890 --> 00:01:28,380
It's something that you might want to
deal with in the context of big data.

32
00:01:30,246 --> 00:01:33,110
So this is an introduction to Apache Tika.

33
00:01:33,110 --> 00:01:38,580
So I'll talk about what, what Apache Tika
is, where did it come from, yeah, what its

34
00:01:38,580 --> 00:01:42,780
current versions are, and how to download
it and use it, and what can it do.

35
00:01:42,780 --> 00:01:46,130
Okay, so, Tika, you know, we've been
talking about content detection and

36
00:01:46,130 --> 00:01:48,090
analysis in the context of big data.

37
00:01:48,090 --> 00:01:51,710
Tika is a content detection and
analysis toolkit, okay?

38
00:01:51,710 --> 00:01:53,030
Ultimately it was written in Java.

39
00:01:53,030 --> 00:01:58,540
It's a set of Java APIs that provide MIME
type, automatic MIME type identification,

40
00:01:58,540 --> 00:02:00,790
language identification,
metadata extraction, and

41
00:02:00,790 --> 00:02:03,620
text extraction, and integration
with various parsing libraries.

42
00:02:03,620 --> 00:02:07,910
In fact, all of the parsing libraries
that support those 1,200 file types or

43
00:02:07,910 --> 00:02:10,610
content types for
Miana are supported by Tika,

44
00:02:10,610 --> 00:02:14,150
including all of the parsers necessary
to extract text and metadata from them.

45
00:02:15,460 --> 00:02:17,420
It has a rich metadata API for

46
00:02:17,420 --> 00:02:20,060
representing sort of key
multivalued metadata models and

47
00:02:20,060 --> 00:02:24,150
different metadata models, and actual
instances of metadata in that domain.

48
00:02:24,150 --> 00:02:25,650
It has a command line interface for

49
00:02:25,650 --> 00:02:29,240
interacting with it as
well as a REST interface.

50
00:02:29,240 --> 00:02:32,960
As well as being ported
in a GUI interface but

51
00:02:32,960 --> 00:02:36,180
it's also been ported to a number of
different sort of downstream libraries.

52
00:02:36,180 --> 00:02:40,710
So Tika exists in,
as a dot net library, it exists as

53
00:02:40,710 --> 00:02:46,430
a Python module it's been ported
to Debian as an RPM module.

54
00:02:47,490 --> 00:02:50,750
Another thing, and this isn't
covered necessarily in the slides.

55
00:02:50,750 --> 00:02:54,840
But I just want to mention this to you
too is, if you have used Drupal or

56
00:02:54,840 --> 00:02:59,120
Alfresco, or Clone, or
most any major content management system.

57
00:02:59,120 --> 00:03:02,250
And you've issued a search against
that content management system,

58
00:03:02,250 --> 00:03:04,020
you've interacted with Tika.

59
00:03:04,020 --> 00:03:07,620
So Tika is part of Apache Solr,
which is, it's sort of de facto, one of

60
00:03:07,620 --> 00:03:12,490
the de facto search engines that are out
there, as part of the Lucene project.

61
00:03:12,490 --> 00:03:16,580
It powers even things like elastic search
and other search engines technologies.

62
00:03:16,580 --> 00:03:20,740
And when you throw files at Solr or
an elastic search engine Lucene and

63
00:03:20,740 --> 00:03:23,240
those files are automatically parsed,
their text and

64
00:03:23,240 --> 00:03:26,410
their metadata is extracted,
the thing that is doing it is Tika.

65
00:03:26,410 --> 00:03:30,340
Okay, so things that we do to Tika
has a lot down stream impact.

66
00:03:30,340 --> 00:03:33,600
It's downloaded thousands of time per day
from the Apache software foundation, and

67
00:03:33,600 --> 00:03:35,480
it's a really sort of prevalence project.

68
00:03:36,620 --> 00:03:41,240
The original idea for Tika came from
myself and a guy named Jerome Charron,

69
00:03:41,240 --> 00:03:44,880
who is a Frenchman,
who's one of the notch project committer.

70
00:03:44,880 --> 00:03:48,240
So, we were working on
notch before Hadoop.

71
00:03:48,240 --> 00:03:52,030
And we saw all the distributed
computing and big data wonks go out and

72
00:03:52,030 --> 00:03:55,130
help to create Hadoop
including some of our efforts.

73
00:03:55,130 --> 00:03:56,990
And we decided that
the content detection and

74
00:03:56,990 --> 00:04:00,370
analysis portions of search engines that
were present in notch at the time just

75
00:04:00,370 --> 00:04:04,552
like Hadoop was, really deserved it's
own you know, first class status.

76
00:04:04,552 --> 00:04:07,770
And it, content detection and
analysis was itself an emerging field, and

77
00:04:07,770 --> 00:04:10,870
it was something that we wanted to
focus on just in our own projects.

78
00:04:10,870 --> 00:04:14,700
So we originally proposed Tika as a sub
project to Apache Lucene in 2006.

79
00:04:14,700 --> 00:04:19,985
Didn't get much traction,
so we had some help and

80
00:04:19,985 --> 00:04:25,530
mentorship from a gentleman by the name
of Jukka Zitting all right Jukka Zitting.

81
00:04:25,530 --> 00:04:31,090
And basically Jukka was really familiar
with Apache he had been involved

82
00:04:31,090 --> 00:04:34,780
in the foundation for a number of years
helping to build the Jackrabbit Project,

83
00:04:34,780 --> 00:04:36,600
which was a content management system.

84
00:04:36,600 --> 00:04:40,000
And he came because he needed these
types of content detection and

85
00:04:40,000 --> 00:04:43,820
analysis capabilities, because they were
building CMSs like Jackrabbit at the time,

86
00:04:43,820 --> 00:04:45,680
and things like Alfresco and whatever.

87
00:04:45,680 --> 00:04:47,290
And they needed these
texting capabilities.

88
00:04:47,290 --> 00:04:53,830
So, Jukka helped us basically reformat or
recapitulate our, our proposal for Tika.

89
00:04:53,830 --> 00:04:57,540
And we were accepted as
a incubator project in,

90
00:04:57,540 --> 00:05:01,780
in I think it's like circa
the 2007 time frame at Apache.

91
00:05:01,780 --> 00:05:06,990
And then after making I think seven or
eight releases we we were graduate,

92
00:05:06,990 --> 00:05:10,030
we graduated to a Lucene
sub-project at the time.

93
00:05:10,030 --> 00:05:12,510
So we became officially
a part of Apache Lucene, and

94
00:05:12,510 --> 00:05:18,500
then in 2010 we graduated to within Tika,
our own full level, or top level project.

95
00:05:18,500 --> 00:05:21,880
Indicating we're not just simply used
within Lucene, but we've, you know,

96
00:05:21,880 --> 00:05:24,850
established a status where we're you know,
managing our own project and

97
00:05:24,850 --> 00:05:28,429
it has a number of uses outside of
just that particular software effort.

98
00:05:29,650 --> 00:05:31,910
So you want to use Tika
just to get started today.

99
00:05:31,910 --> 00:05:33,543
The current version is 1.5.

100
00:05:33,543 --> 00:05:34,760
1.6 will be out soon,

101
00:05:34,760 --> 00:05:39,100
you can actually download it from
the URL here on your website, grab it.

102
00:05:39,100 --> 00:05:44,690
If you're in Unix you can do this in,
in WIndows, as well, but alias the command

103
00:05:44,690 --> 00:05:51,920
Tika to the command java-jar, the
tika-app-1.5, or the version number, .jar.

104
00:05:51,920 --> 00:05:53,700
And then you suddenly have
a command called tika.

105
00:05:53,700 --> 00:05:54,520
Feed it a file.

106
00:05:54,520 --> 00:05:57,190
Type a file, like a word document into it.

107
00:05:57,190 --> 00:05:59,530
If you don't provide it any parameters,
out the other end,

108
00:05:59,530 --> 00:06:02,445
comes all of the extracted
text in XHTML format.

109
00:06:02,445 --> 00:06:06,440
because we represent,
the extracted text internally as XHTML.

110
00:06:06,440 --> 00:06:11,070
We do this because, downstream, we can
use the simple API for XML processing or

111
00:06:11,070 --> 00:06:16,080
SAX, to have sort of
content handlers parse and

112
00:06:16,080 --> 00:06:20,670
extract text and do downstream sort of
pipelines from the extracted content.

113
00:06:20,670 --> 00:06:23,450
So we can create derivative analysis and
so forth.

114
00:06:23,450 --> 00:06:26,080
And we use Sax because it has
low memory footprint, and

115
00:06:26,080 --> 00:06:27,570
it's something that we can pipeline.

116
00:06:27,570 --> 00:06:30,140
It only loads single nodes or
single characters at a time,

117
00:06:30,140 --> 00:06:33,850
instead of loading the entire
structure of the document into memory,

118
00:06:33,850 --> 00:06:36,600
like Dom does within the content of XML.

119
00:06:36,600 --> 00:06:39,190
So give it a file, out comes the text.

120
00:06:39,190 --> 00:06:43,010
Ask for its metadata by passing the -m
flag, and out comes the metadata for that.

121
00:06:43,010 --> 00:06:45,470
So this is interacting with
Tika on the command line,

122
00:06:45,470 --> 00:06:48,840
this is calling the Java API
which exists internally for it.

123
00:06:48,840 --> 00:06:50,970
You can also interact
with Tika in other ways,

124
00:06:50,970 --> 00:06:53,860
like you could, for example,
write a Java program.

125
00:06:53,860 --> 00:06:57,670
So you want to detect MIME types,
automatically from Java.

126
00:06:57,670 --> 00:07:00,140
We provide a facade class
called the Tika Facade,

127
00:07:00,140 --> 00:07:03,100
it's a static class,
you can simply call static methods on it.

128
00:07:03,100 --> 00:07:04,770
So you want the type of a particular file,

129
00:07:04,770 --> 00:07:08,200
like maybe you have an input
stream to that file.

130
00:07:08,200 --> 00:07:10,930
You've created some buffered reader or
something in Java.

131
00:07:10,930 --> 00:07:13,850
Well, feed it in,
feed Tika.detect in input stream, and

132
00:07:13,850 --> 00:07:18,686
out comes the type of the file,
classified along the the hierarchy.

133
00:07:18,686 --> 00:07:23,260
Give it a java.io.fileobject,
give it a URL,

134
00:07:23,260 --> 00:07:27,480
you can even use Tika on URL and
include it in URL file type processing.

135
00:07:27,480 --> 00:07:30,330
So if there's a remote file that
you want to determine what the file

136
00:07:30,330 --> 00:07:33,330
type is for it, give it a java.net.URL.

137
00:07:33,330 --> 00:07:38,780
You can also give Tika a string, which
is a pointer a string path to file, and

138
00:07:38,780 --> 00:07:39,520
it'll give you a,

139
00:07:39,520 --> 00:07:43,810
you know, on your local file system,
and it'll give you a type for that too.

140
00:07:43,810 --> 00:07:47,940
So that type detection is powered
through several detectors,

141
00:07:47,940 --> 00:07:50,380
which heuristically combine
different mime type approaches,

142
00:07:50,380 --> 00:07:53,590
which we talked about in
the second portion of this module.

143
00:07:53,590 --> 00:07:57,950
And it does that by leveraging and
exploiting the full Iona uuh,

144
00:07:57,950 --> 00:08:01,516
MIME registry, which Tika maintains and
actually, arguable, has an even

145
00:08:01,516 --> 00:08:06,883
more up-to-date and well-curated version
of that Iona registry than Iona does.

146
00:08:06,883 --> 00:08:10,880
Tika is one of the projects
that constantly updates it

147
00:08:10,880 --> 00:08:13,930
with new files types, and we're constantly
getting people contacting us and

148
00:08:13,930 --> 00:08:15,170
saying a new file type isn't there.

149
00:08:15,170 --> 00:08:16,020
Can you please add it?

150
00:08:16,020 --> 00:08:17,890
Or do this, or we have other people and

151
00:08:17,890 --> 00:08:21,620
new contributors come into the project and
becoming committees and project committee

152
00:08:21,620 --> 00:08:25,200
management committee members
themselves through their contributions.

153
00:08:25,200 --> 00:08:27,820
And so
we have a very robust representation of

154
00:08:27,820 --> 00:08:29,260
the MIME registry in XML.

155
00:08:29,260 --> 00:08:30,800
It's constantly being added to.

156
00:08:30,800 --> 00:08:34,610
You can also fork or create your own
derivative of this XML registry.

157
00:08:34,610 --> 00:08:38,600
And you may ne, never contribute it back
upstream to us, but just maintain it for

158
00:08:38,600 --> 00:08:39,430
your project if you want.

159
00:08:39,430 --> 00:08:42,330
How do we get text and

160
00:08:42,330 --> 00:08:46,430
parse information out of file types
in Tika from a Java API perspective?

161
00:08:46,430 --> 00:08:47,430
Here you go.

162
00:08:47,430 --> 00:08:50,570
And part of the Tika facade,
there's a parse to string method.

163
00:08:50,570 --> 00:08:56,490
So basically you give it an InputStream
to a file a java.io.fileobject, the URL,

164
00:08:56,490 --> 00:09:00,180
and out comes the extracted text
in the form of a string format.

165
00:09:00,180 --> 00:09:03,190
If it's really big, like you're dealing
with big data, you can get a reader

166
00:09:03,190 --> 00:09:07,440
in a java.io.reader and you can have sort
of a callback mechanism in which you just

167
00:09:07,440 --> 00:09:11,990
read the bytes you know, reads some subset
of the bytes, and process it as you will.

168
00:09:11,990 --> 00:09:14,860
So, if it's a lot of, a lot of text
that you're going to get out from

169
00:09:14,860 --> 00:09:17,670
some big file, then you can deal with it,
with a reader.

170
00:09:17,670 --> 00:09:19,280
How would you do language
detection in Tika?

171
00:09:19,280 --> 00:09:22,810
Language has Tika has
a language identifier class.

172
00:09:22,810 --> 00:09:25,730
You simply create and
instantiate a new language identifier.

173
00:09:25,730 --> 00:09:27,570
You, you give it a language profile.

174
00:09:27,570 --> 00:09:29,710
You give it some snippet of a file.

175
00:09:29,710 --> 00:09:32,390
You can either give it the text from all,
you know,

176
00:09:32,390 --> 00:09:36,900
all of the file of a particular language,
or you can give it some snippet of text.

177
00:09:36,900 --> 00:09:42,940
And what the language identifier does is
compares that text for the file that you

178
00:09:42,940 --> 00:09:48,210
give it, using Ngram detection, and
spits out a language detection for a file.

179
00:09:48,210 --> 00:09:50,870
So basically the,

180
00:09:50,870 --> 00:09:54,280
the Ngram detection mechanism here
originally in Tika came from Nutch.

181
00:09:54,280 --> 00:09:56,840
We're looking at other
Ngram approaches as well,

182
00:09:56,840 --> 00:10:00,840
like Google has an Ngram detection
library that's out there in Google code.

183
00:10:00,840 --> 00:10:06,040
And there are other language
identifications sort of mechanisms and

184
00:10:06,040 --> 00:10:07,250
things like that.

185
00:10:07,250 --> 00:10:09,850
For example there's something called
magic in python which looks at

186
00:10:09,850 --> 00:10:11,870
things like keerset/s and languages and

187
00:10:11,870 --> 00:10:14,560
things like that, that we can
potentially integrate down the road.

188
00:10:15,855 --> 00:10:18,320
Tika has a metadata object for
representing metadata.

189
00:10:18,320 --> 00:10:19,870
It's a key multi valued structure.

190
00:10:19,870 --> 00:10:21,900
So you create a metadata object.

191
00:10:21,900 --> 00:10:24,860
You also have access to all of
the meta data model attributes and

192
00:10:24,860 --> 00:10:27,830
keys that Tika knows about, which is
currently about 20 metadata models

193
00:10:27,830 --> 00:10:31,920
including Dublin Core, including HTTP
headers, Creative Commons metadata,

194
00:10:31,920 --> 00:10:35,000
Climate Forecast metadata is particularly
relevant within the context of

195
00:10:35,000 --> 00:10:38,070
the science domain if you're
dealing with climate model output.

196
00:10:38,070 --> 00:10:39,900
Or remote sensing data and things.

197
00:10:39,900 --> 00:10:42,910
And what you do is you set or
create metadata keys.

198
00:10:42,910 --> 00:10:45,610
And keys could also have
multiple values for them.

199
00:10:45,610 --> 00:10:48,930
You see here in this particular example,
we're adding two values for

200
00:10:48,930 --> 00:10:50,460
the key format.

201
00:10:50,460 --> 00:10:54,570
So we're setting metadata.format equal
to both text HTML and text plain.

202
00:10:54,570 --> 00:10:56,890
And you find this a lot,
this is really relevant to,

203
00:10:56,890 --> 00:11:02,070
for example, a file type a hierarchical or
multiple mind types associated with that.

204
00:11:02,070 --> 00:11:05,800
Okay, can also run Tika from
the command line as a GUI.

205
00:11:05,800 --> 00:11:08,200
This will start up a little
GUI in which you can drop and

206
00:11:08,200 --> 00:11:11,650
drag files onto the GUI and have it
extract the text, the metadata, and

207
00:11:11,650 --> 00:11:15,260
the language and various tabs and
just sort of interact with it that way.

208
00:11:15,260 --> 00:11:19,410
Not a lot of people use Tika in the GUI
form, it's mostly just a debugging thing,

209
00:11:19,410 --> 00:11:23,400
but I thought I'd just show it you know,
for pedagogical purposes.

210
00:11:24,810 --> 00:11:29,010
Can integrate Tika into your application
in a number of different ways?

211
00:11:29,010 --> 00:11:31,490
At its core, it's built using Maven.

212
00:11:31,490 --> 00:11:33,270
Tika's built using Maven,
so you can use Tika,

213
00:11:33,270 --> 00:11:37,010
all of the Tika jars are published
on the central repository for Maven.

214
00:11:37,010 --> 00:11:40,710
So if you have a Maven project in
Java you can integrate Tika into your

215
00:11:40,710 --> 00:11:42,520
project simply by referencing Tika.

216
00:11:44,200 --> 00:11:48,920
Various Tika modules in your Maven
project and various versions like 1.5.

217
00:11:48,920 --> 00:11:51,980
Tika sort of has a layered
architecture as a core library that

218
00:11:51,980 --> 00:11:56,890
includes all the code for part,
includes all of the, the parsing API and

219
00:11:56,890 --> 00:12:00,430
the MIME identification framework and
language identification framework.

220
00:12:00,430 --> 00:12:04,390
Then specific parsers and
all of the various third party parsers and

221
00:12:04,390 --> 00:12:07,200
libraries for handling those twelve
hundred different content types

222
00:12:07,200 --> 00:12:10,740
are part of an, a module on top
of Tika called tika-parsers.

223
00:12:10,740 --> 00:12:13,600
On top of that is tika-app,
that's the command line and

224
00:12:13,600 --> 00:12:16,870
GUI interface to Tika that sits
on top of the parsers and core.

225
00:12:16,870 --> 00:12:22,010
Bundle is an OSGI interface,
t is Tika in OSGI environments on top of,

226
00:12:23,120 --> 00:12:26,620
parsers as well, but not shown in
this diagram is a Tika rest server.

227
00:12:26,620 --> 00:12:30,910
So, it's called tika server and that's the
jacksar server to present Tika as a rest

228
00:12:30,910 --> 00:12:34,900
tpi, and then downstream of these even are
various bundles of Tika, in libraries and

229
00:12:34,900 --> 00:12:35,960
integrations, like in tiki-p-,

230
00:12:35,960 --> 00:12:39,740
the Tika python library, the .NET version.

231
00:12:39,740 --> 00:12:42,490
If you're familiar with MIT's Julia
language, which is a really

232
00:12:42,490 --> 00:12:47,340
popular language that's emerging right
now, there's a project called taro.jl.

233
00:12:47,340 --> 00:12:50,620
And that is effectively Tika
imported to the Julia language.

234
00:12:50,620 --> 00:12:54,632
Okay so you can use it sort of in
that context if you're dealing with

235
00:12:54,632 --> 00:12:59,720
.NET Python, Julia, Java you know, and
then anything that can speak a rest

236
00:12:59,720 --> 00:13:03,650
service can use Tika and incorporate
it into your application that way.

237
00:13:03,650 --> 00:13:04,960
Okay.
You can use it,

238
00:13:04,960 --> 00:13:07,310
incorporate Tika into
your Eclipse project.

239
00:13:07,310 --> 00:13:09,590
If you're using Eclipse,
your Ant project, or whatever.

240
00:13:09,590 --> 00:13:11,610
It's, it's integratable in
a number of different ways.

241
00:13:13,040 --> 00:13:17,030
So, there's lots of information about
Tika on the tika.apache.org website.

242
00:13:17,030 --> 00:13:18,240
We have public mailing lists and

243
00:13:18,240 --> 00:13:21,250
archives at Apache, and
encourage you to check those out.

244
00:13:21,250 --> 00:13:24,240
You can search them via google,
because all of Apache's mailing lists and

245
00:13:24,240 --> 00:13:26,060
communications are archived by google, and

246
00:13:26,060 --> 00:13:28,520
most major search engines,
as well as the mailarchives.com.

247
00:13:28,520 --> 00:13:33,020
If you are thinking about ways to extend
Tika, you might think about doing some

248
00:13:33,020 --> 00:13:37,500
project in Tika either you know, during
the summer school as a side project.

249
00:13:37,500 --> 00:13:39,907
And here are some possible ideas,
adding parsers for

250
00:13:39,907 --> 00:13:41,567
content types are always welcome.

251
00:13:41,567 --> 00:13:44,580
If Tika doesn't support a content
type that you're dealing with in

252
00:13:44,580 --> 00:13:47,240
your big data project,
please add it or you know,

253
00:13:47,240 --> 00:13:51,440
Omnigraphal is one that we have basic
support for but it's not really good.

254
00:13:52,550 --> 00:13:55,230
Expanding the ability to handle
random access file parsing.

255
00:13:55,230 --> 00:13:59,060
Like if there's a file in which
the parsing library for it needs to

256
00:13:59,060 --> 00:14:02,560
load the whole thing into memory, we don't
have a load of good support for that.

257
00:14:02,560 --> 00:14:06,760
We have to deal with file formats
that support random access file for

258
00:14:06,760 --> 00:14:10,070
sort of parsing, so we don't have
a great set of support for that like and

259
00:14:10,070 --> 00:14:12,030
this is common in scientific data formats.

260
00:14:12,030 --> 00:14:15,880
So, any contributions there that you
can make Tika handle scientific data

261
00:14:15,880 --> 00:14:18,450
formats better would be much appreciated.

262
00:14:18,450 --> 00:14:21,960
Improving language and charset detection
as I showed there in the second module,

263
00:14:21,960 --> 00:14:25,820
it could use a lot of improvement,
so it's always welcome.

264
00:14:25,820 --> 00:14:28,540
We have an emerging set of
machine translation API's in Tika

265
00:14:28,540 --> 00:14:32,010
that are going to come out in one sixth if
you're interested in machine translation,

266
00:14:32,010 --> 00:14:33,940
translating from one language to another.

267
00:14:33,940 --> 00:14:36,980
Contributions there would be really
welcomed to, and it's also a good little

268
00:14:36,980 --> 00:14:40,820
side project if you're interested in big
data content detection and analysis.

269
00:14:40,820 --> 00:14:44,910
So, I want to acknowledge the material
that was provided by my collegue he and

270
00:14:44,910 --> 00:14:47,340
I have co authored many talks on Tika,

271
00:14:47,340 --> 00:14:50,960
Jukka Zitting sort of inspired some
of the material behind these talks.

272
00:14:50,960 --> 00:14:54,390
And those slides there on Slideshare
will give you some thoughts and

273
00:14:54,390 --> 00:14:55,380
further references on that.

274
00:14:55,380 --> 00:14:59,960
And some other further references are my
search engines class at USC and it's,

275
00:14:59,960 --> 00:15:03,470
it's on search engines and information
retrieval of which content detection and

276
00:15:03,470 --> 00:15:05,210
analysis is a really huge part.

277
00:15:05,210 --> 00:15:07,970
My home page at USC,
the book on Tika called,

278
00:15:07,970 --> 00:15:10,070
Tika in Action you can
take a look at that.

279
00:15:10,070 --> 00:15:11,810
And then, the Apache Tika website.

280
00:15:11,810 --> 00:15:14,640
So, thanks, I'm Chris Mattmann and
I encourage you to contact me if

281
00:15:14,640 --> 00:15:17,340
you're interested on content detection and
analysis.

282
00:15:17,340 --> 00:15:19,160
And thank you to JPL and Caltech.

283
00:15:19,160 --> 00:15:22,290
And I hope you're enjoying
your stay here at

284
00:15:22,290 --> 00:15:25,468
the JPL Caltech virtual Summer
school in big data analytics.

285
00:15:25,468 --> 00:15:25,968
Thanks

