Hey guys. It's Chris Mattmann. Here we are for the wrap-up, the penultimate, the third portion of the Content Detection and Analysis for Big Data module at the JPL-Caltech Virtual Summer School on Big Data Analytics. So thank you, my time is shortly about to end with you guys, so we'll try and wrap it up and make it fun. Let's get into it. We're talking about content detection and analysis. I talked to you guys about the information landscape, all the different file formats that are out there. Like I said by some estimates 18,000 to 51,000. We talked about why it's important to be able to detect content types, because we want to parse text and metadata, and language information out of them. And it's really important because there's all these different uses and search engines, and your browser, and big data systems for translating things from one language to another or from one content type to another. We talked about some of the challenges in doing all this. Ranging from the fact that integrating third party parsing libraries are hard. The different file types typically have different software that extracts information from them. We have the fact that detecting language and software is, or and content, is hard. Metadata is typically hard. There are many metadata models that are out there. So we talked about a lot of the issues, the kind of fundamentals of content detection and analysis. And in this lecture, we're going to talk about specific, a specific approach and a specific technology called Apache Tika. Which has a number of the sort of tools that, will give you a number of your tools in your tool belt to deal with content detection and analysis systems. It's something that you might want to deal with in the context of big data. So this is an introduction to Apache Tika. So I'll talk about what, what Apache Tika is, where did it come from, yeah, what its current versions are, and how to download it and use it, and what can it do. Okay, so, Tika, you know, we've been talking about content detection and analysis in the context of big data. Tika is a content detection and analysis toolkit, okay? Ultimately it was written in Java. It's a set of Java APIs that provide MIME type, automatic MIME type identification, language identification, metadata extraction, and text extraction, and integration with various parsing libraries. In fact, all of the parsing libraries that support those 1,200 file types or content types for Miana are supported by Tika, including all of the parsers necessary to extract text and metadata from them. It has a rich metadata API for representing sort of key multivalued metadata models and different metadata models, and actual instances of metadata in that domain. It has a command line interface for interacting with it as well as a REST interface. As well as being ported in a GUI interface but it's also been ported to a number of different sort of downstream libraries. So Tika exists in, as a dot net library, it exists as a Python module it's been ported to Debian as an RPM module. Another thing, and this isn't covered necessarily in the slides. But I just want to mention this to you too is, if you have used Drupal or Alfresco, or Clone, or most any major content management system. And you've issued a search against that content management system, you've interacted with Tika. So Tika is part of Apache Solr, which is, it's sort of de facto, one of the de facto search engines that are out there, as part of the Lucene project. It powers even things like elastic search and other search engines technologies. And when you throw files at Solr or an elastic search engine Lucene and those files are automatically parsed, their text and their metadata is extracted, the thing that is doing it is Tika. Okay, so things that we do to Tika has a lot down stream impact. It's downloaded thousands of time per day from the Apache software foundation, and it's a really sort of prevalence project. The original idea for Tika came from myself and a guy named Jerome Charron, who is a Frenchman, who's one of the notch project committer. So, we were working on notch before Hadoop. And we saw all the distributed computing and big data wonks go out and help to create Hadoop including some of our efforts. And we decided that the content detection and analysis portions of search engines that were present in notch at the time just like Hadoop was, really deserved it's own you know, first class status. And it, content detection and analysis was itself an emerging field, and it was something that we wanted to focus on just in our own projects. So we originally proposed Tika as a sub project to Apache Lucene in 2006. Didn't get much traction, so we had some help and mentorship from a gentleman by the name of Jukka Zitting all right Jukka Zitting. And basically Jukka was really familiar with Apache he had been involved in the foundation for a number of years helping to build the Jackrabbit Project, which was a content management system. And he came because he needed these types of content detection and analysis capabilities, because they were building CMSs like Jackrabbit at the time, and things like Alfresco and whatever. And they needed these texting capabilities. So, Jukka helped us basically reformat or recapitulate our, our proposal for Tika. And we were accepted as a incubator project in, in I think it's like circa the 2007 time frame at Apache. And then after making I think seven or eight releases we we were graduate, we graduated to a Lucene sub-project at the time. So we became officially a part of Apache Lucene, and then in 2010 we graduated to within Tika, our own full level, or top level project. Indicating we're not just simply used within Lucene, but we've, you know, established a status where we're you know, managing our own project and it has a number of uses outside of just that particular software effort. So you want to use Tika just to get started today. The current version is 1.5. 1.6 will be out soon, you can actually download it from the URL here on your website, grab it. If you're in Unix you can do this in, in WIndows, as well, but alias the command Tika to the command java-jar, the tika-app-1.5, or the version number, .jar. And then you suddenly have a command called tika. Feed it a file. Type a file, like a word document into it. If you don't provide it any parameters, out the other end, comes all of the extracted text in XHTML format. because we represent, the extracted text internally as XHTML. We do this because, downstream, we can use the simple API for XML processing or SAX, to have sort of content handlers parse and extract text and do downstream sort of pipelines from the extracted content. So we can create derivative analysis and so forth. And we use Sax because it has low memory footprint, and it's something that we can pipeline. It only loads single nodes or single characters at a time, instead of loading the entire structure of the document into memory, like Dom does within the content of XML. So give it a file, out comes the text. Ask for its metadata by passing the -m flag, and out comes the metadata for that. So this is interacting with Tika on the command line, this is calling the Java API which exists internally for it. You can also interact with Tika in other ways, like you could, for example, write a Java program. So you want to detect MIME types, automatically from Java. We provide a facade class called the Tika Facade, it's a static class, you can simply call static methods on it. So you want the type of a particular file, like maybe you have an input stream to that file. You've created some buffered reader or something in Java. Well, feed it in, feed Tika.detect in input stream, and out comes the type of the file, classified along the the hierarchy. Give it a java.io.fileobject, give it a URL, you can even use Tika on URL and include it in URL file type processing. So if there's a remote file that you want to determine what the file type is for it, give it a java.net.URL. You can also give Tika a string, which is a pointer a string path to file, and it'll give you a, you know, on your local file system, and it'll give you a type for that too. So that type detection is powered through several detectors, which heuristically combine different mime type approaches, which we talked about in the second portion of this module. And it does that by leveraging and exploiting the full Iona uuh, MIME registry, which Tika maintains and actually, arguable, has an even more up-to-date and well-curated version of that Iona registry than Iona does. Tika is one of the projects that constantly updates it with new files types, and we're constantly getting people contacting us and saying a new file type isn't there. Can you please add it? Or do this, or we have other people and new contributors come into the project and becoming committees and project committee management committee members themselves through their contributions. And so we have a very robust representation of the MIME registry in XML. It's constantly being added to. You can also fork or create your own derivative of this XML registry. And you may ne, never contribute it back upstream to us, but just maintain it for your project if you want. How do we get text and parse information out of file types in Tika from a Java API perspective? Here you go. And part of the Tika facade, there's a parse to string method. So basically you give it an InputStream to a file a java.io.fileobject, the URL, and out comes the extracted text in the form of a string format. If it's really big, like you're dealing with big data, you can get a reader in a java.io.reader and you can have sort of a callback mechanism in which you just read the bytes you know, reads some subset of the bytes, and process it as you will. So, if it's a lot of, a lot of text that you're going to get out from some big file, then you can deal with it, with a reader. How would you do language detection in Tika? Language has Tika has a language identifier class. You simply create and instantiate a new language identifier. You, you give it a language profile. You give it some snippet of a file. You can either give it the text from all, you know, all of the file of a particular language, or you can give it some snippet of text. And what the language identifier does is compares that text for the file that you give it, using Ngram detection, and spits out a language detection for a file. So basically the, the Ngram detection mechanism here originally in Tika came from Nutch. We're looking at other Ngram approaches as well, like Google has an Ngram detection library that's out there in Google code. And there are other language identifications sort of mechanisms and things like that. For example there's something called magic in python which looks at things like keerset/s and languages and things like that, that we can potentially integrate down the road. Tika has a metadata object for representing metadata. It's a key multi valued structure. So you create a metadata object. You also have access to all of the meta data model attributes and keys that Tika knows about, which is currently about 20 metadata models including Dublin Core, including HTTP headers, Creative Commons metadata, Climate Forecast metadata is particularly relevant within the context of the science domain if you're dealing with climate model output. Or remote sensing data and things. And what you do is you set or create metadata keys. And keys could also have multiple values for them. You see here in this particular example, we're adding two values for the key format. So we're setting metadata.format equal to both text HTML and text plain. And you find this a lot, this is really relevant to, for example, a file type a hierarchical or multiple mind types associated with that. Okay, can also run Tika from the command line as a GUI. This will start up a little GUI in which you can drop and drag files onto the GUI and have it extract the text, the metadata, and the language and various tabs and just sort of interact with it that way. Not a lot of people use Tika in the GUI form, it's mostly just a debugging thing, but I thought I'd just show it you know, for pedagogical purposes. Can integrate Tika into your application in a number of different ways? At its core, it's built using Maven. Tika's built using Maven, so you can use Tika, all of the Tika jars are published on the central repository for Maven. So if you have a Maven project in Java you can integrate Tika into your project simply by referencing Tika. Various Tika modules in your Maven project and various versions like 1.5. Tika sort of has a layered architecture as a core library that includes all the code for part, includes all of the, the parsing API and the MIME identification framework and language identification framework. Then specific parsers and all of the various third party parsers and libraries for handling those twelve hundred different content types are part of an, a module on top of Tika called tika-parsers. On top of that is tika-app, that's the command line and GUI interface to Tika that sits on top of the parsers and core. Bundle is an OSGI interface, t is Tika in OSGI environments on top of, parsers as well, but not shown in this diagram is a Tika rest server. So, it's called tika server and that's the jacksar server to present Tika as a rest tpi, and then downstream of these even are various bundles of Tika, in libraries and integrations, like in tiki-p-, the Tika python library, the .NET version. If you're familiar with MIT's Julia language, which is a really popular language that's emerging right now, there's a project called taro.jl. And that is effectively Tika imported to the Julia language. Okay so you can use it sort of in that context if you're dealing with .NET Python, Julia, Java you know, and then anything that can speak a rest service can use Tika and incorporate it into your application that way. Okay. You can use it, incorporate Tika into your Eclipse project. If you're using Eclipse, your Ant project, or whatever. It's, it's integratable in a number of different ways. So, there's lots of information about Tika on the tika.apache.org website. We have public mailing lists and archives at Apache, and encourage you to check those out. You can search them via google, because all of Apache's mailing lists and communications are archived by google, and most major search engines, as well as the mailarchives.com. If you are thinking about ways to extend Tika, you might think about doing some project in Tika either you know, during the summer school as a side project. And here are some possible ideas, adding parsers for content types are always welcome. If Tika doesn't support a content type that you're dealing with in your big data project, please add it or you know, Omnigraphal is one that we have basic support for but it's not really good. Expanding the ability to handle random access file parsing. Like if there's a file in which the parsing library for it needs to load the whole thing into memory, we don't have a load of good support for that. We have to deal with file formats that support random access file for sort of parsing, so we don't have a great set of support for that like and this is common in scientific data formats. So, any contributions there that you can make Tika handle scientific data formats better would be much appreciated. Improving language and charset detection as I showed there in the second module, it could use a lot of improvement, so it's always welcome. We have an emerging set of machine translation API's in Tika that are going to come out in one sixth if you're interested in machine translation, translating from one language to another. Contributions there would be really welcomed to, and it's also a good little side project if you're interested in big data content detection and analysis. So, I want to acknowledge the material that was provided by my collegue he and I have co authored many talks on Tika, Jukka Zitting sort of inspired some of the material behind these talks. And those slides there on Slideshare will give you some thoughts and further references on that. And some other further references are my search engines class at USC and it's, it's on search engines and information retrieval of which content detection and analysis is a really huge part. My home page at USC, the book on Tika called, Tika in Action you can take a look at that. And then, the Apache Tika website. So, thanks, I'm Chris Mattmann and I encourage you to contact me if you're interested on content detection and analysis. And thank you to JPL and Caltech. And I hope you're enjoying your stay here at the JPL Caltech virtual Summer school in big data analytics. Thanks