It's Chris Mattmann, I'm back. It's time for part two of content detection and analysis for big data. Welcome JPL Caltech, virtual summer school big data analytics, really appreciate that. In the first part of this module we covered sort of the importance of file types, how there are so many by some estimates by 18 to 51,000 different types that are out there. And growing nowadays, so the information landscape. And then why it's important to do sort of content type detection, and what, why you want to do it, to parse text, to parse metadata, to identify language. In all of these different contents that you find in your big data systems. So, in this lecture, we're going to cover a little bit more on, on why it's important to content detection by looking at the mime hierarchy and the mime database, in a more sort of kind of fine grained estimate of how many content types that are out there. Then, we'll look at some of the challenges with doing things like extracting text and metadata from content types. We'll study some of those challenges and think about some of the approaches to sort of mitigate that. So, so, I've been sort of flippantly saying content types, and so forth, and file types. And kind of using them interchangeably here in the first portion of this module. We'll kind of get a little bit more formal here in part two, and we'll define file types by MIME, MIME types. Which are multi sort of multi-media internet message exchange, or MIME, which was a a if you call it sort of a standard or an RFC. They came out of the people that were working in the internet mail community. And they were interested in when you send mail, like SMTP mail being able to classify and identify the different attachments that you have. So, eventually this RFC was taken over by IANA. The Internet Assigned Numbers Authority, if you've ever registered a DNS site through now one of the many providers, like GoDaddy or Google or whatever. What's happening in the background is that eventually they're registering your site with the domain registrar. And that's sort of governed by this international body called IANA, the Internet Assigned Numbers Authority. Well, one thing they also did at IANA was they defined a hierarchy or taxonomy of MIME types, okay? They broke it down into initially a seven-layer, or seven sort of category classification for audio, images, multipart messages, video, text and application and things like that that you see here in your, in your diagram in, in terms of a mail message. And then after they sort of broke it down like that as part of this they find this sort of taxonomy and ability to have sort of sub lease on these sort of top level types. And eventually a tree that included around 1200, which are the canonical sort of vial, and kinds of text that you see out there on the internet nowadays. So along with this hierarchy that they developed, and 1,200 is probably a more accurate estimate than simply looking at the URLs like, and, and the extension or the end of the URLs like file.txt does. 1200 is, is a lot more accurate. These are richly curated types. They are the primary and subtype, so there's a parental sort of classification and a hierarchical classification for that. The IANA's MIME type's registry also includes different ways of detecting,uh, these MIME types or file types, or content types. They include information about sort of what we call Glob extensions or the extension pattern of the, the file like startup.txt or startup.pdf. What you might see in terms of a URL? Like what file EXT uses to determine if it's a PDF file or a text file, or whatever. You might see it look at the end of so these are also defined in the IANA registry. Magic bytes which basically correspond to digital file signatures. So, so all files have some sort of digital fingerprint to them, and they're usually in the form of, of things like what character set the file is in. Most US most Western countries don't use something called UTF8 where as the rest of the world does. Which is particular encoding format. So it just, immediately we can tell a lot about, for example, a files type, its language or whatever just based on the encoding that its using. So that's a, a element of sort of magic byte. But even more so files typically start at offsets with a particular byte sequence that are authored at a particular type. So, for example, PDF, all PDF 7 files start with the characters bang or exclamation mark PDF 7 you know, at offset zero. Okay? And these types of magic bytes or digital finger prints, our file signatures are also defined in this IANA MIME registry. Okay, for that. So you can also use combinations or different combinations, and combine them heuristically to be able to accurately detect a file's type. Or you can say, look at the URL. If you can't discern from that, see if there's a glob pattern, and at the end of the day try Magic bytes on maybe the first 1024 bytes of a particular file that you get, or something like that. So, so classifying and being able automatically use this information to identify what the MIME type of a file is, allows you to target your interaction with that file. Like I said, how to parse it, what applications can read or write it, or you know, a number of the different things that you need to do downstream with that. So basically the goals for sort of exploiting this information are to be able to sort of accurately have a means for representing this MIME type registry. And we're going to talk later about a technology called Apache Tika that sort of fully realizes and implements this MIME type registry, okay, for that. And then allows you to use that information to detect files, to extract text and metadata from these files, and to basically exploit that information and content detection and analysis, okay? So, Tika, is something that we're going to talk about later, and it's something that sort of is, a softer technology that allows you to deal with the number of these issues that you face. One thing that you face if you're trying to deal with content detection and analysis, deal with all of these different file types, whether you are using the IANA registry or you're just trying other simplistic techniques like filext.com does is. You're faced with the fact that many times to, once you've targeted and figured out what type of file it is, to extract and get information out is typically dependent on the type of file that it is. For example, to read an Office file, like a Word, a, or PowerPoint, or an Excel file, you need Microsoft Office. Right? So you deal with a Photoshop file or a particular type of image, or a JPG or a GIF, you might need Photoshop. Or you might need some other library to read image files that can, like, like Gimp, if for example, you're a Linux person or, so on and so forth. You might need Adobe Reader or Acrobat Reader to read PDF files. So on and so forth, our preview on Mac. So, there are many, many, many, many custom file types that accompany with them custom applications and tools that read and write those file types. So, that's something that we want to exploit when we deal with content detection analysis. Like, we want to exploit the fact that there are these readers and writers out there. So most of the custom applications for these file types, which also in turn create the proliferation of content types and file types that are out there, typically include readers and writers that let you get information from it. For example, Photoshop needs to have a reader for JPG files in order to be able to present you information about a JPG file, to visualize it, to allow you to look at the meta data properties, and so on and so forth. Same goes for things like Office. So, those are commercial and proprietary examples. It's always nice if you can find reader, in this case, reader libraries which we're concerned with associated with file types that are, that are open source. And as it turns out, there are many open source libraries out there that accompany the software products to read these particular file types. For example, Microsoft Office. There's a software called Apache POI which reads and writes difference Office file formats and types. Excel files, document files, and things like that. So you can take POI and basically use it as a parser to extract text and metadata from these various files. There's something called Apache PDFBox, and FontBox, which reads PDF files, and extracts metadata and information and in some cases text and other information From these file types. So it's good to be able to take these sort of reader and writer libraries and then exploit them, and then use them, right, to pull out the information. So that's good, but what we find when we do that, is that not all of these sort of parsing libraries for text and metadata all deal with text in the same way. Or metadata in this same way. So more better at extracting text from certain types of PDF files, or HTML files, or whatever. Some aren't as good. Some don't get all the text. Some miss information. Some incorrectly extract text, or they extract it in weird character sets. Some are faster than others. Some take a long time to run, and, you know, they're, they're things you want to avoid. Some of are more or less reliable than others. Some parsing libraries, are libraries, such that, every time or fifth time you call them, they cause a big memory crash on your computer. If you have an out of memory thing. Some don't, and they represent the content types efficiently or inefficiently. So, dealing with these issues is one of the challenges in dealing with the integration of these existing parsing libraries, to get information out of text and metadata files. To be able to handle those 1200 files in the IANA registry, you need to mitigate all of these different parsers to parse those 1200 files and bring them together in a particular way. Thinking about metadata, metadata has its own sort of issues. You guys are likely, I, I believe here in the summer school are going to be covering metadata as a topic amongst the other lectures and things like that. But just in the context relevant to our own lecture, the, I was again, thinking of metadata as data about data. It's really important to understand the different content types correspond to different metadata models. For example, Word has its own metadata model, about 192 different metadata elements, like, like again, author, number of pages. What slide you're on, if it's a PowerPoint file, and things like that. So Microsoft Office actually defines this sort of canonical metadata representation. EXIF is a metadata model that's about images. It defines things like the number of frames that are actually in your image, number of bits and pixels in which your image was taken. In some cases, EXIF used to define the geographic latitude and longitude by where and which your image was taken before people started to get freaked out when Facebook could tell where you are when they took pictures of you. And people were uploading it to Flickr and finding out that people could suddenly understand where people were. So, EXIF has these model X and P as a metadata model published by Adobe to represents for the Photoshop and the other Adobe family of file formats. So, the metadata model typically also corresponds to the content type that you're dealing with, okay. And there's lots of standards and models out there, and, and typically they correspond with content type. And we need ways of extracting not just the models and their attributes and so forth, but their values, understanding what units. They are like understanding that for example Word in, in, in, or in Microsoft office metadata, number of pages is an integer, okay? And not a strain and things like that. Okay? And then it may have a value range. So thinking about metadata in the context of actual example that maybe relevant to big data, you might think about cancer research, okay. This is a, a, actual image, a slide image related to sort of looking at, at cancer cells originally taken from a bright a white light bronchoscopy. Thinking about lung cancer, this is a cell's image taken related to that. And this is it's associated metadata, okay? And so this metadata is going to have things here represented in RDF format. And extracted that way, is going to have things like attributes. Again, these are the properties of metadata. This might be number of pages or, or things like that and a particular value. And the metadata is also going to have relationships, which are relationships between the different attributes. Okay? Here, in this particular cancer research example, this is a biomarker if you will. There's a biomarker that's related to this particular cancerous image, and that biomarker's recorded as is properties about this particular image, who has access to it and relationships related to that. Okay? So these are all important things that a content detection and analysis framework, in that the entire realm of content detection and analysis. These are what you want to be thinking about when you're thinking about how to extract and capture metadata from there. It's also important to understand language, okay? So it's hard, you know, when you're parsing text and metadata out of different file types, you really need to understand the language that they're in, right? So you have a French document, you know, j'aime la classe de CS 572 that I use at my, in my search engines class at USC. Right? And maybe the publisher in terms of the metadata of this is L'University de Californie en Etas-Unis de Sud, right? And the English equivalent is I love the CS 572 class which is the text that's present in this document. The metadata's publisher is the University of Southern California. Okay? So how do you compare the extracted text and the metadata from these two different documents without understanding that one's in French and one's in English? Okay, but they're effectively equivalent doc, documents, all right. So it's really important from a content detection and analysis perspective to have means, and hopefully automated means of making these type of language identifications and detection. So that we can act and understand that, in fact, these are the same content that we're analyzing from a big data perspective. There are different methods for language identification. They basically break down to computational methods. And non-computational approaches, the very common computational approach is N-grams. Which is basically looking at N size sub-sequences of words, or grams or character sequences in, in snippets of text. And using those character, sequence or snippets, those N-sized or those n-word sized snippets, to basically determine whether or not these snippets of text actually belong or correspond to a language. As it turns out there are only so many three grams, if you will, in, in the context of English and in the context of in French, or things like that. Statistically, where very rapidly by examining a snippet of text and comparing it against these sort of trained and built N-grams models. You can very rapidly determine whether or not this is a French or an English document, based on those sub sequences of words. It's a computational technique based on what, what language and what text. And how much sort of sample data, and how many how much text you have to see before you can sort of statistically in a good way, detect the language. But it's a very sort of automated technique which tends to lend itself to a lot of people wanting to use these types of approaches, trading accuracy usually for that. Non-computational approaches for language identification and detection are tagging. Either having a human or some type of automated process. All right, you know, you may tag n-grams on content after you've run it through a particular model. Or you may have humans simply go in and classify text or documents or content as particular language, and then use an act based on those classifications, either try to learn a classifier for it. Or simply just you know, have humans maintain a large repository of these taggings for languages and notification. Which of course is not really computationally effective on large amounts of data, and so on and so forth. Identifying the language of content is really important, because once you identified the language, you might be able to do something called machine translation if you have a model. So once you detected the language, you can automatically translate from a source language, say, English to French, or from French to English, and so on and so forth. And there's an entire field of statistical machine translation that we, I'm not going to cover in this lecture and, and I don't think we're, we're covering here in the summer school. But I encourage you guys to take a look into this field, because it's really sort of emerging especially within the realm of content detection and analysis. Right? And there are many APIs and tool kits to take a look at, for example, Google translate, Bing translate, Lingo 24. These are all API based machine translation services. You give it text, it gives you back text in one language and you tell it what language you want it to translate to, and it will do that. Or it can translate from a language to another language and, and so on and so forth. And then there are toolkits which you can download free and open source today. Many of them, very popular within the MT community, are things like Moses and Joshua Decoder. And you can take these and you can use them to train a model, based on data that you give it, to then perform machine translation statistically after that. So, there's lots of challenges when you deal of course with language identification, text and metadata extraction, machine translation. I'll just cover a couple of these challenges here and point them out. Scalability is a really big challenge. Well, first, the ability to uniformly extract and present metadata is difficult because there are so many metadata models. There aren't as many metadata models as there are file types. But there are near, almost the amount of metadata models. You know, almost each file type a lot of times presents metadata models. They're not a lot of use between them. Scale is really important when you're doing content detection and analysis on large numbers of documents. Being able to do things automatically is really important, especially as we sort of scale out on that. Having humans in the loop is really prohibitive in this environment. We need approaches for automatically doing these types of activities. Integrating third-party parsing libraries is really difficult, for the reasons, in the aforementioned reasons that I have. Like I said you know, some perform in different ways. But also, a number of these parsing libraries have intrinsic dependencies on one another. Like some depend on all sorts of other software that when you start to use these parsing libraries, you basically end up having a sort of big, bloated you know, software that is suddenly hard to download and very hard to install on other machines. And then, again, there, there's not, there's a lack of sort of uniform ways, or extraction interfaces, or bringing these libraries together, to all bring, take out text and metadata and language and so fort from different content. Another one of the challenges is that this is a graph that shows one of, probably benefacto to look at for the content detection analysis Tika. And it shows the quality of it's ability to do charset detection and language detection, and it's actually really difficult. This is on the Y or on the X axis for this is. The number of pages that it's looking at within a sample based on this crawl part of the public terabyte dataset project. And on the on the y axis is the percentage correct of those number of documents that Tika was looking at from this particular web crawl. The percentage graph how often it got corrected, what the associated character set or, or language was correctly. And, and you can tell that the, the character sets that appear most commonly in sorts of downloads or sample of web pages, like, for example, ISO85591. There is a very commonly occurring character set type on this graph. Tika actually detects very poorly. If you look at the percentage correct that it got there on the y y axis. And, you know, the ones that it detects very correctly are, are typically the ones that don't occur very much. And so language and character detection is hard. We're looking at this, you know, in content detection, and analysis, libraries. We're looking at sort of non computational approaches. We're looking at more accurate curated approaches, and things like that. And that's something that we're going to have to work on and get better at. main, maintaining a MIME database is really difficult, especially as more and more content types are being added in identifying and things like that. So, just ensure that data base MIME information can be kept up to date is really difficult. Making sure that basically various sort of down stream applications beyond simply search engines. But things like, you know, web browsers and web servers and so forth, are actively using and, and leveraging this MIME information in the content detection and analysis approaches. It's really difficult because there's, like everybody needs to understand content nowadays. And they don't always know about the right ways to perform this type of dete, detection and analysis. And they don't always know about the libraries and toolkits that are out there to do it. And so a lot of times, they roll their own. And so that's a real big challenge nowadays. It's just dealing with the fact that, that's going on right now. So, just real quick, to wrap up on this second part of the content detection and analysis module here. We covered sort of different MIME detection, what MIME is. We covered parsing, and integrating parsing libraries. We covered language identification, machine translation. Common metadata models and formats, and then we talked about kind of the challenges in each of these areas. And in the final portion of this module, we're going to cover specific technology called Apache Tika, which fully realizes and implements a number of, of the types of tools that you'll need to do content detection and analysis. So thanks, and we'll cover that in the next the final portion here of this module.