To a different module. You might have seen my Big Data Architecture Fundamental module here earlier in the class. So welcome to a completely different module, three part series here. We're going to talk about Content Detection and Analysis for Big Data okay. So we're going to cover a number of different topics. We're going to talk about sort of the landscape of all of the different content types that are out there when you're building big data systems. Like search engines or just data analysis systems, what types of file formats are out there, what their meaning is, what their importance is. Why it's important to be able to automatically and rapidly detect them. We'll talk about some of the challenges in doing so, in detecting them, in extracting text and information from them, and extracting their language. Different approaches to doing that and sort of the final the part of this sort of three part model here will cover a specific technology called Apache Tika. Which is really good to have in your tool belt for dealing with content detection and analysis. So, this particular first module will cover the information landscape and the importance all the different content types out there and the importance of automatic appro, approaches for detecting those types, and for eventually extracting texts, metadata, and information from them. So, just some quick notes again. This talk is optimized for breadth, not depth. We're going to cover a lot of things in this talk, so we're going to, I'm going to talk, I'm going to try not to talk too fast. I might talk fast. But just understand that if there's something that you didn't catch along the way, I have a website in which I teach a full version of this class. It's at the University of Southern California. The link is there, I teach a class on search engines and information retrieval, so I encourage you to check that out, there are lecture notes and slides. So, any of the things you don't feel like you get enough information about just kind of go to that site, and you can get more information there. And then also, feel free to ask me questions, again, like if you saw my other lecture here, and my other module on Big Data Fundamentals. Feel free to reach out to me via my e-mail, or via Twitter, or, you know, however, to reach out to me if you have a question. I'd be happy to try and ans, answer it. And again, I want to welcome you to JPL and Caltech's sort of virtual summer school on Big Data Analytics. So let's get right to it. You're dealing with content. You're searching for it in a Big Data System. You're processing it. And there's a lot of different file types that are out there. There's, that you might have to deal with even in a big data scientific data system. You might have to parse information or get information out of PowerPoint slides. You know, definitely if you're building a search engine you're going to have to get it out of web documents. Heck nowadays, you gotta get it out of PDF documents binary files, code, there's all kinds of ways to search for code now. Okay, charts, graphs, all these different types of things. By some estimates, there's between 16,000 and 51,000 different types of content that's out there on the internet, okay? So this is sort of a more, kind of course grained estimate from a website called filext.com. And that website takes a look at web logs and search logs, and so forth, for various web companies like Google and so forth. And what people and URLs people are going to. Looks at the end of the URL and tries to determine simply based on extension and URL extension if it is a new content type. So if you create a dot caltec big data file that's going to be even if it is a text file as a new type by this sort screen estimate. There are other estimates, we'll talk about them sort of later, in this module, that are a little bit more finer grain, and probably more accurate that sets the amount of content types about 1200. Richly curated types out there. So what do you want to do with content? Well, once you know that there are all these types, the common thing you want to do in big data systems a lot of times is to parse and to extract information from it. To extract their text, texts out of these different file types and content types, and maybe their structure, for example, this text is important it's emphasized or its heading. You know, this text is a matrix actually, it actually corresponds to numbers and this role in the matrix and things. You want to index or maybe capture or do something with the meta data. Meta data is data about data. The way I like to describe it is if your data is a book, like a PDF file, a digital book. Meta data might be properties about that book, it's author. If you're dealing with Tika. in Action it might be Chris Mattman. It's title, might be Tika, in Action, it might have a meta data property ISBN, which is its, international standardized book number okay for that. So meta data are properties about data. And you definitely want to extract that from the content somehow. But then again, there are all these different types so how do we do that. You might also want to identify what language that content is in. Especially these large scale big data systems that you guys are dealing with. You're dealing with, basically a lot of the problem, you know, today is that everything just isn't in English. Right? You know, everything is in a number of different languages. Defense projects are really interested in Arabic. Okay, especially as we deal with sort of international terrorism and things like that. We might, have financial interests in decoding or understanding languages from Chinese documents or it might, you know, need to take a French document from NASA which represents some spacecraft design and extract. It's information and I understand it in English. So we might need to, we, it's really important to identify the language that text, the metadata, come with. And then potentially act on it, like translate it or things like that. So, content types are really important. Not to mention the fact that there are so many of them and there's all these things we want to do with them but, you know, they're actually really important. Take a look at Google. I mean this is sort of common place now, but, you know, maybe two, two years ago with the first appearance in Google of this little kind of circle here. You see this little red circle on the slide, of a little content type by your search result. When you get back a PDF document, and your search result's going to tell you it's a PDF, right? That's actually really important. When search engines like Google and Bing started identifying what the content type are in search results, okay. Beyond that we need to identify what the content types are because all these down stream applications act on, what the content type is. For example, content type detection is so important that in your browser, your browser the wa, the way that it knows what to do, like Firefox or Chrome or. You know, Internet Explorer, whatever your browser is, there are mappings inside of your browser knowing that when you click on a video file for example that there's a set of applications or handlers to deal with that video file. Like load Quicktime if it's a movie file, right. Pull it into Excel if it's an Excel file. Well to be able to tell that it's an Excel file we need some mapping. Of the content types, and some ability to detect the content types, and then to map them to a specific application to deal with that. Okay? So, it's so important, it's actually appearing in the software that, you know, tens of millions of people are using nowadays. Content detection is important. It's also detect important in the context of search engines. There are multiple places in a search engine that really rely on content detection. This is sort of the conical search engine architecture from the Apache nutch project. Nutch is sort of in an open source web search engine. It implements the the architecture and anatomy, if you will, of a large scale web hypertextural search engine is defined by Britain and Page of their conical paper and computer networks and ISDN systems on Google. Right? So, search engines have things like, like protocol frameworks to download content over different protocols like FTP, HTTP. They have things like parsing frameworks, which when they get content or download them over a particular protocol they've gotta parse the text on the meta data out. They have indexing frameworks which decide. How to in, take the text and how to take the meta data and how to put it in a search index, to make it available later, for search. They have ranking sort of elements and components to it, that allow it rank the results that come back from a particular query. And so they have URL filtering and filtering frameworks to decided, which URLs to go fetch, and which to kind of throw out and discard. So there are a number of places in the search engine architecture where content detection and analysis is important. First, start with filtering. URL filtering. We may only want to build a vertical search engine that goes after movies. Because maybe we're building a search engine for Netflix and we don't necessarily care in our search engine to allow you to search for URLs like HTML. URLs we just want present movie files to you. So it's important to know and to filter. URLs that are only for movies in that case. Take Parson for example. We may only want to parse things like author or number of pages or you know headings or things like that out of things that are documents, like PDF documents or Word documents. Author, number of pages and things like that have no applicability for example a JPEG image. Image. Okay, so it's really important to know how to parse content in a search engine framework, by its content type. And thus, we have to know the content type to be able to detect that, so being able to detect it is really important. Take indexing, in a search engine. Indexing is really important. We know if it's, a Microsoft Office file, that it has a property in it called number of pages, that we may want to search for later. Well, we only maybe want to store that or store that particular metadata field if it's a Microsoft Office file. We need ability to detect that sort of as well and that's sort of where content detection and analysis comes in for that. So that's sort of the end of the first part of of this mod, module on content detection and analysis. In the next part, we're going to talk about some of the, sort of more information about things like mime types and mime hierarchies, more information about parsing and why that's important, more information about ways and methodologies for doing that. And so we'll head right into that next module or right now. Thanks.