[MUSIC]. So, let's consider an example of a paper gram/g. And see how expressing the program in terms of these high-level operations, offers optimization opportunities that the system can exercise unilaterally without the programmer having to specify it. Okay. So, in this example, we're looking at traffic web, web log traffic data with sort of three columns [SOUND] here. The IP address, the time and the URL being accessed. Okay, so in the first step, we load the data and here we don't need to use the using clause to specify our own parsing function because presumably the traffic.dat is in some format that they already naively know how to parse. So in some sort of limited format, okay, so it's expecting to find three columns. and if it doesn't, it'll be an error. Alright. So the second command is GROUP. So we group A by IP address. And then third we say, well, for each of those grouped tuples, I just want the IP address. And then I want the number of web log entries associated with that IP address. Okay. So this is counting the number of accesses by a particular IP address out there on the internet coming in. Okay? And the next step, well, we, maybe in this particular task we're only interested in sort of looking at traffic that originated from two particular gateways on our local network. Work and so we fill the [INAUDIBLE] with the filter command. And then we store the result into a file [INAUDIBLE] and the observation here is that you know, if you look at this carefully you can see that as written this is some what inefficient, right. We load a very large data set. We apply a grouping operation to a very large data set. We do some manipulation of those groups across the very large data set. And then we filter it down to something that's much smaller but they're only [INAUDIBLE] these two IP addresses. So, what would be nice is if we could apply the filter first. And so, one thing you could do is just rewrite this program. And in this case that's probably reasonable, you probably notice that this is going on and rewrite it yourself manually. But in a more complicated program you may or may not and this is not the only example of an optimization that can be done, okay? And so what pig will do is actually move the filter automatically before hand, before the group. And use, I guess, I guess I didn't modify the program here. I modified the out, the abstract plan. It won't actually, it wont' actually produce the text here, the ASCII is done. Remember, remember we talked about an execution plan, the minute will manipulate that. And so it moves the filter beforehand because it know it's safe to do so. And this is again, I keep coming back to this but this is the, you know, you know, one advantage of a high level language might be that it's easier to express your program in. But that may or may not be true in many cases it could be actually harder to express your program in terms of these, you know, limited operators. so you know, you might prefer kind of general, a Java program might prefer a full-width Java interface where they can just write whatever code they want. But the point is that if you can, you know, tie one arm behind your back and express your task in terms of these operations. The system can take over and make manipulations that it would be very difficult for it to do if you wrote in a lower level language, okay. So, the other point to make here is that remember this is what we call lazy evaluation. So th, a, at, when this command is, is executed, that's when it makes an op, that's when it has an opportunity to apply these optimizations. Right, so we did nothing but sort of build up your, your execution plan as you were executing these commands. It didn't actually, it didn't actually do any work. And then finally when you say okay, now I really want the result, it can say aha I see, I see your list of sequences. Or sorry, I see your sequence of commands, I'm going to do some manipulation of those and figure out the right way to execute these, these commands before I produce the result. Okay. Fine, so were not done, we now have to map this abstract execute command down into a sequence of map produce jobs. And so the way it does this is it first identifies all the group and CoGroup operators in the plan and assigns a, an individual MapReduce job to each one of those operators. And then incrementally what it's going to try to do is walk forward and backward and put as much work as it can into the same MapReduce job, okay. So, the Group command needs to be it's own MapReduce job because it's going to shuffle things across the network. But filtering, I could apply, I could apply the filtering condition as a [INAUDIBLE] data of a disc, or as I process a tuple I can check its condition to make sure it's from one of these IP addresses before I actually apply the group. So I might as well do that in the same map reduced job, I don't need a whole independent map produced job I will just do that. Further the load command is just reading things off disk it doesn't actually do anything difficult at all I can put that on the same map reduce job as well. And so it ends up with this one map function that's going to look something like you know as you process a value its going to say you know pause the data. And then just going to say if bow dot IP equals this or a bow IP equals this. And emit IP val or val IP. Val, right so this little program essentially gets generated by the compiler as one as just one map file function. And then gets executed. And on the reduce side you can do a couple things. You can crih, construct the groups but you don't actually care about the groups because you're immediately going to count the results. And so it's smart enough to see that group plus foreach is really going to just produce tuples like this: IP and Count. And not, you know, not IP plus some big group. And that's actually pretty significant savings, because constructing these objects and sort of passing them around is, is, is pretty expensive. So you don't actually need them you don't want to use them. So there's sort of two levels of optimization. One is it moves operators around, which I like, y'know because they get this algebraic flavor. And two it compresses logical operations into single physical operations, and the overall name of the game here is to use to reduce the total number of map reduce jobs being executed because they are expensive. Okay. And then finally, you can write the things out and that can also happen in the same reduce space. So this entire task ends up being just one map reduce job Job. In other cases you may need need to just do sort of multiple groupings or maybe just short or so on. And each one of those requires its own map reduce job. So for example, certain commands always require a map reduce job like sorting. Okay. So what did we talk about in the last several segments? Well we described no sequel systems. And argue that they're important for a data scientist to understand, in part because you may be using them but also because you may be asked to weigh in on their relative strengths and weaknesses compared to other systems. And so we talk about NoSQL sort of meaning no schema, and no transactions, and no language. Language and may be less about specifically bad SQL and overall they're kind of a reboot of data systems zeroing in on just high-throughput reads and writes. Okay. Right, and so now the design space of these large scale data sets is sort of being more fully exploit with different permutations and combinations of particular features. But over all there's a clear trend back towards re-introducing schemas and transactions and languages and so on, so we talked about Google spanner system as an example of this trend. So, no SQL is not, it's an evolving of concept, the main thing to realize is the entire space of possibilities in large scale databases is being explored and no sequel represents one segment relational database is representing another. And there's new segments emerging so you may hear the term new sequal for example which is kind of a architecture database that does try to achieve some of these a other features okay. So fine then we talked about pig in particular as an example not so much as another sequel. Full system but as a, a analytic system that puts a layer on top of MapReduce. And we chose pig because it has this relational algebra like layer on top of Hadoop and so I want to make the point that relational algebra comes up in a lot of different context. You know, and in, in this one it's sort of a little bit interesting because although it has a clear relational algebra flavor, it's not actually a pure relational data model. You got this sort of. These structures. And the point here is that you can kind of cherry-pick some of these concepts and techniques pioneered by databases. And you don't actually have to use them in the context of databases. So you don't need to sort of throw the baby out with the bath water if a database doesn't appear to be meeting your needs and these systems are sort of doing that. So, as you review different systems for merging, as a data scientist, you can start to understand the particular set of feature they offer. And understand the problems with those features and the history there. And the pros and cons therein. Okay. And so Pig, also. The, you know, one take away of it is that it has this sort of schema-on-read fl, flavor rather than schema-on-write, meaning that you can work with in situ data. Okay, I've mentioned this a couple of times. I just wanted to point out that that's really one of the new requirements associated with NoSQL as well as SQL [UNKNOWN] is that you have to be able to work with data in it's kind of native format. And the reason for that is it's just to big to transform all over the place.