[MUSIC]. Okay so, I want to talk a little bit about the system Pig, and there's a couple reasons for this. One is it represents this language layer on top of map produce, which I think is pretty important as I've, as I've mentioned before. And it also represents a more direct interface to relational algebra. Which I've argued is really the important abstraction from databases. SQL happens to be sort-of intergalactic data-speak, and it is you know the universal language of programming databases and manipulating databases. But really, the secret sauce, the technology that makes it effective, is this underlying formulas and relation algebra. And Pig and a few other systems sort of exposes as a, as, as a first class citizen. And it, it, it's not necessarily the only way to do it or the better way to do it, but since you can see relational algebra I think it's a nice thing to be familiar with, okay. and also I think makes the point that you know, you, you can cherry pick a little bit what features you bring. So, while it's very obviously inspired by relational algebra, and of course if you look in the sort of people that developed it, you can see why. It definitely changes some things about the relational model and I, and I think that's an important point too, is that we can use this relation algebra of abstraction independently of some of the systems that it came from. We can make whatever changes we need and keep, you know keep what we like and throw out what we don't like. And so I want, you know I want you to sort of put your relation algebra goggles on and see the world in terms of relation algebra. I think it's pretty effective for working with not just data in general, but especially big data and reason they are very scalable algorithms. Okay, so fine, what is Pig? So, Pig is a engine for executing programs on top of Hadoop. Okay, so when you are going to write a pig program in Pig, and it's going to generate a sequence of map produced jobs, that implement that program, okay. And the language here is called Pig Latin, and it's an Apache open sourced project and you can read more about it. Okay, and so, why bother doing this? As well, if you're coming from, if you're a map produce program, or if you're happy with map produce, why would you bother doing this? Well, suppose you have data in one file and data from websites, sorry, user data in one file and website data in another, and you need find the top most visited sites by user in a particular age range. And so if you're thinking in terms of SQL, SQL you can probably think about how you express this task as a query. and here's a pictorial, rep, representation of it. That sort of shows a, a kind of data flow here. We're going to load users, we'll load the pages, we'll filter by age, do some kind of join on name, group, count up the clicks, order by clicks, and then take the top five. This may or may not be the way you draw it on the whiteboard if you were to express it yourself, but it's a, but it's a reasonable way of describing the problem. Well in MapReduce there's 170 lines of code, and at least in one example coming from the papers here, it took someone four hours to write it. These kind of metrics of how long development time are not never too compelling, because it depends on their background and skill set, and so on. Still, a decent amount of time, while the same program is, in Pig Latin is just nine lines of code and takes, you know, at least in this case, 15 minutes to write. And further, you know, it does sort of capture some of this more abstract description of what's going on. Now, this is a little bit contrived, because these boxes obviously correspond to sort of particular pig commands. But I'd argue that this, that this even non, people that don't know Pig, if they were to ask sort of what the tasks are in eval, in evaluating this English question, they might come up with something along these lines. They may or may not say join, but some of these steps would be, would be present. So, I think it, you can wave your hands a little bit and say that it, that it captures a little bit of the natural language description of the task, okay. You may or may not buy that. Fine, so here is what it looks like, load the user data according to a particular schema, filter the user database on the age range, load the pages data, perform a join group. You say, well, for each group, count the number of clicks and we'll talk. This one's probably look the least like relation algebra. And we'll talk about it. And then sort the thing, and then just take the top five, and store that result into a new file, okay. And so, the point here is that you could also think, well boy I could just write all this in python, right. That would be much easier, why do I need to use this new language. Well the point that each one of these steps well not actually not each one. Groups of these steps correspond to individual map reduce jobs, which mean they scale really, really well, right. Does, doesn't matter how big your user's table is, and it doesn't matter how big your page's table is this program will work. Okay, so how does the system work? Well the programmer's going to enter a Pig Latin program by writing these commands [SOUND]. and assigning the results to variables. And then you can refer to the variables in future commands. So this is, you know, we've, we've loaded data into this variable A, loaded it into variable B, and then you can filter by reference to A then sort of the result and C, and so on, okay. And then, we'll make this point again later, but nothing actually happens. No work is actually done until you try to actually write out the results. So, this is sort of what we call lazy, lazy evaluation. Alright, so fine. So, what happens when you write this program? Well the Pig parser will turn the syntax of the program into an abstract representation of the, program called an execution plan. And so this is using parlance from databases, and the lead of the Pig project is a, is a great database guy named Chris Olson, okay. So, this abstract presentation of the plan, you know, these operators, these operations that I'm going to call operators your going to use in database [UNKNOWN] are sort of one-to-one with the commands. Although they need not always be. Okay, but we're not done there. That's just the execution plan, it's still sort of abstract we can't actually execute that, what we need to do is compile that down into map reduced jobs. And the game here is going to sort of be to minimize the number of map reduced jobs you need, because there's a lot of overhead to executing one of these. So, you don't want to run a whole map reduced job, just to filter a data set. You might as well lump that in with other work that you're already doing. As long as you're scanning the data and reading off the disk, you might as well do as much as you can with it. And so for example, in this case, this program can all be implemented as just one single map reduce job. One, in the map phase you load the data, you filter it, and you also load the other data set, and the reduce phase you do the join, okay. And we'll see another example of this, a little later on. And then finally, this, these map reduce jobs are scheduled on a Hadoop cluster as usual, end run as usual, okay. So, these performance [LAUGH] results are, a, a little funny but I sort of like, I sort of, I, I still like the argument to be made. This is, Pig Performance versus Map-Reduce. Okay, so this is a little funny, because Pig is built on top of Map-Reduce. But the point is, is that the first version of Pig in September 11th, 2008, was a lot slower that map, than just writing the thing raw against Map-Reduce. But with some various improvements it got better and better over time, until eventually it was just as fast as the handwritten Map-Reduce. And I think you see this sort of pattern quite a bit. that, that there is a cost to abstraction. But that you can, you typically can recover a lot of the performance of the hand coded stuff. Meanwhile having gained some measure of program productivity by offering a high level interface, okay. And so I just like the fact that they are actually doesn't bothered telling the story, admitting they were quite slow in the beginning and they got faster over time, okay. So, what's the data model here, so there is four types involved here. One is the atom, which is just a, a primitive, right, an integer, a string and so on. And then there's three different collection types. One is a tuple, and these are, should be familiar from our first assignment, we worked with Python. Alright, so there's, these three all exist in, well I shouldn't say that actually, bag is a little funny thing in terms of Python. Lemme just say, the first is a tuple which is a sequence of fields. And every field can be of any type. They need not be atoms, okay. And then there's a bag which is a collection of tuples, always tuples. But those tuples need not be the same type. This is different than a table, different than a relation, okay. And it's a bag not a set. And, you know, I guess I'll, I can ask you what's the difference between a bag and a set? Well, a bag allows duplicates, okay. And then the third type is a map, which is a little bit confusing because we're talking about Map Reduce, but this is a dictionary in Python, right? So, string literal keys map to any other type, okay. So, let's look at an example of this. So, what is this thing? Well depends on my notation I guess, but assuming that you don't mind my angle brackets referring to tuple. Then you can see that this outer thing is a tuple, and it has three fields, it has a integer which is just an atom, and then it has this thing which is a bag, the curly braces I'm going to use to indicate bag. And then it's got this thing, whoop excuse me, which is a map, a dictionary that has just one element in it. Mapping the key apache to the value, search, okay. So, let's name these guys f1, f2 and f3, and we'll point out in the, in a couple, in a couple slides where those names come from. There's a few different places they can come from, and I'll point out one of those sources soon. But right now, let's just assume that we have them, okay. Okay, so let's consider these expressions over this type. So, we can write $0, and what that means is give me the first field in the tuple. All right, so, in this case, it's just the number 1. We can also access it by name, so if we access f2 that will give us the second field, just because we gave it that explicit name. There's nothing magic about f, nothing magic about 2. And that'll give us a bag with these values in it. By the way we should say what are the elements of this bag? Well, they're tuples with two integers each. Okay, now you can also access f2 to give me, so what does this expression do? Well, f2 gives me the bag, and then you can write .$0 and I don't love this notation, it's a bit of an abusive notation. But what it means is for every element of the bag, project out the first element. And so, this will give me another bag, but it only has 2 and 4 and 5. It has the zeroth element of every tuple, okay. Then you've got this magic hash symbol here which means finding the value associated with the key I'm about to provide. So, look at f3 which is the, the third field. And hash into it, and give me the key apache. If it's not a map, you're going to get an error, but in this case it is a map, and so what gets returned is the atom search, okay. And you can also write functions, you can say sum up all of these values, and this expression is the same one we saw here, which will be a sequence of integers. And so it gives you 2 plus 4 plus 5, and there's a few different other ways to manipulate these things. But the first thing to notice is this is non-relational, right? You've got these nested data structures, and you got a few different data types.