[MUSIC]. Alright, so lets go over some of the commands in pig. So, the first one is load, which is how you get data into the system. And here the, you know, data is on HDFS in a hadoop cluster, and the logical data model, the, the each input data is assumed to be a bag, which is a sequence of tuples, okay. So, you can specify, you know, you can just say LOAD, but you can also specify the function you're going to use to parse the data with this keyword USING. And this is, seems like sort of just like a trivial interpretation, but those are kind of an interesting point, and really gets to the heart of one of the problems with many database products. And one of the advantages of these map reduced base systems is that sometimes you're presented with data that's in its raw form. And you need to parse it yourself. Okay, and so, the database value proposition is well look, you know, design a schema and load all the data into the schema, and after that you'll be able to get all these benefits from querying. Well look, who's going to do that loading task, right? That's a big parallel job, if I've got 20 terabytes of text files, somehow I've gotta manage that computations. And so this is where Hadoop and also therefore Pig and Hive come into play. And so a lot of times, what you'll see is that the application for, a say data science application system, you know, data science application will be architected. where Hadoop and its extensions are being used to do the initial processing, parsing, you know, and loading, and then the result of that is loaded into more conventional database for kind of ad hoc querying. And we saw this in the very first segment, we talked about the Obama campaign using this architecture as well. They released an article you can sort of infer that they had used a similar article. They mention Hadoop for the ETL work load, extract transform load, another piece of jargon, and then they mention there's a vertical database for the slicing and dicing, okay. And you see this, you see this, it's almost becoming a, a commit a standard, of this. I hate to use the word standard, but it's the best practice in designing these kind of data analysis architecture at scale. Okay, so fine. So, what's nice about this is that you can specify your own parsing function, so you can work with data In it's raw format, alright. And then further you can specify a schema if your parsing function is capable of producing lots of different things, you can say, here are the names of the columns that I want to supply. So, this is one source of where these column is they're going to come from, as you specify them right there in the load commands. So, this is kind of schema on read, if you will, right? There's no schema associated with the data, it's not self describing. But you can impose a schema on it as you read it into memory. Okay, fine. And so, maybe this is what we get back, we get back a sequence of tuples, and we'll use this as a running example. Alright, so the next command is Filter, which is pretty simple, you just get rid of some of these tuples. And you can have arbitrary Boolean conditions, and you can do kind of regular expressions, because everything's very text based in that you do [UNKNOWN] which is in some cases, a limitation. And the syntax looks like this. You say, filter some big data set by some condition. And so, in this case, filter where f1 equals 8. Remember that f1 was the name we ap, we gave it in the load command the, the name we gave to the very first column. And so this finds all the tuples where the first position is equal to 8, fine. So the command is group which is bringing data together. [LAUGH] Okay. And so, there's a couple different forms of this. but we're going to focus on, so, just, just right now, we're just talking about group, okay. There's a code group command that I'm going to talk about in, in a moment, alright. So, group says, group this large data said by some sequence of columns, okay. So, this looks a lot like the SQL clause, that has just the same, they got the same sort of flavor. But it does something pretty different, okay. So gr, if this is A and we group by a f1, well, what we get out is tuples. But now the first, column is the same, because we grouped by f1. But the second column, the second field I should say, is now a bag. A little group representation, of all the tuples that were associated with this, with this particular key. And so, if you look over here there's only one tuple with, with one, and that tuple appears in a group. There's two tuples with f1 equal to 4, and so this bag has two tuples in it and so on, okay. So, this is how you and, you can start off with something that's kind of a flat structure and you can build up a nested structure for various reasons. And we'll see why. Now the other thing to keep in mind, and this is another thing I don't really love about pig. Because of doing things sort of implicitly rather than explicitly, is that the name of this field, every field, you end up having to have a name. The name of this field is defaulted to the, to the name group. yeah, sorry, I just lied. The first field is defaulted to the name group. And the second field is defaulted to the name of the original data set. And so, I find this a lil, a little confusing, when you're writing pig yourself, but you can sort of see why they did it. The reason is is that just these bags altogether, actually have the same information as the original data set, just nested in a certain way. And so, calling it A sort of makes sense. It's got a listing of information. So for example, notice that, notice that the, the value of 1 is now repeated twice, once in the group and it's still in the tuple. Okay, so fine. So those are command DISTINCT, that does just what you might imagine, it gets rid of all duplicates. And, just for a simple example here, if you've got two different elements in the bag that have the same value, you'll, output will be, will be this, okay. I might make a claim that DISTINCT A is equal to Grouping of A by all three columns at once. So first of all, why am I making that claim? Well, remember that the group operator puts out a single tuple for every unique value of the grouping columns. So, that sounds about right. and you can also maybe think from, from, if you, if you know sequel that this is sort of true there, right. You can use this distinct keyword in, in SQL and you can also group by all the columns, and you'll end up getting the same result. So, are these two expressions the same? Do they produce the same output? Well, not quite, because the grouping structure. So, the group command produces a group field, which is now the whole tuple. Excuse me [SOUND]. And the, the A field, which is a bag with a single tuple in it. So, you get this sort of repetition of information, so this distinct is much more concise. So, you need to be careful, and sort of make sure you understand what these things are going to produce. So, how does group work? Well, we already saw, you know, going back to our map reduce schematic, or our parallel processing schematic. We break the data into pieces,we apply a map function that assigns each tuple to its key, and groups the tuple to, in the value as well. And so here if we group by f1, then f1 becomes the key, and the value is all, is all three of the elements and the tuple. And those are shuffled across the network to the reduce side, and the reduce side will construct this bag type out of the set of tuples, okay. So, this is a single map reduce job, alright. So the for each command is almost certainly the most complex one. So, here you're basically going to manipulate each tuple in a bag. So, you write it like this. You'll say FOREACH A GENERATE something, and here we're generating a tuple with two fields. One is f0, and one is the sum of f1 and f2. And here you can call user-defined functions or you can write other kinds of arithmetic expressions. You can sort of do lots of things, okay? So another example is, first we group A by f1, just like we did in the previous slide. excuse me, this should be y. And then for each Y, generate group. So what's group? Remember that's the magic assign, given to the group field. And then Y, which is the magic name given to the bag field, dot a list of projection columns. And so this is the second element and the third element from each of those tuples. Okay, so what does this look like? Well, X which is this first element here, well we get the first field from A, which is this column here. And then we get the sum of the second two. So, 2 plus 3 is equal to 5, 2 plus 1 is equal to 3, 3 plus 4 is equal to 7 and so on. Okay, and down here for Z, remember that Y looks just like it did in the previous slide. Here, up, wait oh no, sorry I didn't actually do, I didn't actually do the complete example of Y. So, Y is all of A grouped by the first field. but then we're going to generate the grouping field, which is the same. And we're going to project out the second two columns, we're going to ignore the first column in each one of the bags. Okay, and so the point here, being that you can manipulate these nested objects by writing these kinds of expressions. But I think the other, sort of, lurking point here, is that it's a little bit complicated to think about what's going on, because of this extra flexibility you get with a nested data model, okay. So, you'll get a chan, in the assignment, you'll get a chance to try some of these out, and make sure going to understand what they're doing. Okay, so then there's this key word FLATTEN. And it's not really it's own operator, you use it in the context of for each. And so here, you know, because of the complexity here, what you might really want is, look, I just wanted to get rid of the. If you want to recover the flat version of the message structures created and stored in the variable Z. Or you can do something like this FOREACH X GENERATE group and then the FLATTEN of X. And the result you'll get out here is but regardless, what it does is sort of peel out this bag and produced some extra tuple for each. And I don't really like this, because it's sort of changes, just using the keyword, FLATTEN, changes the semantics of the FOREACH generate and I think it's a very, very confusing way of doing it. So the idea, it's hard to explain the principle going on here, just sort of memorized what it does, practice with it and memorized what it does, alright?