1
23:59:59,500 --> 00:00:05,482
[MUSIC]. 

2
00:00:05,482 --> 00:00:07,706
Alright, so lets go over some of the 
commands in pig. 

3
00:00:07,706 --> 00:00:10,068
So, the first one is load, which is how 
you get data into the system. 

4
00:00:10,068 --> 00:00:15,268
And here the, you know, data is on HDFS 
in a hadoop cluster, and the logical data 

5
00:00:15,268 --> 00:00:20,068
model, the, the each input data is 
assumed to be a bag, which is a sequence 

6
00:00:20,068 --> 00:00:25,793
of tuples, okay. 
So, you can specify, you know, you can 

7
00:00:25,793 --> 00:00:29,185
just say LOAD, but you can also specify 
the function you're going to use to parse 

8
00:00:29,185 --> 00:00:34,371
the data with this keyword USING. 
And this is, seems like sort of just like 

9
00:00:34,371 --> 00:00:38,247
a trivial interpretation, but those are 
kind of an interesting point, and really 

10
00:00:38,247 --> 00:00:43,572
gets to the heart of one of the problems 
with many database products. 

11
00:00:43,572 --> 00:00:47,220
And one of the advantages of these map 
reduced base systems is that sometimes 

12
00:00:47,220 --> 00:00:50,960
you're presented with data that's in its 
raw form. 

13
00:00:50,960 --> 00:00:56,558
And you need to parse it yourself. 
Okay, and so, the database value 

14
00:00:56,558 --> 00:00:59,798
proposition is well look, you know, 
design a schema and load all the data 

15
00:00:59,798 --> 00:01:03,200
into the schema, and after that you'll be 
able to get all these benefits from 

16
00:01:03,200 --> 00:01:06,844
querying. 
Well look, who's going to do that loading 

17
00:01:06,844 --> 00:01:09,408
task, right? 
That's a big parallel job, if I've got 20 

18
00:01:09,408 --> 00:01:14,340
terabytes of text files, somehow I've 
gotta manage that computations. 

19
00:01:14,340 --> 00:01:19,980
And so this is where Hadoop and also 
therefore Pig and Hive come into play. 

20
00:01:19,980 --> 00:01:23,824
And so a lot of times, what you'll see is 
that the application for, a say data 

21
00:01:23,824 --> 00:01:27,606
science application system, you know, 
data science application will be 

22
00:01:27,606 --> 00:01:32,388
architected. 
where Hadoop and its extensions are being 

23
00:01:32,388 --> 00:01:36,294
used to do the initial processing, 
parsing, you know, and loading, and then 

24
00:01:36,294 --> 00:01:40,324
the result of that is loaded into more 
conventional database for kind of ad hoc 

25
00:01:40,324 --> 00:01:45,278
querying. 
And we saw this in the very first 

26
00:01:45,278 --> 00:01:51,740
segment, we talked about the Obama 
campaign using this architecture as well. 

27
00:01:51,740 --> 00:01:53,143
They released an article you can sort of 
infer that they had used a similar 

28
00:01:53,143 --> 00:01:55,695
article. 
They mention Hadoop for the ETL work 

29
00:01:55,695 --> 00:01:59,985
load, extract transform load, another 
piece of jargon, and then they mention 

30
00:01:59,985 --> 00:02:05,275
there's a vertical database for the 
slicing and dicing, okay. 

31
00:02:05,275 --> 00:02:09,367
And you see this, you see this, it's 
almost becoming a, a commit a standard, 

32
00:02:09,367 --> 00:02:12,474
of this. 
I hate to use the word standard, but it's 

33
00:02:12,474 --> 00:02:17,010
the best practice in designing these kind 
of data analysis architecture at scale. 

34
00:02:17,010 --> 00:02:19,386
Okay, so fine. 
So, what's nice about this is that you 

35
00:02:19,386 --> 00:02:22,238
can specify your own parsing function, so 
you can work with data In it's raw 

36
00:02:22,238 --> 00:02:25,531
format, alright. 
And then further you can specify a schema 

37
00:02:25,531 --> 00:02:28,435
if your parsing function is capable of 
producing lots of different things, you 

38
00:02:28,435 --> 00:02:32,230
can say, here are the names of the 
columns that I want to supply. 

39
00:02:32,230 --> 00:02:34,731
So, this is one source of where these 
column is they're going to come from, as 

40
00:02:34,731 --> 00:02:37,430
you specify them right there in the load 
commands. 

41
00:02:37,430 --> 00:02:41,290
So, this is kind of schema on read, if 
you will, right? 

42
00:02:41,290 --> 00:02:43,910
There's no schema associated with the 
data, it's not self describing. 

43
00:02:43,910 --> 00:02:47,848
But you can impose a schema on it as you 
read it into memory. 

44
00:02:47,848 --> 00:02:51,176
Okay, fine. 
And so, maybe this is what we get back, 

45
00:02:51,176 --> 00:02:55,950
we get back a sequence of tuples, and 
we'll use this as a running example. 

46
00:02:55,950 --> 00:02:58,722
Alright, so the next command is Filter, 
which is pretty simple, you just get rid 

47
00:02:58,722 --> 00:03:01,648
of some of these tuples. 
And you can have arbitrary Boolean 

48
00:03:01,648 --> 00:03:04,486
conditions, and you can do kind of 
regular expressions, because everything's 

49
00:03:04,486 --> 00:03:08,710
very text based in that you do [UNKNOWN] 
which is in some cases, a limitation. 

50
00:03:08,710 --> 00:03:13,109
And the syntax looks like this. 
You say, filter some big data set by some 

51
00:03:13,109 --> 00:03:16,540
condition. 
And so, in this case, filter where f1 

52
00:03:16,540 --> 00:03:20,005
equals 8. 
Remember that f1 was the name we ap, we 

53
00:03:20,005 --> 00:03:26,827
gave it in the load command the, the name 
we gave to the very first column. 

54
00:03:26,827 --> 00:03:33,200
And so this finds all the tuples where 
the first position is equal to 8, fine. 

55
00:03:33,200 --> 00:03:36,678
So the command is group which is bringing 
data together. 

56
00:03:36,678 --> 00:03:38,858
[LAUGH] Okay. 
And so, there's a couple different forms 

57
00:03:38,858 --> 00:03:42,152
of this. 
but we're going to focus on, so, just, 

58
00:03:42,152 --> 00:03:47,170
just right now, we're just talking about 
group, okay. 

59
00:03:47,170 --> 00:03:49,671
There's a code group command that I'm 
going to talk about in, in a moment, 

60
00:03:49,671 --> 00:03:54,741
alright. 
So, group says, group this large data 

61
00:03:54,741 --> 00:04:02,932
said by some sequence of columns, okay. 
So, this looks a lot like the SQL clause, 

62
00:04:02,932 --> 00:04:07,242
that has just the same, they got the same 
sort of flavor. 

63
00:04:07,242 --> 00:04:11,950
But it does something pretty different, 
okay. 

64
00:04:11,950 --> 00:04:19,044
So gr, if this is A and we group by a f1, 
well, what we get out is tuples. 

65
00:04:19,044 --> 00:04:23,930
But now the first, column is the same, 
because we grouped by f1. 

66
00:04:23,930 --> 00:04:27,470
But the second column, the second field I 
should say, is now a bag. 

67
00:04:27,470 --> 00:04:31,890
A little group representation, of all the 
tuples that were associated with this, 

68
00:04:31,890 --> 00:04:36,090
with this particular key. 
And so, if you look over here there's 

69
00:04:36,090 --> 00:04:40,228
only one tuple with, with one, and that 
tuple appears in a group. 

70
00:04:40,228 --> 00:04:45,808
There's two tuples with f1 equal to 4, 
and so this bag has two tuples in it and 

71
00:04:45,808 --> 00:04:49,940
so on, okay. 
So, this is how you and, you can start 

72
00:04:49,940 --> 00:04:52,865
off with something that's kind of a flat 
structure and you can build up a nested 

73
00:04:52,865 --> 00:04:56,700
structure for various reasons. 
And we'll see why. 

74
00:04:56,700 --> 00:05:00,426
Now the other thing to keep in mind, and 
this is another thing I don't really love 

75
00:05:00,426 --> 00:05:03,367
about pig. 
Because of doing things sort of 

76
00:05:03,367 --> 00:05:06,985
implicitly rather than explicitly, is 
that the name of this field, every field, 

77
00:05:06,985 --> 00:05:11,411
you end up having to have a name. 
The name of this field is defaulted to 

78
00:05:11,411 --> 00:05:15,063
the, to the name group. 
yeah, sorry, I just lied. 

79
00:05:15,063 --> 00:05:19,430
The first field is defaulted to the name 
group. 

80
00:05:19,430 --> 00:05:24,640
And the second field is defaulted to the 
name of the original data set. 

81
00:05:24,640 --> 00:05:27,718
And so, I find this a lil, a little 
confusing, when you're writing pig 

82
00:05:27,718 --> 00:05:31,070
yourself, but you can sort of see why 
they did it. 

83
00:05:31,070 --> 00:05:34,970
The reason is is that just these bags 
altogether, actually have the same 

84
00:05:34,970 --> 00:05:40,200
information as the original data set, 
just nested in a certain way. 

85
00:05:40,200 --> 00:05:43,190
And so, calling it A sort of makes sense. 
It's got a listing of information. 

86
00:05:43,190 --> 00:05:46,476
So for example, notice that, notice that 
the, the value of 1 is now repeated 

87
00:05:46,476 --> 00:05:50,173
twice, once in the group and it's still 
in the tuple. 

88
00:05:50,173 --> 00:05:54,234
Okay, so fine. 
So those are command DISTINCT, that does 

89
00:05:54,234 --> 00:05:58,700
just what you might imagine, it gets rid 
of all duplicates. 

90
00:05:58,700 --> 00:06:01,934
And, just for a simple example here, if 
you've got two different elements in the 

91
00:06:01,934 --> 00:06:06,280
bag that have the same value, you'll, 
output will be, will be this, okay. 

92
00:06:06,280 --> 00:06:15,897
I might make a claim that DISTINCT A is 
equal to Grouping of A by all three 

93
00:06:15,897 --> 00:06:23,078
columns at once. 
So first of all, why am I making that 

94
00:06:23,078 --> 00:06:27,363
claim? 
Well, remember that the group operator 

95
00:06:27,363 --> 00:06:35,214
puts out a single tuple for every unique 
value of the grouping columns. 

96
00:06:35,214 --> 00:06:39,681
So, that sounds about right. 
and you can also maybe think from, from, 

97
00:06:39,681 --> 00:06:44,230
if you, if you know sequel that this is 
sort of true there, right. 

98
00:06:44,230 --> 00:06:47,516
You can use this distinct keyword in, in 
SQL and you can also group by all the 

99
00:06:47,516 --> 00:06:51,240
columns, and you'll end up getting the 
same result. 

100
00:06:51,240 --> 00:06:54,550
So, are these two expressions the same? 
Do they produce the same output? 

101
00:06:54,550 --> 00:06:57,740
Well, not quite, because the grouping 
structure. 

102
00:06:57,740 --> 00:07:01,687
So, the group command produces a group 
field, which is now the whole tuple. 

103
00:07:01,687 --> 00:07:06,540
Excuse me [SOUND]. 
And the, the A field, which is a bag with 

104
00:07:06,540 --> 00:07:11,952
a single tuple in it. 
So, you get this sort of repetition of 

105
00:07:11,952 --> 00:07:16,120
information, so this distinct is much 
more concise. 

106
00:07:16,120 --> 00:07:19,353
So, you need to be careful, and sort of 
make sure you understand what these 

107
00:07:19,353 --> 00:07:22,480
things are going to produce. 
So, how does group work? 

108
00:07:22,480 --> 00:07:27,088
Well, we already saw, you know, going 
back to our map reduce schematic, or our 

109
00:07:27,088 --> 00:07:33,047
parallel processing schematic. 
We break the data into pieces,we apply a 

110
00:07:33,047 --> 00:07:37,857
map function that assigns each tuple to 
its key, and groups the tuple to, in the 

111
00:07:37,857 --> 00:07:42,772
value as well. 
And so here if we group by f1, then f1 

112
00:07:42,772 --> 00:07:49,160
becomes the key, and the value is all, is 
all three of the elements and the tuple. 

113
00:07:49,160 --> 00:07:55,324
And those are shuffled across the network 
to the reduce side, and the reduce side 

114
00:07:55,324 --> 00:08:02,274
will construct this bag type out of the 
set of tuples, okay. 

115
00:08:02,274 --> 00:08:06,430
So, this is a single map reduce job, 
alright. 

116
00:08:06,430 --> 00:08:10,220
So the for each command is almost 
certainly the most complex one. 

117
00:08:10,220 --> 00:08:14,465
So, here you're basically going to 
manipulate each tuple in a bag. 

118
00:08:14,465 --> 00:08:19,370
So, you write it like this. 
You'll say FOREACH A GENERATE something, 

119
00:08:19,370 --> 00:08:24,220
and here we're generating a tuple with 
two fields. 

120
00:08:24,220 --> 00:08:28,440
One is f0, and one is the sum of f1 and 
f2. 

121
00:08:28,440 --> 00:08:30,582
And here you can call user-defined 
functions or you can write other kinds of 

122
00:08:30,582 --> 00:08:35,998
arithmetic expressions. 
You can sort of do lots of things, okay? 

123
00:08:35,998 --> 00:08:42,235
So another example is, first we group A 
by f1, just like we did in the previous 

124
00:08:42,235 --> 00:08:48,879
slide. 
excuse me, this should be y. 

125
00:08:56,390 --> 00:09:01,815
And then for each Y, generate group. 
So what's group? 

126
00:09:01,815 --> 00:09:08,730
Remember that's the magic assign, given 
to the group field. 

127
00:09:08,730 --> 00:09:14,772
And then Y, which is the magic name given 
to the bag field, dot a list of 

128
00:09:14,772 --> 00:09:20,841
projection columns. 
And so this is the second element and the 

129
00:09:20,841 --> 00:09:25,270
third element from each of those tuples. 
Okay, so what does this look like? 

130
00:09:25,270 --> 00:09:32,340
Well, X which is this first element here, 
well we get the first field from A, which 

131
00:09:32,340 --> 00:09:37,835
is this column here. 
And then we get the sum of the second 

132
00:09:37,835 --> 00:09:40,090
two. 
So, 2 plus 3 is equal to 5, 2 plus 1 is 

133
00:09:40,090 --> 00:09:43,325
equal to 3, 3 plus 4 is equal to 7 and so 
on. 

134
00:09:43,325 --> 00:09:47,095
Okay, and down here for Z, remember that 
Y looks just like it did in the previous 

135
00:09:47,095 --> 00:09:51,128
slide. 
Here, up, wait oh no, sorry I didn't 

136
00:09:51,128 --> 00:09:57,790
actually do, I didn't actually do the 
complete example of Y. 

137
00:09:57,790 --> 00:10:00,100
So, Y is all of A grouped by the first 
field. 

138
00:10:00,100 --> 00:10:04,750
but then we're going to generate the 
grouping field, which is the same. 

139
00:10:07,920 --> 00:10:11,950
And we're going to project out the second 
two columns, we're going to ignore the 

140
00:10:11,950 --> 00:10:16,596
first column in each one of the bags. 
Okay, and so the point here, being that 

141
00:10:16,596 --> 00:10:21,060
you can manipulate these nested objects 
by writing these kinds of expressions. 

142
00:10:21,060 --> 00:10:24,718
But I think the other, sort of, lurking 
point here, is that it's a little bit 

143
00:10:24,718 --> 00:10:28,789
complicated to think about what's going 
on, because of this extra flexibility you 

144
00:10:28,789 --> 00:10:34,234
get with a nested data model, okay. 
So, you'll get a chan, in the assignment, 

145
00:10:34,234 --> 00:10:37,086
you'll get a chance to try some of these 
out, and make sure going to understand 

146
00:10:37,086 --> 00:10:41,545
what they're doing. 
Okay, so then there's this key word 

147
00:10:41,545 --> 00:10:44,822
FLATTEN. 
And it's not really it's own operator, 

148
00:10:44,822 --> 00:10:49,280
you use it in the context of for each. 
And so here, you know, because of the 

149
00:10:49,280 --> 00:10:53,824
complexity here, what you might really 
want is, look, I just wanted to get rid 

150
00:10:53,824 --> 00:10:58,290
of the. 
If you want to recover the flat version 

151
00:10:58,290 --> 00:11:03,644
of the message structures created and 
stored in the variable Z. 

152
00:11:03,644 --> 00:11:08,468
Or you can do something like this FOREACH 
X GENERATE group and then the FLATTEN of 

153
00:11:08,468 --> 00:11:12,392
X. 
And the result you'll get out here is but 

154
00:11:12,392 --> 00:11:16,928
regardless, what it does is sort of peel 
out this bag and produced some extra 

155
00:11:16,928 --> 00:11:21,976
tuple for each. 
And I don't really like this, because 

156
00:11:21,976 --> 00:11:26,461
it's sort of changes, just using the 
keyword, FLATTEN, changes the semantics 

157
00:11:26,461 --> 00:11:33,340
of the FOREACH generate and I think it's 
a very, very confusing way of doing it. 

158
00:11:33,340 --> 00:11:36,390
So the idea, it's hard to explain the 
principle going on here, just sort of 

159
00:11:36,390 --> 00:11:41,070
memorized what it does, practice with it 
and memorized what it does, alright? 

