1
23:59:59,500 --> 00:00:05,344
[MUSIC]. 

2
00:00:05,344 --> 00:00:07,734
So, let's consider an example of a paper 
gram/g. 

3
00:00:07,734 --> 00:00:11,702
And see how expressing the program in 
terms of these high-level operations, 

4
00:00:11,702 --> 00:00:16,042
offers optimization opportunities that 
the system can exercise unilaterally 

5
00:00:16,042 --> 00:00:20,280
without the programmer having to specify 
it. 

6
00:00:20,280 --> 00:00:22,106
Okay. 
So, in this example, we're looking at 

7
00:00:22,106 --> 00:00:26,778
traffic web, web log traffic data with 
sort of three columns [SOUND] here. 

8
00:00:26,778 --> 00:00:32,086
The IP address, the time and the URL 
being accessed. 

9
00:00:32,086 --> 00:00:35,319
Okay, so in the first step, we load the 
data and here we don't need to use the 

10
00:00:35,319 --> 00:00:38,552
using clause to specify our own parsing 
function because presumably the 

11
00:00:38,552 --> 00:00:44,117
traffic.dat is in some format that they 
already naively know how to parse. 

12
00:00:44,117 --> 00:00:48,270
So in some sort of limited format, okay, 
so it's expecting to find three columns. 

13
00:00:48,270 --> 00:00:50,064
and if it doesn't, it'll be an error. 
Alright. 

14
00:00:50,064 --> 00:00:52,180
So the second command is GROUP. 
So we group A by IP address. 

15
00:00:52,180 --> 00:00:58,420
And then third we say, well, for each of 
those grouped tuples, I just want the IP 

16
00:00:58,420 --> 00:01:03,964
address. 
And then I want the number of web log 

17
00:01:03,964 --> 00:01:10,440
entries associated with that IP address. 
Okay. 

18
00:01:10,440 --> 00:01:15,000
So this is counting the number of 
accesses by a particular IP address out 

19
00:01:15,000 --> 00:01:18,944
there on the internet coming in. 
Okay? 

20
00:01:18,944 --> 00:01:21,623
And the next step, well, we, maybe in 
this particular task we're only 

21
00:01:21,623 --> 00:01:24,631
interested in sort of looking at traffic 
that originated from two particular 

22
00:01:24,631 --> 00:01:29,133
gateways on our local network. 
Work and so we fill the [INAUDIBLE] with 

23
00:01:29,133 --> 00:01:33,278
the filter command. 
And then we store the result into a file 

24
00:01:33,278 --> 00:01:38,110
[INAUDIBLE] and the observation here is 
that you know, if you look at this 

25
00:01:38,110 --> 00:01:46,010
carefully you can see that as written 
this is some what inefficient, right. 

26
00:01:46,010 --> 00:01:50,548
We load a very large data set. 
We apply a grouping operation to a very 

27
00:01:50,548 --> 00:01:53,623
large data set. 
We do some manipulation of those groups 

28
00:01:53,623 --> 00:01:56,480
across the very large data set. 
And then we filter it down to something 

29
00:01:56,480 --> 00:01:59,693
that's much smaller but they're only 
[INAUDIBLE] these two IP addresses. 

30
00:01:59,693 --> 00:02:04,410
So, what would be nice is if we could 
apply the filter first. 

31
00:02:04,410 --> 00:02:07,240
And so, one thing you could do is just 
rewrite this program. 

32
00:02:07,240 --> 00:02:10,516
And in this case that's probably 
reasonable, you probably notice that this 

33
00:02:10,516 --> 00:02:13,545
is going on and rewrite it yourself 
manually. 

34
00:02:13,545 --> 00:02:16,717
But in a more complicated program you may 
or may not and this is not the only 

35
00:02:16,717 --> 00:02:20,350
example of an optimization that can be 
done, okay? 

36
00:02:20,350 --> 00:02:25,030
And so what pig will do is actually move 
the filter automatically before hand, 

37
00:02:25,030 --> 00:02:28,039
before the group. 
And use, I guess, I guess I didn't modify 

38
00:02:28,039 --> 00:02:29,560
the program here. 
I modified the out, the abstract plan. 

39
00:02:29,560 --> 00:02:31,420
It won't actually, it wont' actually 
produce the text here, the ASCII is done. 

40
00:02:31,420 --> 00:02:33,603
Remember, remember we talked about an 
execution plan, the minute will 

41
00:02:33,603 --> 00:02:39,114
manipulate that. 
And so it moves the filter beforehand 

42
00:02:39,114 --> 00:02:47,492
because it know it's safe to do so. 
And this is again, I keep coming back to 

43
00:02:47,492 --> 00:02:50,807
this but this is the, you know, you know, 
one advantage of a high level language 

44
00:02:50,807 --> 00:02:54,830
might be that it's easier to express your 
program in. 

45
00:02:54,830 --> 00:02:57,470
But that may or may not be true in many 
cases it could be actually harder to 

46
00:02:57,470 --> 00:03:01,275
express your program in terms of these, 
you know, limited operators. 

47
00:03:01,275 --> 00:03:05,405
so you know, you might prefer kind of 
general, a Java program might prefer a 

48
00:03:05,405 --> 00:03:11,050
full-width Java interface where they can 
just write whatever code they want. 

49
00:03:11,050 --> 00:03:15,430
But the point is that if you can, you 
know, tie one arm behind your back and 

50
00:03:15,430 --> 00:03:20,216
express your task in terms of these 
operations. 

51
00:03:20,216 --> 00:03:24,377
The system can take over and make 
manipulations that it would be very 

52
00:03:24,377 --> 00:03:30,422
difficult for it to do if you wrote in a 
lower level language, okay. 

53
00:03:30,422 --> 00:03:35,702
So, the other point to make here is that 
remember this is what we call lazy 

54
00:03:35,702 --> 00:03:39,510
evaluation. 
So th, a, at, when this command is, is 

55
00:03:39,510 --> 00:03:42,316
executed, that's when it makes an op, 
that's when it has an opportunity to 

56
00:03:42,316 --> 00:03:46,609
apply these optimizations. 
Right, so we did nothing but sort of 

57
00:03:46,609 --> 00:03:50,790
build up your, your execution plan as you 
were executing these commands. 

58
00:03:50,790 --> 00:03:53,360
It didn't actually, it didn't actually do 
any work. 

59
00:03:53,360 --> 00:03:55,960
And then finally when you say okay, now I 
really want the result, it can say aha I 

60
00:03:55,960 --> 00:03:59,080
see, I see your list of sequences. 
Or sorry, I see your sequence of 

61
00:03:59,080 --> 00:04:01,303
commands, I'm going to do some 
manipulation of those and figure out the 

62
00:04:01,303 --> 00:04:04,932
right way to execute these, these 
commands before I produce the result. 

63
00:04:04,932 --> 00:04:07,928
Okay. 
Fine, so were not done, we now have to 

64
00:04:07,928 --> 00:04:14,592
map this abstract execute command down 
into a sequence of map produce jobs. 

65
00:04:14,592 --> 00:04:19,472
And so the way it does this is it first 
identifies all the group and CoGroup 

66
00:04:19,472 --> 00:04:25,392
operators in the plan and assigns a, an 
individual MapReduce job to each one of 

67
00:04:25,392 --> 00:04:31,456
those operators. 
And then incrementally what it's going to 

68
00:04:31,456 --> 00:04:35,238
try to do is walk forward and backward 
and put as much work as it can into the 

69
00:04:35,238 --> 00:04:40,612
same MapReduce job, okay. 
So, the Group command needs to be it's 

70
00:04:40,612 --> 00:04:47,390
own MapReduce job because it's going to 
shuffle things across the network. 

71
00:04:47,390 --> 00:04:50,520
But filtering, I could apply, I could 
apply the filtering condition as a 

72
00:04:50,520 --> 00:04:53,734
[INAUDIBLE] data of a disc, or as I 
process a tuple I can check its condition 

73
00:04:53,734 --> 00:04:56,958
to make sure it's from one of these IP 
addresses before I actually apply the 

74
00:04:56,958 --> 00:05:00,343
group. 
So I might as well do that in the same 

75
00:05:00,343 --> 00:05:02,071
map reduced job, I don't need a whole 
independent map produced job I will just 

76
00:05:02,071 --> 00:05:04,326
do that. 
Further the load command is just reading 

77
00:05:04,326 --> 00:05:06,861
things off disk it doesn't actually do 
anything difficult at all I can put that 

78
00:05:06,861 --> 00:05:11,755
on the same map reduce job as well. 
And so it ends up with this one map 

79
00:05:11,755 --> 00:05:17,835
function that's going to look something 
like you know as you process a value its 

80
00:05:17,835 --> 00:05:29,069
going to say you know pause the data. 
And then just going to say if bow dot IP 

81
00:05:29,069 --> 00:05:46,942
equals this or a bow IP equals this. 
And emit IP val or val IP. 

82
00:05:46,942 --> 00:05:53,953
Val, right so this little program 
essentially gets generated by the 

83
00:05:53,953 --> 00:06:01,053
compiler as one as just one map file 
function. 

84
00:06:01,053 --> 00:06:04,926
And then gets executed. 
And on the reduce side you can do a 

85
00:06:04,926 --> 00:06:07,250
couple things. 
You can crih, construct the groups but 

86
00:06:07,250 --> 00:06:10,130
you don't actually care about the groups 
because you're immediately going to count 

87
00:06:10,130 --> 00:06:13,146
the results. 
And so it's smart enough to see that 

88
00:06:13,146 --> 00:06:16,656
group plus foreach is really going to 
just produce tuples like this: IP and 

89
00:06:16,656 --> 00:06:23,827
Count. 
And not, you know, not IP plus some big 

90
00:06:23,827 --> 00:06:28,918
group. 
And that's actually pretty significant 

91
00:06:28,918 --> 00:06:33,533
savings, because constructing these 
objects and sort of passing them around 

92
00:06:33,533 --> 00:06:38,395
is, is, is pretty expensive. 
So you don't actually need them you don't 

93
00:06:38,395 --> 00:06:43,730
want to use them. 
So there's sort of two levels of 

94
00:06:43,730 --> 00:06:48,982
optimization. 
One is it moves operators around, which I 

95
00:06:48,982 --> 00:06:52,031
like, y'know because they get this 
algebraic flavor. 

96
00:06:52,031 --> 00:06:56,119
And two it compresses logical operations 
into single physical operations, and the 

97
00:06:56,119 --> 00:06:59,703
overall name of the game here is to use 
to reduce the total number of map reduce 

98
00:06:59,703 --> 00:07:04,055
jobs being executed because they are 
expensive. 

99
00:07:04,055 --> 00:07:05,166
Okay. 
And then finally, you can write the 

100
00:07:05,166 --> 00:07:07,110
things out and that can also happen in 
the same reduce space. 

101
00:07:07,110 --> 00:07:10,400
So this entire task ends up being just 
one map reduce job Job. 

102
00:07:10,400 --> 00:07:13,280
In other cases you may need need to just 
do sort of multiple groupings or maybe 

103
00:07:13,280 --> 00:07:16,004
just short or so on. 
And each one of those requires its own 

104
00:07:16,004 --> 00:07:17,932
map reduce job. 
So for example, certain commands always 

105
00:07:17,932 --> 00:07:20,988
require a map reduce job like sorting. 
Okay. 

106
00:07:20,988 --> 00:07:25,980
So what did we talk about in the last 
several segments? 

107
00:07:25,980 --> 00:07:32,672
Well we described no sequel systems. 
And argue that they're important for a 

108
00:07:32,672 --> 00:07:35,885
data scientist to understand, in part 
because you may be using them but also 

109
00:07:35,885 --> 00:07:39,251
because you may be asked to weigh in on 
their relative strengths and weaknesses 

110
00:07:39,251 --> 00:07:43,710
compared to other systems. 
And so we talk about NoSQL sort of 

111
00:07:43,710 --> 00:07:46,902
meaning no schema, and no transactions, 
and no language. 

112
00:07:46,902 --> 00:07:50,142
Language and may be less about 
specifically bad SQL and overall they're 

113
00:07:50,142 --> 00:07:53,598
kind of a reboot of data systems zeroing 
in on just high-throughput reads and 

114
00:07:53,598 --> 00:07:59,260
writes. 
Okay. 

115
00:07:59,260 --> 00:08:02,676
Right, and so now the design space of 
these large scale data sets is sort of 

116
00:08:02,676 --> 00:08:06,204
being more fully exploit with different 
permutations and combinations of 

117
00:08:06,204 --> 00:08:10,708
particular features. 
But over all there's a clear trend back 

118
00:08:10,708 --> 00:08:14,348
towards re-introducing schemas and 
transactions and languages and so on, so 

119
00:08:14,348 --> 00:08:19,480
we talked about Google spanner system as 
an example of this trend. 

120
00:08:19,480 --> 00:08:23,320
So, no SQL is not, it's an evolving of 
concept, the main thing to realize is the 

121
00:08:23,320 --> 00:08:27,340
entire space of possibilities in large 
scale databases is being explored and no 

122
00:08:27,340 --> 00:08:33,704
sequel represents one segment relational 
database is representing another. 

123
00:08:33,704 --> 00:08:37,724
And there's new segments emerging so you 
may hear the term new sequal for example 

124
00:08:37,724 --> 00:08:41,684
which is kind of a architecture database 
that does try to achieve some of these a 

125
00:08:41,684 --> 00:08:46,286
other features okay. 
So fine then we talked about pig in 

126
00:08:46,286 --> 00:08:49,494
particular as an example not so much as 
another sequel. 

127
00:08:49,494 --> 00:08:53,843
Full system but as a, a analytic system 
that puts a layer on top of MapReduce. 

128
00:08:53,843 --> 00:08:57,253
And we chose pig because it has this 
relational algebra like layer on top of 

129
00:08:57,253 --> 00:09:00,718
Hadoop and so I want to make the point 
that relational algebra comes up in a lot 

130
00:09:00,718 --> 00:09:04,794
of different context. 
You know, and in, in this one it's sort 

131
00:09:04,794 --> 00:09:06,800
of a little bit interesting because 
although it has a clear relational 

132
00:09:06,800 --> 00:09:09,800
algebra flavor, it's not actually a pure 
relational data model. 

133
00:09:09,800 --> 00:09:11,720
You got this sort of. 
These structures. 

134
00:09:11,720 --> 00:09:14,366
And the point here is that you can 
kind of cherry-pick some of these 

135
00:09:14,366 --> 00:09:17,400
concepts and techniques pioneered by 
databases. 

136
00:09:17,400 --> 00:09:19,770
And you don't actually have to use them 
in the context of databases. 

137
00:09:19,770 --> 00:09:22,010
So you don't need to sort of throw the 
baby out with the bath water if a 

138
00:09:22,010 --> 00:09:24,610
database doesn't appear to be meeting 
your needs and these systems are sort of 

139
00:09:24,610 --> 00:09:28,931
doing that. 
So, as you review different systems for 

140
00:09:28,931 --> 00:09:34,244
merging, as a data scientist, you can 
start to understand the particular set of 

141
00:09:34,244 --> 00:09:38,783
feature they offer. 
And understand the problems with those 

142
00:09:38,783 --> 00:09:43,451
features and the history there. 
And the pros and cons therein. 

143
00:09:43,451 --> 00:09:44,850
Okay. 
And so Pig, also. 

144
00:09:44,850 --> 00:09:47,454
The, you know, one take away of it is 
that it has this sort of schema-on-read 

145
00:09:47,454 --> 00:09:50,226
fl, flavor rather than schema-on-write, 
meaning that you can work with in situ 

146
00:09:50,226 --> 00:09:52,917
data. 
Okay, I've mentioned this a couple of 

147
00:09:52,917 --> 00:09:55,208
times. 
I just wanted to point out that that's 

148
00:09:55,208 --> 00:09:58,663
really one of the new requirements 
associated with NoSQL as well as SQL 

149
00:09:58,663 --> 00:10:02,226
[UNKNOWN] is that you have to be able to 
work with data in it's kind of native 

150
00:10:02,226 --> 00:10:06,608
format. 
And the reason for that is it's just to 

151
00:10:06,608 --> 00:10:09,230
big to transform all over the place. 

