1
23:59:59,500 --> 00:00:05,857
[MUSIC]. 

2
00:00:05,857 --> 00:00:11,083
Okay, so let's talk a little bit about 
other large scale data processing systems 

3
00:00:11,083 --> 00:00:15,217
besides MapReduce. 
And as a step to that, let's think about 

4
00:00:15,217 --> 00:00:20,414
the design space of possibilities here. 
So, this is a breakdown that was proposed 

5
00:00:20,414 --> 00:00:25,100
by Michael Isard who's developed a system 
called Dryad at Microsoft, which is a 

6
00:00:25,100 --> 00:00:31,990
very nice system with the same sort of 
motivations as, as MapReduce, okay. 

7
00:00:31,990 --> 00:00:35,488
And so, he divided this space into these 
three axes where he worried about, sort 

8
00:00:35,488 --> 00:00:40,640
of, low latency, very interactive sort of 
speeds, quick turnaround time. 

9
00:00:40,640 --> 00:00:43,727
Versus things that maximize kind of 
throughput, you know massive batch jobs 

10
00:00:43,727 --> 00:00:47,725
operating on you know, thousands and 
thousands of computers at once. 

11
00:00:47,725 --> 00:00:51,181
Versus another access here is sort of 
whether it's in a private data center or 

12
00:00:51,181 --> 00:00:54,699
whether it's scaled out widely over the 
internet. 

13
00:00:56,860 --> 00:01:01,621
and then maybe the third access is data 
parallel versus the shared memory, okay 

14
00:01:01,621 --> 00:01:05,234
that we talked about. 
So the, the areas that we're mostly 

15
00:01:05,234 --> 00:01:08,260
concerned with are going to be here which 
is what we're talking about currently. 

16
00:01:08,260 --> 00:01:12,724
And then in a couple of segments we're 
going to talk about these Low-latency, 

17
00:01:12,724 --> 00:01:19,180
smaller operations which you can think of 
as the no, the NoSQL systems. 

18
00:01:19,180 --> 00:01:21,290
Okay. 
And then maybe where, where Michael 

19
00:01:21,290 --> 00:01:24,640
placed older ar, relational databases, 
although he didn't label it the same way 

20
00:01:24,640 --> 00:01:27,790
I have. 
Is down here in this quadrant where 

21
00:01:27,790 --> 00:01:31,490
they're mostly in shared memories with 
the space and low latency. 

22
00:01:31,490 --> 00:01:35,577
I, I say older databases to try to point 
out that not all relational databases 

23
00:01:35,577 --> 00:01:39,629
operate in the space. 
Many are, most are in fact are data para, 

24
00:01:39,629 --> 00:01:44,375
most parallel databases are. 
Are data-parallel, as well. 

25
00:01:44,375 --> 00:01:49,803
And so this is, you know, MySQL and 
PostgreSQL, if you're familiar with 

26
00:01:49,803 --> 00:01:55,190
those, are probably in this space. 
Shared memory, shared disk. 

27
00:01:55,190 --> 00:01:57,411
Alright. 
And here, HPC means high performance 

28
00:01:57,411 --> 00:02:01,401
computing, where you know, it's private 
data center and it's a big mass of shared 

29
00:02:01,401 --> 00:02:05,120
memory. 
but there's, but, but it's a batch job 

30
00:02:05,120 --> 00:02:08,290
submission system, right? 
You say, you're on your big compute, 

31
00:02:08,290 --> 00:02:10,640
compute-intensive job and submit it to 
the machine. 

32
00:02:10,640 --> 00:02:15,660
And it processes on it and eventually 
turns the results. 

33
00:02:15,660 --> 00:02:18,315
And then, I'm not going to talk too much 
about this, but this notion of grid 

34
00:02:18,315 --> 00:02:21,375
computing was really sort of focused on 
connecting up clusters of computers from 

35
00:02:21,375 --> 00:02:25,894
different universities and letting them 
all sort of talk to each other. 

36
00:02:25,894 --> 00:02:28,500
And so that's why he's pushed this up on 
this other axis. 

37
00:02:28,500 --> 00:02:31,620
Increasingly, you're seeing these systems 
also being pushed up in the in this 

38
00:02:31,620 --> 00:02:34,668
access. 
Internet scale, planetary, you know, 

39
00:02:34,668 --> 00:02:38,223
distributed hash tables. 
With different kinds of layers for 

40
00:02:38,223 --> 00:02:43,340
guaranteeing certain kinds of semantic. 
So the spanner system from Google fairly 

41
00:02:43,340 --> 00:02:46,950
recently is, is a nice example of this. 
Okay. 

42
00:02:46,950 --> 00:02:52,735
So, sort of wrap up what we talked about 
last time, large scale data processing 

43
00:02:52,735 --> 00:03:00,540
you know, many tasks need to process big 
data and produce big data. 

44
00:03:00,540 --> 00:03:02,650
And so you want to use hundreds of 
thousands of. 

45
00:03:02,650 --> 00:03:07,130
The CPUs are, and hundreds of thousand or 
tens of thousands of computers to solve 

46
00:03:07,130 --> 00:03:11,350
these problems, but this needs to be much 
easier. 

47
00:03:11,350 --> 00:03:13,100
And so there are such things as parallel 
database. 

48
00:03:13,100 --> 00:03:16,805
We talked about Databases and we sort of 
extolled the virtues of programming in 

49
00:03:16,805 --> 00:03:20,185
that model, and they exist but they're 
often. 

50
00:03:20,185 --> 00:03:22,784
Expensive. 
Well, they, they almost exclusively are 

51
00:03:22,784 --> 00:03:25,810
expensive. 
And they're difficult to set up. 

52
00:03:25,810 --> 00:03:29,026
And it's actually not totally clear that 
many of the parallel databases scale to 

53
00:03:29,026 --> 00:03:31,980
really hundreds or thousands of, of 
machines. 

54
00:03:31,980 --> 00:03:35,550
Ok. 
And so MapReduce. 

55
00:03:38,090 --> 00:03:41,850
Came around at a time at a time as a, as 
a bit of response to this scenario. 

56
00:03:41,850 --> 00:03:44,402
So it's more a light weight framework 
featuring, you know, automatic 

57
00:03:44,402 --> 00:03:47,526
parallelization and distribution that 
we've been talking about, featuring I, a 

58
00:03:47,526 --> 00:03:50,617
fault-tolerance. 
And I mentioned a couple of other things 

59
00:03:50,617 --> 00:03:52,473
here that are probably less important but 
the I/O scheduling, the status and 

60
00:03:52,473 --> 00:03:55,107
monitoring. 
So really sort of strip everything down 

61
00:03:55,107 --> 00:03:58,568
to just parallel processing. 
Not all the features of parallel database 

62
00:03:58,568 --> 00:04:01,214
is offered, just parallel processing, 
with the added benefit of 

63
00:04:01,214 --> 00:04:04,566
fault-tolerance. 
Okay, and this really seemed to scratch a 

64
00:04:04,566 --> 00:04:07,200
niche with people. 
I'll, I've, I'll argue here and I'll 

65
00:04:07,200 --> 00:04:11,028
probably mention this again at some point 
that it's not totally clear to me that 

66
00:04:11,028 --> 00:04:14,392
MapReduce would have been quite so 
popular had there been available a 

67
00:04:14,392 --> 00:04:20,358
parallel open source. 
relational database product, but all the 

68
00:04:20,358 --> 00:04:23,800
open source databases were all not, not 
parallel. 

69
00:04:23,800 --> 00:04:26,484
In fact they were not even single 
threaded for query processing, you know 

70
00:04:26,484 --> 00:04:30,075
only a single thread was working on an 
individual query at a time. 

71
00:04:30,075 --> 00:04:38,150
But that's, that's, that's speculation. 
Well, actually, I have a little a bit of 

72
00:04:38,150 --> 00:04:40,650
evidence for that that I'll lay out in a 
bit. 

73
00:04:40,650 --> 00:04:42,308
Okay. 
So now I want to talk about, maybe, I 

74
00:04:42,308 --> 00:04:45,505
guess I'm building up to that argument. 
I want to talk about parallel databases 

75
00:04:45,505 --> 00:04:48,066
and how they work. 
And hopefully, show that there's some 

76
00:04:48,066 --> 00:04:51,126
similarities, show where there are 
similarities and where there. 

77
00:04:51,126 --> 00:04:52,990
There are differences. 
Okay. 

78
00:04:52,990 --> 00:04:57,870
So we'll call that a key idea of. 
Relational databases was this notion of a 

79
00:04:57,870 --> 00:05:00,980
relational algebra, where you could write 
sort of plans like this. 

80
00:05:00,980 --> 00:05:05,335
And that this top-level language called 
SQL was the most common way of Producing 

81
00:05:05,335 --> 00:05:09,853
a Relational Algebra plan, right? 
You wrote the query and sequel and it was 

82
00:05:09,853 --> 00:05:12,950
automatically turned into a Relational 
Algebra plan by the system. 

83
00:05:12,950 --> 00:05:15,230
OK. 
So this is kind of thrown out the window 

84
00:05:15,230 --> 00:05:19,256
with MapReduce, arguably in favor of 
sort of flexibility in providing the 

85
00:05:19,256 --> 00:05:25,556
program with more, more control. 
But let's go back to this model for a 

86
00:05:25,556 --> 00:05:27,921
bit. 
Fine. 

87
00:05:27,921 --> 00:05:30,363
So. 
Now we want to, we want to evaluate these 

88
00:05:30,363 --> 00:05:33,270
queries, we want to do it in parallel 
now. 

89
00:05:33,270 --> 00:05:36,042
And so there's two different terms that I 
want you to be familiar with, one is 

90
00:05:36,042 --> 00:05:39,034
distributed query and one is parallel 
query or distributed query processing and 

91
00:05:39,034 --> 00:05:42,632
parallel processing. 
And so they're both ways of sort of 

92
00:05:42,632 --> 00:05:45,964
taking advantage of more computing 
resources for the same query but they're, 

93
00:05:45,964 --> 00:05:50,724
behave a little differently. 
So the distributed query, what you're 

94
00:05:50,724 --> 00:05:55,332
doing is taking a, a single large table 
and distributing it across a cluster, 

95
00:05:55,332 --> 00:06:01,260
just like we talked about. 
And then you're breaking your query into 

96
00:06:01,260 --> 00:06:07,620
individual pieces to operate on each of 
those partitions of the file. 

97
00:06:07,620 --> 00:06:09,750
Okay. 
So this sounds like, well isn't that 

98
00:06:09,750 --> 00:06:11,720
basically just the same thing as map 
reduce? 

99
00:06:11,720 --> 00:06:16,072
It is, except for the fact that, all the 
results of those individual pieces are 

100
00:06:16,072 --> 00:06:19,464
all sent back to the head node to a 
single server to sort of finish 

101
00:06:19,464 --> 00:06:23,850
processing. 
So for example if you're doing a, a, a 

102
00:06:23,850 --> 00:06:28,537
count, right, you want to count all the 
records that match some criteria. 

103
00:06:28,537 --> 00:06:32,112
Well if you have a very large file that's 
split across several machines, these 

104
00:06:32,112 --> 00:06:35,632
distributed query systems, Microsoft 
Sequel Server in particular is, is an 

105
00:06:35,632 --> 00:06:39,372
example of this, is smart enough to break 
the query into a bunch a little pieces 

106
00:06:39,372 --> 00:06:44,266
and run each of these of pieces in 
parallel. 

107
00:06:44,266 --> 00:06:47,896
But as they start to string tuples out to 
be counted they'll send them all back to 

108
00:06:47,896 --> 00:06:51,002
the head node. 
Actually, I guess that may not actually 

109
00:06:51,002 --> 00:06:52,974
be true so maybe it's one of the thongs 
that you can just count things in 

110
00:06:52,974 --> 00:06:55,998
parallel and add them. 
The map, but it's not hard to construct a 

111
00:06:55,998 --> 00:06:58,644
query where, where you have this 
bottleneck of, of sending everything back 

112
00:06:58,644 --> 00:07:01,951
to a single server. 
So it's essentially, you can think of it 

113
00:07:01,951 --> 00:07:05,060
as having the map phase, but not really 
the reduce phase. 

114
00:07:05,060 --> 00:07:09,956
Now, parallel query, every individual 
operator in the relation to algebra is 

115
00:07:09,956 --> 00:07:13,786
implemented in parallel. 
So when you're doing joins, you're doing 

116
00:07:13,786 --> 00:07:16,903
joins across. 
A bunch of nodes when you're doing 

117
00:07:16,903 --> 00:07:19,743
groupings. 
You're doing grouping across a bunch of 

118
00:07:19,743 --> 00:07:21,996
nodes. 
And we're seen how to implement 

119
00:07:21,996 --> 00:07:26,154
relational join in map reduce and it's 
not too far off from how it's actually 

120
00:07:26,154 --> 00:07:30,491
implemented inside databases. 
Okay, so if we how to implmenet join, 

121
00:07:30,491 --> 00:07:33,515
that's usually the harder one. 
One, trust me that you can implement the 

122
00:07:33,515 --> 00:07:37,200
other ones that way. 
Well, now we have a way to do parallel 

123
00:07:37,200 --> 00:07:41,310
query processing with the relational 
algebra. 

124
00:07:41,310 --> 00:07:43,270
You know, so why not do that? 
Well, the answer is that people do do 

125
00:07:43,270 --> 00:07:45,250
that, and we'll come back to that in one 
second. 

126
00:07:45,250 --> 00:07:48,820
Okay, so for a distributed query, I guess 
I was waving my hands a second ago trying 

127
00:07:48,820 --> 00:07:52,698
to explain this, when, when it was all on 
the next slide. 

128
00:07:52,698 --> 00:07:57,045
You can imagine constructing a view, and 
we talked about views, if you don't 

129
00:07:57,045 --> 00:08:01,461
recall what that is, it's you know a, a 
named query that can be then accessed as 

130
00:08:01,461 --> 00:08:07,692
a single table. 
So we say that the sales table is really 

131
00:08:07,692 --> 00:08:12,970
the union of a bunch of smaller sales 
tables, one for each month, and in 

132
00:08:12,970 --> 00:08:18,703
particular you could put each one of 
these sales table on a different disc or 

133
00:08:18,703 --> 00:08:28,060
even a different server all together. 
Right? 

134
00:08:28,060 --> 00:08:31,781
And then the, the user who is querying 
the sales table doesn't have to care 

135
00:08:31,781 --> 00:08:36,850
about the fact this is actually 
distributed distributed table. 

136
00:08:36,850 --> 00:08:38,995
They don't have to go gather up all the 
results from January, and then gather up 

137
00:08:38,995 --> 00:08:41,140
all the results from February, and the 
gather up all the results from March and 

138
00:08:41,140 --> 00:08:44,812
put them all together. 
That's done automatically by the system. 

139
00:08:44,812 --> 00:08:49,436
However right, and so this, this is, this 
is the create table statement that we 

140
00:08:49,436 --> 00:08:54,060
didn't talk about for constructing the 
sales table for sa, individual March, you 

141
00:08:54,060 --> 00:08:58,435
know, okay. 
But again, however, when you process this 

142
00:08:58,435 --> 00:09:01,360
stuff in parallel that works great but 
when you get the results you need to send 

143
00:09:01,360 --> 00:09:05,240
them all back to the single node to for, 
to finish processing. 

144
00:09:05,240 --> 00:09:06,900
And that's the limitation Distributed 
query. 

145
00:09:06,900 --> 00:09:09,319
So it's great that you get some 
parallelism, but it can't do everything 

146
00:09:09,319 --> 00:09:11,661
in parallel. 
And you, you can, you can see this when 

147
00:09:11,661 --> 00:09:15,874
you run performance experiments. 
But a true parallel query example, would 

148
00:09:15,874 --> 00:09:20,028
be, for example from a system called 
Teradata, which is a database company 

149
00:09:20,028 --> 00:09:24,383
that many folks haven't heard of because 
they're selling very, very high-end 

150
00:09:24,383 --> 00:09:29,679
databases to very, very high-end 
customers. 

151
00:09:29,679 --> 00:09:33,967
And so they don't sort of need to have 
much word of mouth in the popular news 

152
00:09:33,967 --> 00:09:40,120
media. 
But what's happening here is that as 

153
00:09:40,120 --> 00:09:43,945
every. 
Individual rows inserted into the 

154
00:09:43,945 --> 00:09:49,560
parallel database. 
It will be assigned to some particular 

155
00:09:49,560 --> 00:09:52,090
server using a hash function. 
Okay. 

156
00:09:52,090 --> 00:09:55,410
Fine. 
So, everything is automatically 

157
00:09:55,410 --> 00:10:01,350
partitioned more or less randomly across 
the cluster and then whenever you're 

158
00:10:01,350 --> 00:10:07,560
running queries on this, all of this, all 
the machines will access their, their 

159
00:10:07,560 --> 00:10:13,419
data in parallel. 
Okay, and you can see this has a little 

160
00:10:13,419 --> 00:10:17,460
bit of the flavor of how we did the 
relational join in in MapReduce. 

161
00:10:17,460 --> 00:10:21,175
And that should be coming more clear in a 
second. 

162
00:10:21,175 --> 00:10:24,355
Okay, so remember this is our query 
orders and line items, and this is the 

163
00:10:24,355 --> 00:10:28,860
plan that we're going to do. 
We're going to select some orders and 

164
00:10:28,860 --> 00:10:32,740
then join the orders with the items. 
Alright. 

165
00:10:32,740 --> 00:10:38,350
So how this starts is these parallel 
processing units, units called amps in 

166
00:10:38,350 --> 00:10:45,110
teradata terms, will each contain a piece 
of the data. 

167
00:10:45,110 --> 00:10:48,690
A chunk of the data, and the chunk was to 
find sort of randomly by hashing. 

168
00:10:48,690 --> 00:10:51,893
Okay. 
And so they all in parallel should begin 

169
00:10:51,893 --> 00:10:57,622
to scan their, their individual chunk. 
And then they'll all in parallel apply 

170
00:10:57,622 --> 00:11:01,091
the filtering conditions to throw out 
certain records that they don't want. 

171
00:11:01,091 --> 00:11:07,993
And then they'll all in parallel hash on 
the so this isn't right. 

172
00:11:07,993 --> 00:11:11,680
This should hash on the order. 
ID. 

173
00:11:11,680 --> 00:11:12,610
Not the item ID. 
[BLANK_AUDIO]. 

174
00:11:18,110 --> 00:11:20,252
So this is the join key, right. 
We're going to join on order. 

175
00:11:20,252 --> 00:11:23,218
Here. 
and I probably had this wrong back here 

176
00:11:23,218 --> 00:11:25,441
too. 
Yeah, this is wrong as well, this should 

177
00:11:25,441 --> 00:11:27,564
be join on order ID. 
Doesn't quite make sense to call this 

178
00:11:27,564 --> 00:11:32,272
item. 
Same thing here. 

179
00:11:32,272 --> 00:11:43,115
Okay, so then they're all in parallel 
hash on the join, join attri/g. 

180
00:11:43,115 --> 00:11:46,834
Tribute the order ID. 
And that will shuffle it, for lack of a 

181
00:11:46,834 --> 00:11:51,066
better term to another set of amps. 
Perhaps the same set of amps, but 

182
00:11:51,066 --> 00:11:55,080
typically another set of amps that will 
do the next step. 

183
00:11:55,080 --> 00:11:58,695
And actually compete the join. 
And so this should look like a MapReduce 

184
00:11:58,695 --> 00:12:02,514
job, right. 
You've got a map function that's scanning 

185
00:12:02,514 --> 00:12:06,770
and selecting and, and then, and actually 
then hashing. 

186
00:12:06,770 --> 00:12:09,668
And then we gotta re, reduce function 
coming up to actually produce the join. 

187
00:12:09,668 --> 00:12:13,542
Okay. 
And for the other relation the same thing 

188
00:12:13,542 --> 00:12:18,578
happens, you scan the items. 
And then hash on. 

189
00:12:18,578 --> 00:12:25,708
Again, this is order, order, order. 
Right, so you're scanning the order, 

190
00:12:25,708 --> 00:12:31,540
scanning and selecting on the orders and 
just scanning on the items. 

191
00:12:31,540 --> 00:12:34,310
And then both are hashed on the 
appropriate joint attribute. 

192
00:12:34,310 --> 00:12:37,710
And, lo and behold, all the items and all 
the orders that correspond to the same 

193
00:12:37,710 --> 00:12:40,960
Order ID to the same join attribute, same 
join key end up on the same machine and 

194
00:12:40,960 --> 00:12:46,460
you can actually process the join, just 
like the MapReduce example we saw. 

195
00:12:46,460 --> 00:12:50,110
Alright. 
So then at the end of these two steps 

196
00:12:50,110 --> 00:12:57,550
that I've shown you, amp four will have 
all the orders and all the line items 

197
00:12:57,550 --> 00:13:09,030
where hash of order, goodness equals 1. 
And this AMP5 will have have all the 

198
00:13:09,030 --> 00:13:15,930
orders and items where hash of order 
equals 2. 

199
00:13:15,930 --> 00:13:19,375
And this one will have all the orders and 
line items where hash of order equals 3 

200
00:13:19,375 --> 00:13:22,502
and now it has enough enough to 
individually and in parallel finish the 

201
00:13:22,502 --> 00:13:28,815
join and actually produce the result. 
And all these other the orders as well. 

202
00:13:28,815 --> 00:13:35,960
Alright. 
So, fine, so the point is, is that the 

203
00:13:35,960 --> 00:13:40,430
same machinery already exists in these 
parallel databases. 

204
00:13:40,430 --> 00:13:44,212
And in Map Reduce you know, if you're 
interested in, in doing a join, you're 

205
00:13:44,212 --> 00:13:49,084
sort of implementing this yourself. 
And so this observation was not lost on 

206
00:13:49,084 --> 00:13:52,568
people that, you know, hey, it might be 
nice if there was sort of a standard way 

207
00:13:52,568 --> 00:13:57,027
of doing join in MapReduce. 
And we didn't have to sort of rewrite it 

208
00:13:57,027 --> 00:14:00,563
out ourselves every time and in fact, you 
know, there's, there's libraries on top 

209
00:14:00,563 --> 00:14:05,742
of MapReduce that do this. 
And so there's a library called Pig from 

210
00:14:05,742 --> 00:14:12,250
Yahoo, that encourages you to check out 
that Is recognizably relational algebra. 

211
00:14:12,250 --> 00:14:15,780
Alright, there are operators called join, 
there are operators called group by. 

212
00:14:15,780 --> 00:14:18,074
It does have a bit of a funny data model, 
where you're allowed to have kind of 

213
00:14:18,074 --> 00:14:22,240
complicated nesting. 
As opposed to just straight tuples and 

214
00:14:22,240 --> 00:14:24,861
straight relations. 
But the relation algebra is there, and in 

215
00:14:24,861 --> 00:14:28,242
fact, this is sort of one of the points I 
want to make, is that. 

216
00:14:28,242 --> 00:14:31,634
It, it, you know, it's important to sort 
of be able to modularize the concepts 

217
00:14:31,634 --> 00:14:35,750
that come out of various communities and 
especially databases. 

218
00:14:35,750 --> 00:14:38,744
They, it tends to be true that you know, 
it's kind of all or nothing. 

219
00:14:38,744 --> 00:14:41,223
If you're interested in using databases, 
well then you have to take everything, 

220
00:14:41,223 --> 00:14:43,390
you have, you have to take the whole 
package. 

221
00:14:43,390 --> 00:14:45,854
You know, it's all or nothing. 
But increasingly what you're finding is 

222
00:14:45,854 --> 00:14:48,385
that these concepts are leaking out into 
other systems. 

223
00:14:48,385 --> 00:14:52,600
Which is why I'm really, emphasizing this 
relational algebra piece a lot. 

224
00:14:52,600 --> 00:14:57,220
Is that you can use these concepts 
independently of buying in to a strict 

225
00:14:57,220 --> 00:15:00,300
relational model. 
Okay. 

226
00:15:00,300 --> 00:15:03,732
And certainly not a strict, a strict 
adherence to, to a particular information 

227
00:15:03,732 --> 00:15:06,683
of it. 
Now, another system called HIVE is 

228
00:15:06,683 --> 00:15:08,708
literally SQL on top of. 
[INAUDIBLE]. 

229
00:15:08,708 --> 00:15:11,636
So it's goes one step even higher. 
Instead of just stopping the Relational 

230
00:15:11,636 --> 00:15:14,105
Algebra level, it actually provides a 
sequel interface. 

231
00:15:14,105 --> 00:15:17,785
Impala's a more recent system from 
Cloudera I should mention here. 

232
00:15:17,785 --> 00:15:20,435
Cloudera by the way, is a company that 
has. 

233
00:15:20,435 --> 00:15:24,323
align themselves pretty closely with the, 
the Hadoop stack, and so they have their 

234
00:15:24,323 --> 00:15:27,725
own fork of the Hadoop system and a bunch 
of great tools for working with that 

235
00:15:27,725 --> 00:15:31,235
ecosystem, and Impala is a new system 
that they produced that is, provides SQL 

236
00:15:31,235 --> 00:15:37,905
over HDFS, and actually uses a lot of the 
code from the HIVE system. 

237
00:15:37,905 --> 00:15:41,324
Okay. 
Cascading is another system that's maybe 

238
00:15:41,324 --> 00:15:43,680
a little bit less common, but it's also 
very recognizably to be relational 

239
00:15:43,680 --> 00:15:46,500
algebra. 
The Dryad system I mentioned has nothing 

240
00:15:46,500 --> 00:15:49,600
to do with MapReduce directly except for 
a similar motivation, but it very 

241
00:15:49,600 --> 00:15:54,326
obviously has relational algebra there. 
the Clustera system I, I mentioned, it's 

242
00:15:54,326 --> 00:15:58,910
more a research project and is not clear 
to me that the code is available. 

243
00:15:58,910 --> 00:16:00,520
But it's also very clearly relational 
algebra. 

244
00:16:00,520 --> 00:16:03,160
So, you know, when you put your 
relational algebra goggles on, you start 

245
00:16:03,160 --> 00:16:07,000
to see the world in this way, and it 
starts to come up everywhere, okay? 

246
00:16:07,000 --> 00:16:11,372
So it's good to go back and understand 
those operations. 

247
00:16:11,372 --> 00:16:18,584
All right. 

