1
23:59:59,500 --> 00:00:04,704
[MUSIC]. 

2
00:00:04,704 --> 00:00:07,613
Okay. 
Yes, so going back to the grid we can 

3
00:00:07,613 --> 00:00:10,892
highlight those systems that were based 
on MapReduce itself. 

4
00:00:10,892 --> 00:00:15,917
You know, the MapReduce paper itself in 
2004 and these language layers on top Pig 

5
00:00:15,917 --> 00:00:21,747
and Hive in 2008, where Hive is SQL. 
And Pig is a relational algebra looking 

6
00:00:21,747 --> 00:00:25,285
language that we'll talk about in some 
detail in the next few segments and 

7
00:00:25,285 --> 00:00:30,696
Tenzing which is also SQL. 
And Impala, which is also SQL, or Tenzing 

8
00:00:30,696 --> 00:00:34,728
is from Google and Impala is from a 
company called Cloudera, that's a pretty 

9
00:00:34,728 --> 00:00:41,560
eager evangelist of MapReduce and Hadoop, 
based technologies in general, okay. 

10
00:00:41,560 --> 00:00:45,400
So, one trend I think you see is that 
these declarative languages on top of the 

11
00:00:45,400 --> 00:00:50,950
parallel processing primitive of 
MapReduce are really here to stay, right. 

12
00:00:50,950 --> 00:00:56,186
So, people that were relatively against 
these kind of languages certainly doing 

13
00:00:56,186 --> 00:00:59,611
it. 
Now, it's also fair to say that the 

14
00:00:59,611 --> 00:01:06,180
enterprises in general have made a pretty 
significant investment in SQL expertise. 

15
00:01:07,190 --> 00:01:11,143
So, even if they're attracted to the, to 
the advantages that Hadoop might bring, 

16
00:01:11,143 --> 00:01:14,528
they're, they're pretty much demanding 
SQL. 

17
00:01:14,528 --> 00:01:18,000
So, this may be a response to this 
inertia from having invested in SQL in 

18
00:01:18,000 --> 00:01:21,276
the past. 
I think that is certainly true, however, 

19
00:01:21,276 --> 00:01:24,624
it is also true that the desire for 
declarative languages reasonably well 

20
00:01:24,624 --> 00:01:28,201
founded for reasons we've already talked 
about. 

21
00:01:28,201 --> 00:01:30,093
Okay. 
So, you can put these systems on this 

22
00:01:30,093 --> 00:01:33,387
time line and the only point I want to 
make about this, is that there is a bit 

23
00:01:33,387 --> 00:01:38,080
of a gap between the paper in 2004 and 
the systems in 2008. 

24
00:01:38,080 --> 00:01:41,664
But as soon as we had Hadoop the system 
itself, developed at Yahoo and, and 

25
00:01:41,664 --> 00:01:45,066
released as an Apache open sourced 
project. 

26
00:01:45,066 --> 00:01:49,275
You sort of immediately see an ecosystem 
start to emerge with extensions to it, 

27
00:01:49,275 --> 00:01:53,822
that add these, in particular, add these 
languages on top. 

28
00:01:53,822 --> 00:01:57,665
And so I think that the, the need for 
high level interface is motivated by how 

29
00:01:57,665 --> 00:02:02,640
quickly they came around as soon as 
Hadoop wa, was, was out. 

30
00:02:02,640 --> 00:02:06,355
And again, it, it didn't stop with these 
later systems a few years later. 

31
00:02:06,355 --> 00:02:10,130
Okay. 
And actually you know, not on this page 

32
00:02:10,130 --> 00:02:15,380
there's potentially hundreds of, if you 
include you know, research projects based 

33
00:02:15,380 --> 00:02:21,383
on extensions to MapReduce. 
There are really, really a lot okay and 

34
00:02:21,383 --> 00:02:25,680
so this is some of the most popular ones 
alright. 

35
00:02:25,680 --> 00:02:31,110
So, another subset of this grid that you 
can look at is just the no SQL system. 

36
00:02:31,110 --> 00:02:34,134
Now the whole last few segments have 
sensibly been about NoSQL but I've also 

37
00:02:34,134 --> 00:02:37,295
included these kinds of analytic systems 
in here. 

38
00:02:37,295 --> 00:02:42,080
MapReduce base systems and a few others 
for example Dremel and Spark and Shark. 

39
00:02:42,080 --> 00:02:47,435
Dremel is a system from Google that is 
the back end of a query, as a service 

40
00:02:47,435 --> 00:02:53,480
system called Google BigQuery, which is 
pretty nice. 

41
00:02:53,480 --> 00:02:56,180
I, I recommend taking a look at it, you 
can sort of upload data and put it in 

42
00:02:56,180 --> 00:02:59,195
there and it doesn't matter how big it is 
and you can kind of query it at, at very 

43
00:02:59,195 --> 00:03:04,560
low latency speed. 
Spark and Shark, come from the AMPLab at 

44
00:03:04,560 --> 00:03:11,411
Berkeley and are part of the Berkeley 
Data Analytics Stack, BDAS or BDAS. 

45
00:03:11,411 --> 00:03:17,021
And Spark is a language label on top of, 
it's not MapReduce, but on top of a 

46
00:03:17,021 --> 00:03:25,760
parallel processing system and Shark is a 
SQL layer even on top of that, okay. 

47
00:03:25,760 --> 00:03:28,116
And so a couple of distinguishing 
features of Spark is it loads everything 

48
00:03:28,116 --> 00:03:32,090
to memory. 
Process everything there when possible. 

49
00:03:32,090 --> 00:03:35,038
Writing things out to disk only for fault 
tolerance reasons, so much, much, less 

50
00:03:35,038 --> 00:03:38,621
often than MapReduce. 
and it also supports iterative processing 

51
00:03:38,621 --> 00:03:41,806
which is pretty important and we're going 
to come back to that a little later in 

52
00:03:41,806 --> 00:03:44,868
the course. 
And then Shark again is just SQL on top 

53
00:03:44,868 --> 00:03:46,270
of, on top of this. 
All right. 

54
00:03:46,270 --> 00:03:51,292
So, within these NoSQL systems one thing 
you can look at, there's not much I 

55
00:03:51,292 --> 00:03:58,157
want to, excuse me. 
There's not much I want to say about this 

56
00:03:58,157 --> 00:04:02,966
diagram. 
except that, you know, to point out that 

57
00:04:02,966 --> 00:04:06,830
there's been sort of of a cambrian 
explosion. 

58
00:04:06,830 --> 00:04:10,106
So, first, you had sort of memcached, 
which is again just a caching layer for 

59
00:04:10,106 --> 00:04:13,498
really the real system. 
And the real system was a bunch of MySQL 

60
00:04:13,498 --> 00:04:16,578
databases that weren't really working all 
that well for the requirements, they're 

61
00:04:16,578 --> 00:04:19,852
having them used for. 
But you can bring this in the memory and 

62
00:04:19,852 --> 00:04:23,285
keep it there, looking it up by name. 
And it was just like a performance 

63
00:04:23,285 --> 00:04:26,135
enhancement, a free performance 
enhancement, if you invest in this 

64
00:04:26,135 --> 00:04:29,287
system. 
The real you know, approach of throwing 

65
00:04:29,287 --> 00:04:34,496
out everything you have and replacing it 
with a NoSQL system came a little later. 

66
00:04:34,496 --> 00:04:36,936
So the only couple points I want to make 
is that it's been kind of a cambrian 

67
00:04:36,936 --> 00:04:39,620
explosion of a different systems around 
this time. 

68
00:04:39,620 --> 00:04:44,880
And that you know, this space in here is 
nowhere near as empty as it looks. 

69
00:04:44,880 --> 00:04:49,296
I've picked out a few systems here, but 
really the emergence of new systems in 

70
00:04:49,296 --> 00:04:53,740
the space hasn't really slowed down much 
at all. 

71
00:04:53,740 --> 00:05:00,080
So, really filling out the design space 
here since around 2006. 

72
00:05:00,080 --> 00:05:03,056
And the only other point I made is that 
these, these quite popular document 

73
00:05:03,056 --> 00:05:06,368
oriented data models systems, CouchDB and 
MongoDB, have actually been around for 

74
00:05:06,368 --> 00:05:10,157
quite a while. 
So, they've been, you know, they were 

75
00:05:10,157 --> 00:05:14,730
released in 2005 and 2000 seven. 
So, sometimes I think they seem like 

76
00:05:14,730 --> 00:05:18,980
newer systems, but, but they've have some 
maturity. 

77
00:05:18,980 --> 00:05:21,616
Okay. 
So, then finally one, one other point I 

78
00:05:21,616 --> 00:05:28,490
want to make is about this column called, 
that I've called Joins and Analytics. 

79
00:05:28,490 --> 00:05:32,150
And so, one of the distinguishing 
features of NoSQL systems is that they, 

80
00:05:32,150 --> 00:05:35,890
they typically don't support any notion 
of joins. 

81
00:05:35,890 --> 00:05:38,066
And you know, they would, many of the 
proponents of these systems would argue 

82
00:05:38,066 --> 00:05:40,751
that that's okay, that their joining 
aren't really necessary. 

83
00:05:40,751 --> 00:05:45,587
So, the argument about why you can get 
away from joins, goes something like 

84
00:05:45,587 --> 00:05:49,658
this. 
When you're joining two tables you sort 

85
00:05:49,658 --> 00:05:55,920
of have one record and a bunch of 
corresponding records in another table. 

86
00:05:55,920 --> 00:05:59,454
Well, if I put all those corresponding 
records and co-locate them with the 

87
00:05:59,454 --> 00:06:02,473
parent record, right? 
Then I can access everything all at once 

88
00:06:02,473 --> 00:06:04,420
and I don't actually need to compute the 
join, okay. 

89
00:06:04,420 --> 00:06:07,954
So, that works and it should sound 
somewhat sort a familiar, because it was 

90
00:06:07,954 --> 00:06:11,659
part of these network and hierarchical 
data models that we talked about in the 

91
00:06:11,659 --> 00:06:16,860
discussion of how we motivated relational 
databases. 

92
00:06:16,860 --> 00:06:19,996
And if you remember, the down side of 
this was that you had kind of one 

93
00:06:19,996 --> 00:06:25,542
dominant access path into the data. 
But any other access path was either not 

94
00:06:25,542 --> 00:06:29,835
supported at all or was inefficient, 
okay. 

95
00:06:29,835 --> 00:06:33,615
And if you want to reorganize your data 
to support a different access path or 

96
00:06:33,615 --> 00:06:37,575
because your, your needs changed well, a 
lot of your code might of broken and you 

97
00:06:37,575 --> 00:06:42,700
would need to sort of rewrite it. 
Okay. 

98
00:06:42,700 --> 00:06:46,420
And so this tension between buying 
yourself a little bit of performance by 

99
00:06:46,420 --> 00:06:50,085
organizing the data in a, in a pretty 
rigid way. 

100
00:06:50,085 --> 00:06:54,045
Versus the development time that you save 
by not having to rewrite your code 

101
00:06:54,045 --> 00:06:59,588
whenever you want to reorganize the data. 
That equation sort of balance a certain 

102
00:06:59,588 --> 00:07:05,280
way in the early 70s and I would claim 
that it still balances the same way now. 

103
00:07:05,280 --> 00:07:08,408
Now, this is not to say that commercial 
databases, relational databases, as they 

104
00:07:08,408 --> 00:07:11,750
stand today, aren't necessarily meeting 
everybody's needs. 

105
00:07:11,750 --> 00:07:15,710
I think they're clearly not for a variety 
of reasons, but to sort of thorough out 

106
00:07:15,710 --> 00:07:18,899
what we know. 
And, you know, give up on this 

107
00:07:18,899 --> 00:07:22,559
flexibility that was earned from the 
relational data model, in favour of 

108
00:07:22,559 --> 00:07:27,555
upfront decisions about one dominant way 
of organizing the data. 

109
00:07:27,555 --> 00:07:31,896
I'm not sure that's the right one either. 
In order to, so that's one point I want 

110
00:07:31,896 --> 00:07:34,825
to make. 
Another point is, that there is not 

111
00:07:34,825 --> 00:07:38,790
necessarily one right way of decomposing 
things, or one right way of even 

112
00:07:38,790 --> 00:07:43,072
evaluating the join even if you support 
them. 

113
00:07:43,072 --> 00:07:46,000
Okay, and there's a point I made before 
but I want to bring it up again in the 

114
00:07:46,000 --> 00:07:49,178
context of NoSQL. 
So, in a, in a pretty classical web 

115
00:07:49,178 --> 00:07:52,289
application scenario, if you want to show 
all comments by a user named Sue 

116
00:07:52,289 --> 00:07:56,037
associated with any blog post by a user 
named Jim. 

117
00:07:56,037 --> 00:08:00,525
There's a couple different ways you could 
do this, you could look up all blog posts 

118
00:08:00,525 --> 00:08:07,030
associated with Jim, and then fetch all 
corresponding comments filtering for Sue. 

119
00:08:07,030 --> 00:08:10,068
Or you could go the other way, you could 
find all comments by Sue and then for 

120
00:08:10,068 --> 00:08:13,870
each of those comments, look up all the 
blog posts by Jim. 

121
00:08:13,870 --> 00:08:16,407
And either one of these ways may or may 
not be available to you if you have 

122
00:08:16,407 --> 00:08:19,270
already organized your data in a 
particular way. 

123
00:08:19,270 --> 00:08:24,000
But even if they are available to you, 
it's not clear which one's the right one. 

124
00:08:24,000 --> 00:08:28,092
And there may even be a third method that 
sounds a little wild but it's filter all 

125
00:08:28,092 --> 00:08:33,128
comments by Sue and then independently 
filter all posts for Jim. 

126
00:08:33,128 --> 00:08:36,910
Sort them, by some sort of blog id, and 
then pull one from each list and walk 

127
00:08:36,910 --> 00:08:40,705
over the data that way. 
So, that sounds a little wacky, but 

128
00:08:40,705 --> 00:08:43,495
that's just a sort merge join that you 
may or may not be familiar with, and if 

129
00:08:43,495 --> 00:08:47,072
you're not that's okay. 
But the point is all three of these are 

130
00:08:47,072 --> 00:08:50,446
perfectly valid. 
And the right one depends on the details 

131
00:08:50,446 --> 00:08:53,930
of the data that you use a programmer, as 
an application programmer may or may not 

132
00:08:53,930 --> 00:08:58,082
have access to and that may change from 
time to time. 

133
00:08:58,082 --> 00:08:59,680
Time. 
And so, really only the system is 

134
00:08:59,680 --> 00:09:04,325
equipped to make this make this decision. 
And so being, you know, sort of 

135
00:09:04,325 --> 00:09:09,752
succumbing to the tyranny of your, of the 
initial design decision at the time the 

136
00:09:09,752 --> 00:09:15,680
database was built or designed is one 
problem. 

137
00:09:15,680 --> 00:09:19,060
Another problem is that even if you have 
some flexibility, leaving it up to the 

138
00:09:19,060 --> 00:09:24,140
programmer to make the choice over the 
right way to access the, access the data. 

139
00:09:24,140 --> 00:09:27,624
is asking him to do, to make a decision 
for which they aren't equipped to, they 

140
00:09:27,624 --> 00:09:31,558
aren't equipped to make. 
Okay, and this is something that does not 

141
00:09:31,558 --> 00:09:35,328
happen with re, relational databases. 
Neither one of these problems exists is 

142
00:09:35,328 --> 00:09:38,768
in the same way. 
Okay, so, I think trying to inject some 

143
00:09:38,768 --> 00:09:43,356
of that smarts back into these NoSQL 
system is, is a good idea and we see that 

144
00:09:43,356 --> 00:09:46,848
trend happening. 
Okay. 

145
00:09:46,848 --> 00:09:50,976
So, that's maybe the takeaway that I want 
you to have. 

146
00:09:50,976 --> 00:09:53,070
Okay. 
So, that's all I want to say about this 

147
00:09:53,070 --> 00:09:56,604
grid, but let me give you a little bit of 
a character of a response to NoSQL by 

148
00:09:56,604 --> 00:10:00,993
Mike Stonebraker, in communications with 
the ASM. 

149
00:10:00,993 --> 00:10:04,946
And there's a couple of different blog 
posts here, that talk about two different 

150
00:10:04,946 --> 00:10:09,374
arguments in favour of NoSQL. 
And they responded to each one of those 

151
00:10:09,374 --> 00:10:13,270
arguments and so I'm mostly going to 
focus on the second one. 

152
00:10:13,270 --> 00:10:16,033
So, there's two pro NoSQL arguments are 
these. 

153
00:10:16,033 --> 00:10:19,841
So [UNKNOWN] points out that there's two 
value propositions offered by the NoSQL 

154
00:10:19,841 --> 00:10:22,726
community. 
One is performance and story as Mike 

155
00:10:22,726 --> 00:10:26,246
says, that I more or less agree with is 
that you know, these people started out 

156
00:10:26,246 --> 00:10:31,249
with a MySQL deployment of some kind. 
And they had a hard time scaling it out 

157
00:10:31,249 --> 00:10:35,110
in a distributed environment and then 
they had two choices. 

158
00:10:35,110 --> 00:10:39,004
You know, either they could invest in a 
large scale relational database and pay 

159
00:10:39,004 --> 00:10:42,780
of course, many license fees or they 
could do something different, use one of 

160
00:10:42,780 --> 00:10:46,326
these NoSQL systems. 
Alright. 

161
00:10:46,326 --> 00:10:51,201
And then the flexibility argument is 
like, well look, my data doesn't conform 

162
00:10:51,201 --> 00:10:56,076
to a rigid schema, so in the performance 
argument, I'm not going to spend too much 

163
00:10:56,076 --> 00:11:01,807
time on this. 
Because he talked about the Trade-offs 

164
00:11:01,807 --> 00:11:07,051
associated with different choices in 
transactional guarantees in these various 

165
00:11:07,051 --> 00:11:11,446
systems. 
And the other part of the argument that 

166
00:11:11,446 --> 00:11:16,045
that Mike makes, has more to do with 
database internals than we'd initially 

167
00:11:16,045 --> 00:11:21,680
covered, so let me focus on the 
flexibility argument. 

168
00:11:21,680 --> 00:11:25,328
So, an observation that he makes that I 
think I probably agree with is that you 

169
00:11:25,328 --> 00:11:29,150
know, who are the customers of these 
NoSQL systems? 

170
00:11:29,150 --> 00:11:31,650
And so it's a lot of start ups, a lot of 
web start ups. 

171
00:11:31,650 --> 00:11:34,935
And it's not quite such the same 
penetration in the enterprise. 

172
00:11:34,935 --> 00:11:36,910
Okay? 
At least not yet. 

173
00:11:36,910 --> 00:11:39,407
So why is this? 
Well, one argument is that most of the 

174
00:11:39,407 --> 00:11:42,862
applications in the enterprise are 
traditional OLTP. 

175
00:11:42,862 --> 00:11:46,139
And if you don't know, OLTP means Online 
Transaction Processing. 

176
00:11:46,139 --> 00:11:49,559
So these are sort of you know, bank 
records alright things were doing the 

177
00:11:49,559 --> 00:11:54,579
transactions right really matters okay. 
And those further results would be on 

178
00:11:54,579 --> 00:11:58,609
structured organized data and so there is 
a few other applications around the 

179
00:11:58,609 --> 00:12:03,210
edges, but they're perhaps considered 
less important. 

180
00:12:03,210 --> 00:12:06,330
Right, it's okay to take a high risk 
system because the application itself is, 

181
00:12:06,330 --> 00:12:09,150
of less interest to, you know, 
executives. 

182
00:12:09,150 --> 00:12:11,714
Okay. 
And so really no, no asset compliance, 

183
00:12:11,714 --> 00:12:18,070
you know, no transactions is equivalent 
to not much interest in the system. 

184
00:12:18,070 --> 00:12:21,791
You know it's not okay to screw up 
mission critical data, but for that, you 

185
00:12:21,791 --> 00:12:25,878
know, for some other application on the 
edge, maybe it's okay to experiment with 

186
00:12:25,878 --> 00:12:30,544
a NoSQL system. 
Okay, and another, and a second point he 

187
00:12:30,544 --> 00:12:34,314
makes is that you know, relying on these 
low level query interfaces is a, a real 

188
00:12:34,314 --> 00:12:38,392
tough cell. 
And, you know, he calls up CODASYL, which 

189
00:12:38,392 --> 00:12:42,336
is an early data manipulation language 
that pre-dated the declarative language 

190
00:12:42,336 --> 00:12:46,467
that we've been talking about. 
But we've sort of been down that road, 

191
00:12:46,467 --> 00:12:49,420
and it's tough, and it's why these high 
level languages were invented. 

192
00:12:49,420 --> 00:12:52,318
And again, we see the, we see it with 
MapReduce, which is middle [INAUDIBLE] 

193
00:12:52,318 --> 00:12:56,059
analytics, as opposed to, to NoSQL. 
But we see that the value of these high 

194
00:12:56,059 --> 00:13:01,512
level interface pays off, okay. 
And the third point he makes is that, you 

195
00:13:01,512 --> 00:13:07,360
know, NoSQL means, we, all bets are off, 
right. 

196
00:13:07,360 --> 00:13:11,204
There's, there's no, there's no sort of 
homogeneity at all between all the 

197
00:13:11,204 --> 00:13:14,358
different deployments. 
And so if, in a, in a typical enterprise 

198
00:13:14,358 --> 00:13:15,838
you have, you know maybe 10,000 
databases. 

199
00:13:15,838 --> 00:13:19,996
You already have enough trouble trying to 
integrate data from these databases 

200
00:13:19,996 --> 00:13:24,027
because of the heterogeneity with in 
their schemas. 

201
00:13:24,027 --> 00:13:28,248
But at least now that your always working 
with rows and columns, and at least have 

202
00:13:28,248 --> 00:13:32,854
kind of a standard interface to 
manipulating them, okay? 

203
00:13:32,854 --> 00:13:35,584
And so having, you know, the number of 
design decisions that you have to make to 

204
00:13:35,584 --> 00:13:38,428
encode your data in one of these new SQL 
systems. 

205
00:13:38,428 --> 00:13:41,483
You know, what becomes the key and what 
becomes the value is, are the blog posts 

206
00:13:41,483 --> 00:13:45,840
nested under the comments or are the 
comments nested under the blog posts. 

207
00:13:45,840 --> 00:13:50,600
Right, do users keep their own wall or do 
the you know, the, the messages on their, 

208
00:13:50,600 --> 00:13:54,952
on their front page or does the person 
who wrote the message keep access to it, 

209
00:13:54,952 --> 00:13:58,000
or both? 
Right. 

210
00:13:58,000 --> 00:14:02,550
All these different design decisions of 
nesting layers and so forth complicate 

211
00:14:02,550 --> 00:14:06,200
integration and complicate standards. 
Okay. 

212
00:14:06,200 --> 00:14:10,316
So, it's a, it's a, it's a tough sale. 
You know, the other, the other point that 

213
00:14:10,316 --> 00:14:13,176
I guess I'd like to make that's related 
to this, this third one here is that, you 

214
00:14:13,176 --> 00:14:17,517
know, there's no real free lunch. 
Either the complexity's going to be in 

215
00:14:17,517 --> 00:14:20,689
the system that you use to model the 
data, or you're going to sort of hand it 

216
00:14:20,689 --> 00:14:25,724
over to the application. 
But in some sense, they're always, these 

217
00:14:25,724 --> 00:14:31,120
schemas, right, this application business 
model is going to be encoded somewhere. 

218
00:14:32,210 --> 00:14:37,247
So, having it centralized in, in the data 
system, as opposed to hidden more than 

219
00:14:37,247 --> 00:14:44,391
once in various applications that access 
the data system, seems like a good idea. 

