1
23:59:59,500 --> 00:00:05,655
[MUSIC]. 

2
00:00:05,655 --> 00:00:10,596
Okay so, Google BigTable again, had a lot 
of influence and one of the major 

3
00:00:10,596 --> 00:00:15,780
outcomes is just like MapReduce it was 
taken up and implemented as an open 

4
00:00:15,780 --> 00:00:22,235
source project called HBase. 
And so it, where it, where BigTable's 

5
00:00:22,235 --> 00:00:26,167
compatible with MapReduce. 
HBase was compatible with Hadoop and in 

6
00:00:26,167 --> 00:00:29,660
one slide there's not much difference 
here. 

7
00:00:29,660 --> 00:00:31,820
But I just want to mention the 
terminology so you've seen it. 

8
00:00:31,820 --> 00:00:36,565
That there's table at the top level and 
then region store and then Mem Store and 

9
00:00:36,565 --> 00:00:40,365
Store File. 
And so the exact names of these 

10
00:00:40,365 --> 00:00:44,964
structures are a little bit different, 
and they did insert one more layer of 

11
00:00:44,964 --> 00:00:49,411
abstraction which is this region. 
Fine. 

12
00:00:49,411 --> 00:00:52,315
And then, how this sort of is compatible 
with Map Reduce is that, you know, each, 

13
00:00:52,315 --> 00:00:56,090
each one of your Map functions will 
process a single tablet. 

14
00:00:56,090 --> 00:01:00,059
And so in, sort of, one to one with the 
blocks of data that we talked about when 

15
00:01:00,059 --> 00:01:04,694
we talked about MapReduce. 
And then, I just sort of ask this 

16
00:01:04,694 --> 00:01:08,765
question, sort of open ended, is that 
there's just no such speculative 

17
00:01:08,765 --> 00:01:13,457
execution in MapReduce that we talked 
about whereby for fault tolerance reasons 

18
00:01:13,457 --> 00:01:21,150
it might kick off the same map cast twice 
on two different replica of the data. 

19
00:01:22,240 --> 00:01:25,730
And the reason for this is that if one 
fails, well you have the other one. 

20
00:01:25,730 --> 00:01:27,088
Right? 
So you just don't start over from 

21
00:01:27,088 --> 00:01:30,641
scratch. 
But, you know, in this environment where 

22
00:01:30,641 --> 00:01:36,720
you're now working on data that is being 
actively updated. 

23
00:01:36,720 --> 00:01:40,450
And that's what HBase and BigTable were 
designed to support is, is updates. 

24
00:01:40,450 --> 00:01:44,542
It's not quite as clear to me what's 
going to happen when you have, you know, 

25
00:01:44,542 --> 00:01:48,700
it's possible because of eventual 
consistency that these two tablets will 

26
00:01:48,700 --> 00:01:53,991
not always agree instantly. 
On the same record they'll agree, but 

27
00:01:53,991 --> 00:01:56,641
different records within the Tablet may 
not. 

28
00:01:56,641 --> 00:01:59,422
Okay. 
And so, just an example of when you sort 

29
00:01:59,422 --> 00:02:04,072
of mix these 2 systems and there are no, 
sort of, system wide transactional 

30
00:02:04,072 --> 00:02:08,620
guarantees. 
Or system wide even properties, that you 

31
00:02:08,620 --> 00:02:12,104
can run into, run into trouble. 
And I think that a general theme here, 

32
00:02:12,104 --> 00:02:15,008
with these, sort of, no sequel systems, 
including the ones that are designed by 

33
00:02:15,008 --> 00:02:19,770
Google, is that you kind of offloading 
some of those responsibility to. 

34
00:02:19,770 --> 00:02:22,476
The application to the program that will 
sort of sort this out and make sure it's 

35
00:02:22,476 --> 00:02:26,175
okay. 
Okay and we're going to come back to that 

36
00:02:26,175 --> 00:02:29,930
in, in just a couple minutes. 
All right. 

37
00:02:29,930 --> 00:02:33,710
So after BigTable, several years later 
there's a paper by a bunch of folks in 

38
00:02:33,710 --> 00:02:37,610
Google about a system called megastore 
that I'm not going to spend a lot of time 

39
00:02:37,610 --> 00:02:41,954
on. 
But it's basically you know, they found 

40
00:02:41,954 --> 00:02:46,370
this point I have sort of just made is 
that these loose consistency models can 

41
00:02:46,370 --> 00:02:52,393
complicate application programming. 
And what they want to do is provide a 

42
00:02:52,393 --> 00:02:57,280
little more system support for certain 
kinds of safe updates. 

43
00:02:57,280 --> 00:03:00,400
Okay, so here instead of full 
transactions being safe within a 

44
00:03:00,400 --> 00:03:03,654
individual record as they are in big 
table. 

45
00:03:03,654 --> 00:03:06,930
They've extended it with this notion of 
Entity Groups. 

46
00:03:06,930 --> 00:03:09,140
And so an Entity Group, it should be on 
this slide. 

47
00:03:09,140 --> 00:03:17,135
An Entity Group is a set of records that 
tend to go together, tend to be accessed 

48
00:03:17,135 --> 00:03:22,810
together. 
Okay. 

49
00:03:22,810 --> 00:03:25,995
so maybe again this is the blog and all 
of its comments for example maybe each 

50
00:03:25,995 --> 00:03:28,984
one of these is a record is a 6 interval 
record store so it's okay for them to 

51
00:03:28,984 --> 00:03:33,826
have different schemas but they all tend 
to go together okay. 

52
00:03:33,826 --> 00:03:38,518
And so what they does extend transaction 
support over an entire entity group, you 

53
00:03:38,518 --> 00:03:44,060
know, a set of, a set of related records. 
Okay, so they still get the scalibility 

54
00:03:44,060 --> 00:03:47,531
by not requiring full system wide global, 
you know, synchrony. 

55
00:03:47,531 --> 00:03:50,931
But they allow you to sort of, they, they 
get away from this problem of you know, 

56
00:03:50,931 --> 00:03:54,131
I, very frequently I might need to update 
one record and then update all of its 

57
00:03:54,131 --> 00:03:58,000
sort of children records at the same 
time. 

58
00:03:58,000 --> 00:03:59,720
And I can't do that any kind of safe way. 
Okay. 

59
00:03:59,720 --> 00:04:05,977
So fine. 
Fast forward one more year Alright, and 

60
00:04:05,977 --> 00:04:10,465
so there's a 2012 paper on a system 
called spanner. 

61
00:04:10,465 --> 00:04:15,547
And I just want to mention these quotes 
and then we'll talk a little bit about 

62
00:04:15,547 --> 00:04:20,475
the, the system, and this one is still 
sort of being explored by the online 

63
00:04:20,475 --> 00:04:25,358
community. 
It's not available actually for use, but 

64
00:04:25,358 --> 00:04:28,193
the paper's being explored, and the 
idea's being explored, so for example, 

65
00:04:28,193 --> 00:04:31,600
you don't see an open source, actually, 
that's not true. 

66
00:04:31,600 --> 00:04:34,460
You do see, there has been a couple open 
source implementations of the ideas in 

67
00:04:34,460 --> 00:04:36,880
Spanner, but they're not quite as 
popular, some of the open source 

68
00:04:36,880 --> 00:04:40,023
Googledations of the other Google 
systems. 

69
00:04:40,023 --> 00:04:42,883
Okay so you know it says that even though 
many projects happily use BigTable we 

70
00:04:42,883 --> 00:04:45,743
have also consistently received 
complaints from users that BigTable can 

71
00:04:45,743 --> 00:04:49,344
be difficult to use for certain kinds of 
applications. 

72
00:04:49,344 --> 00:04:55,124
Those that have complex evolving schemas. 
Or those that want strong consistency in 

73
00:04:55,124 --> 00:04:59,659
the presence of wide area replication, 
okay. 

74
00:04:59,659 --> 00:05:02,101
And then we go on to say, we believe it's 
better to have application programmers 

75
00:05:02,101 --> 00:05:04,247
deal with performance problems due to 
overuse of transactions as the 

76
00:05:04,247 --> 00:05:08,140
bottlenecks arise, rather than always 
coding around the lack of transaction. 

77
00:05:08,140 --> 00:05:12,138
And so this, you know, the database 
community could have said, well sure, 

78
00:05:12,138 --> 00:05:15,830
[LAUGH] you know, well done. 
[LAUGH] Alright. 

79
00:05:15,830 --> 00:05:20,989
That's exactly the point is that system 
supply and support for transactions is, 

80
00:05:20,989 --> 00:05:23,500
is always a win. 
Right. 

81
00:05:23,500 --> 00:05:27,335
Because it's a, it's difficult, error 
prone, expensive to try to do this at the 

82
00:05:27,335 --> 00:05:30,809
application level. 
And, more importantly, it's fundamentally 

83
00:05:30,809 --> 00:05:33,227
wrong, in some sense, to do it at the 
application level because it does, the 

84
00:05:33,227 --> 00:05:36,458
application doesn't have global knowledge 
of what's going on. 

85
00:05:36,458 --> 00:05:39,260
Alright. 
Only the system does. 

86
00:05:39,260 --> 00:05:41,138
So, fine. 
So, although Spanner is scalable in the 

87
00:05:41,138 --> 00:05:44,128
number of nodes the final quote here, the 
node-local data structures have 

88
00:05:44,128 --> 00:05:47,256
relatively poor performance on complex 
SQL queries, because they were designed 

89
00:05:47,256 --> 00:05:51,740
for simple key-value accesses. 
And then, algorithms and data structures 

90
00:05:51,740 --> 00:05:54,300
from the database literature could 
improve single node performance a great 

91
00:05:54,300 --> 00:05:57,141
deal. 
Again, you know, it's, it's, somewhat of 

92
00:05:57,141 --> 00:06:00,651
a, of a Google-style approach to the 
problem of reboot everything, rebuild it 

93
00:06:00,651 --> 00:06:05,379
all from scratch and then sort of 
cherry-pick and bring things in. 

94
00:06:05,379 --> 00:06:07,676
So this has been working pretty well. 
And they have fantastic impact in the 

95
00:06:07,676 --> 00:06:10,170
community. 
but there's a lot out there in the 

96
00:06:10,170 --> 00:06:14,386
database literature and in the database 
system that could have been used from the 

97
00:06:14,386 --> 00:06:18,478
start, in fact trying to start from the 
beginning and just say were going build a 

98
00:06:18,478 --> 00:06:24,898
big Google style parallel database may 
have been a good choice. 

99
00:06:24,898 --> 00:06:28,618
Rather than sort of getting completely 
away from it and then coming back 

100
00:06:28,618 --> 00:06:32,880
incrementally and finding yourself in, in 
a SQL system. 

101
00:06:32,880 --> 00:06:36,576
Now I sort of skipped over what Spanner 
is but it's, it's a planet scale database 

102
00:06:36,576 --> 00:06:40,048
system, there is a SQL like language 
should just go back to our, I mean I know 

103
00:06:40,048 --> 00:06:44,140
what I'm missing. 
I'm missing our, our, our table here. 

104
00:06:44,140 --> 00:06:46,780
Let me flip back a few. 
So here it is down here. 

105
00:06:46,780 --> 00:06:54,870
So I'm missing, I'm missing this slide 
here where I showed it. 

106
00:06:54,870 --> 00:07:01,113
So really big scale. 
Primary accesses you can't access by 

107
00:07:01,113 --> 00:07:04,282
other attributes. 
There are transactions in effect their 

108
00:07:04,282 --> 00:07:06,924
global this time their real, real asset 
transactions. 

109
00:07:06,924 --> 00:07:10,256
it's not clear to me whether joins are 
supported I suspect they are cause if you 

110
00:07:10,256 --> 00:07:14,860
talk about sequel but I couldn't find an 
example of whether there is or not. 

111
00:07:14,860 --> 00:07:18,586
there is a notion of schema and they do 
sort of protect against data that doesn't 

112
00:07:18,586 --> 00:07:21,840
perform to the schema. 
There is some notion of logical data 

113
00:07:21,840 --> 00:07:24,780
independence, although they, they don't 
talk about it much. 

114
00:07:24,780 --> 00:07:27,110
There is a sequel like decorative 
language on top of it. 

115
00:07:27,110 --> 00:07:29,860
I didn't see much evidence that they're 
doing a whole lot of fancy optimization. 

116
00:07:29,860 --> 00:07:32,980
And I did just show you that quote of 
where they say that they performance is 

117
00:07:32,980 --> 00:07:38,300
sort of poor on complex analytic queries. 
Bus as soon as the problem would come 

118
00:07:38,300 --> 00:07:43,310
along somewhere quickly okay. 
So fine, so that's, the spanner a high 

119
00:07:43,310 --> 00:07:46,812
level. 
Let me give you a couple more details 

120
00:07:46,812 --> 00:07:51,947
about what this system does so the data 
model here is discussed in directories 

121
00:07:51,947 --> 00:07:58,560
and these are a set of continuous keys, 
with a shared prefix. 

122
00:07:58,560 --> 00:08:01,784
So you can think of it kind of like a 
tablet was in BigTable, but now they have 

123
00:08:01,784 --> 00:08:06,394
this notion of multiple logical tables 
that are sort of interweaved. 

124
00:08:06,394 --> 00:08:09,760
And so if you're not used staring at this 
syntax, don't worry too much, but those 

125
00:08:09,760 --> 00:08:13,075
of you that, who are thinking in terms of 
DDL in, in a relational database, they 

126
00:08:13,075 --> 00:08:17,880
have kind of a create table language that 
looks like this. 

127
00:08:17,880 --> 00:08:20,640
You create a table, Users, with two 
columns, and then you give it this key 

128
00:08:20,640 --> 00:08:24,165
word, directory. 
And then you create table albums with 

129
00:08:24,165 --> 00:08:29,680
some columns, and you have this key word, 
interleave in parent users. 

130
00:08:29,680 --> 00:08:32,624
And what you end up with is something 
like this, where, there's a user with all 

131
00:08:32,624 --> 00:08:35,620
of its albums and a user with all of its 
albums. 

132
00:08:35,620 --> 00:08:38,162
As you can see here that, you know, what 
we've been talking about, all these 

133
00:08:38,162 --> 00:08:40,786
different systems are experimenting with 
ways of getting these nested data 

134
00:08:40,786 --> 00:08:45,275
structures. 
Hierarchical data structures that look a 

135
00:08:45,275 --> 00:08:49,736
lot like what we saw way back in the 60s, 
right? 

136
00:08:49,736 --> 00:08:53,610
And they're motivation is the same as it 
was then. 

137
00:08:53,610 --> 00:08:57,055
It's actually really really fast. 
When you're going to access, when you 

138
00:08:57,055 --> 00:09:02,685
want to pull up a user, and then 
immediately pull up all of its album. 

139
00:09:02,685 --> 00:09:06,030
It's really fast axis to this this way 
right? 

140
00:09:06,030 --> 00:09:11,196
you know, but I probably speculate that 
the reasons why relational, the 

141
00:09:11,196 --> 00:09:16,104
relational approach eventually replaced 
these. 

142
00:09:16,104 --> 00:09:21,270
And, and what can and will happen here as 
well, is that performance is not the 

143
00:09:21,270 --> 00:09:28,856
number one priority, it's minimizing the 
amount of developer headaches. 

144
00:09:28,856 --> 00:09:31,936
Okay. 
So, it remains to be seen, but I, but I 

145
00:09:31,936 --> 00:09:36,243
think that this, this incremental walk 
step towards a big new scalable 

146
00:09:36,243 --> 00:09:41,906
relational database is, is, is underway. 
Now again, that doesn't mean that I'm 

147
00:09:41,906 --> 00:09:44,382
saying, use all the old databases. 
They really were designed for, sort of, a 

148
00:09:44,382 --> 00:09:45,902
different work load and they really 
don't. 

149
00:09:45,902 --> 00:09:48,002
There is really no evidence on these 
scale on some of these, some of these 

150
00:09:48,002 --> 00:09:50,242
levels, but that doesn't mean that you're 
sort of throw out all the, you know, 

151
00:09:50,242 --> 00:09:59,000
everything we learn, okay. 
But that's more me editorializing. 

152
00:09:59,000 --> 00:10:02,472
So fine, how this work is there's a 
universe master at the very, very top and 

153
00:10:02,472 --> 00:10:05,835
this is just a singleton. 
There's only one of these for a 

154
00:10:05,835 --> 00:10:08,254
deployment and they sort of imagine there 
is only one or two of these for 

155
00:10:08,254 --> 00:10:11,042
deployments anywhere so they have sort of 
one for test, one for production/test, 

156
00:10:11,042 --> 00:10:14,750
and one for production. 
And that's it. 

157
00:10:14,750 --> 00:10:17,861
So all, so, so, many different 
applications will use the same deployment 

158
00:10:17,861 --> 00:10:20,706
of, of spanner. 
And so, this is mostly just status about 

159
00:10:20,706 --> 00:10:23,924
status information about the zones. 
It doesn't need a, it doesn't interact 

160
00:10:23,924 --> 00:10:27,522
with clients at all. 
Then there's a placement driver that's 

161
00:10:27,522 --> 00:10:32,390
responsible for moving these directory 
sets of records. 

162
00:10:32,390 --> 00:10:36,660
Around for load balancing purposes, and 
this happened on a scale of, of every few 

163
00:10:36,660 --> 00:10:40,106
minutes, alright. 
And then within a zone there's a, a zone 

164
00:10:40,106 --> 00:10:44,248
master that assigns data to spanservers. 
And there's a location proxy that sort of 

165
00:10:44,248 --> 00:10:48,882
knows where everything is. 
And routes requests to the appropriate 

166
00:10:48,882 --> 00:10:51,772
spanserver. 
And the spanservers themselves serve 

167
00:10:51,772 --> 00:10:56,445
data, and so in here it's starting to 
look a little more like BigTable. 

168
00:10:56,445 --> 00:11:05,640
You know, a zone is essentially a, an 
individual BigTable deployment. 

169
00:11:05,640 --> 00:11:08,540
Okay? 
So inside of a spanserver this is where 

170
00:11:08,540 --> 00:11:13,028
the, the big difference here is this is 
where they're going to try to support 

171
00:11:13,028 --> 00:11:20,218
fully consistent transactions. 
So, across these, you know within a group 

172
00:11:20,218 --> 00:11:27,759
of these replicas. 
They can, they support 2 phase commit. 

173
00:11:27,759 --> 00:11:32,179
This only is needed when a transaction 
actually accesses data that's, that's you 

174
00:11:32,179 --> 00:11:36,274
know in that's across the, is not, you're 
not constrained in one particular 

175
00:11:36,274 --> 00:11:39,280
replica. 
Okay. 

176
00:11:39,280 --> 00:11:44,310
Other than that, it just skips over this, 
this logic, and it doesn't cost anything. 

177
00:11:44,310 --> 00:11:46,006
Okay. 
And then one step down, below this, 

178
00:11:46,006 --> 00:11:50,200
across so this is, sorry, I guess I'm 
using the wrong terminology. 

179
00:11:50,200 --> 00:11:52,960
So it should basically come in as across 
groups, and when all of them, the 

180
00:11:52,960 --> 00:11:55,858
transaction and all these contained in 
one single group then you drop down a 

181
00:11:55,858 --> 00:11:58,756
level and you run the Paxos algorithm 
that I didn't talk about in detail but I 

182
00:11:58,756 --> 00:12:04,939
mentioned exists. 
In order to sort out the reads and writes 

183
00:12:04,939 --> 00:12:08,072
for, in order to handle the write. 
Okay? 

184
00:12:08,072 --> 00:12:14,120
And the only other piece I'll mention 
here is that this term colossus is new. 

185
00:12:14,120 --> 00:12:19,856
It's the successor to Google File System. 
And Google File System is the- original 

186
00:12:19,856 --> 00:12:23,618
turn for the op, you know, the open 
sourcing limitation of HTFS which 

187
00:12:23,618 --> 00:12:28,473
underlies Map Produce and Hadoop. 
Sorry. 

188
00:12:28,473 --> 00:12:31,675
GFS is to Map Produce as HTFS is to 
Hadoop. 

189
00:12:31,675 --> 00:12:36,130
So when I'm throwing these acronyms at 
you, that's how to keep it straight. 

190
00:12:36,130 --> 00:12:42,460
Okay, so that's all I want to say about 
spanner in particular. 

191
00:12:42,460 --> 00:12:44,968
Let's just take a step back and look at 
all the different systems that Google has 

192
00:12:44,968 --> 00:12:47,664
for a second. 
You know, map reduce was a paper in 2004 

193
00:12:47,664 --> 00:12:50,719
that had a ton of impact, BigTable had a 
ton of impact, then there's Megastore, 

194
00:12:50,719 --> 00:12:53,727
there's this tens thing that we didn't 
talk about, but it's a SQL system on top 

195
00:12:53,727 --> 00:12:56,782
of mat produce, much like hive, if your 
familiar with that or if you've heard me 

196
00:12:56,782 --> 00:13:01,470
mention it, and then spanner very 
recently. 

197
00:13:01,470 --> 00:13:04,440
And so you can or sort of organize things 
into a timeline this way, just to kind of 

198
00:13:04,440 --> 00:13:07,708
get a sense of this. 
And because of these systems have had so 

199
00:13:07,708 --> 00:13:10,498
much influence I want you to be aware of 
what they are and sort of how they fit 

200
00:13:10,498 --> 00:13:15,210
together so it doesn't just sound like a 
big jumble of terms. 

201
00:13:15,210 --> 00:13:20,030
So MapReduce was, you know, the, one of 
the earliest ones. 

202
00:13:20,030 --> 00:13:23,780
It wasn't quite the earliest. 
There was actually another one called 

203
00:13:23,780 --> 00:13:28,728
Sawzall, that really didn't get a ton of 
traction, but it was a nice paper. 

204
00:13:28,728 --> 00:13:32,440
and then BigTable came a couple years 
later, and I drew a dotted line there 

205
00:13:32,440 --> 00:13:37,360
representing that there's sort of 
compatible design to go together. 

206
00:13:37,360 --> 00:13:40,840
One was map produced for analytics, 
BigTable is for the, sort of micro 

207
00:13:40,840 --> 00:13:44,261
operations. 
And then both of these, a few years later 

208
00:13:44,261 --> 00:13:48,720
have an open source limitation in Hadoop 
and H-base respectively. 

209
00:13:48,720 --> 00:13:52,437
Fast forward a few more years, and you 
got a mega store in spanner coming very 

210
00:13:52,437 --> 00:13:57,170
quickly one right after the other. 
And this hard, this heavy blue line 

211
00:13:57,170 --> 00:14:01,960
represents you know it's pretty clear 
that he influence is fairly direct. 

212
00:14:01,960 --> 00:14:04,217
In fact, I would suspect that there's a 
lot of code being borrowed and then 

213
00:14:04,217 --> 00:14:06,770
megastore makes plenty of references to 
BigTable and spanner makes references to 

214
00:14:06,770 --> 00:14:11,015
both megastore and, and BigTable. 
And they, the papers have many, many of 

215
00:14:11,015 --> 00:14:14,652
the same co-authors. 
Okay, and then MapReduce depends directly 

216
00:14:14,652 --> 00:14:18,760
on, I'm sorry, excuse me, Tenzing depends 
directly on MapReduce. 

217
00:14:18,760 --> 00:14:21,230
It provides a sequel layer on top of 
MapReduce. 

218
00:14:21,230 --> 00:14:24,225
And then there's some other systems here. 
One is called Dremel which was originally 

219
00:14:24,225 --> 00:14:27,458
for very fast aggregate queries but 
really just aggregate queries but of 

220
00:14:27,458 --> 00:14:31,299
extremely low latency. 
So this is you know, in the analytics 

221
00:14:31,299 --> 00:14:34,638
camp cause your doing these sort of 
aggregate questions as opposed to a sort 

222
00:14:34,638 --> 00:14:39,599
of micro updates. 
but it was extremely low latentcy unlike 

223
00:14:39,599 --> 00:14:43,190
map produce it was more of a batch system 
and so this is this the a great fit and 

224
00:14:43,190 --> 00:14:48,278
its a very nice system. 
And in fact, since they then can do joins 

225
00:14:48,278 --> 00:14:52,502
not just aggregates and more importantly 
this was exposed as a service that you 

226
00:14:52,502 --> 00:14:58,920
can just use directly over the web, even 
in your browser, called Big Query. 

227
00:14:58,920 --> 00:15:02,430
And that's a, that's a, that's an 
important word to watch. 

228
00:15:02,430 --> 00:15:05,875
It's one of the few systems that is do, 
available as a service through, that let 

229
00:15:05,875 --> 00:15:10,170
you, as a cloud service. 
but scales a very, very large data, and 

230
00:15:10,170 --> 00:15:13,470
sports analytics, okay. 
And, then, another one that we'll, we'll 

231
00:15:13,470 --> 00:15:16,870
talk about yeah. 
But we'll come back to is Pregel. 

232
00:15:16,870 --> 00:15:22,360
And this adds the one secret ingredient 
that I, is sort of near and dear to my 

233
00:15:22,360 --> 00:15:27,326
heart which is iteration. 
And what I mean by that is when you, when 

234
00:15:27,326 --> 00:15:30,370
you run MapReduce jobs and do analytics 
you're sort of taking step one. 

235
00:15:30,370 --> 00:15:31,820
And then step two. 
And then step three. 

236
00:15:31,820 --> 00:15:34,348
And you stop. 
But for many kinds of tasks, especially 

237
00:15:34,348 --> 00:15:38,790
in data science, many of these analytics 
tasks, these machine learning tasks. 

238
00:15:38,790 --> 00:15:41,030
You have to do something again and again 
and again and again until some kind of 

239
00:15:41,030 --> 00:15:45,373
convergence condition is reached. 
And Pregel and one of our systems and a 

240
00:15:45,373 --> 00:15:50,104
few other systems, are the ones that 
tried to, you know, notices this 

241
00:15:50,104 --> 00:15:56,340
limitation of map reduce and extend it. 
So people were doing this with map 

242
00:15:56,340 --> 00:15:59,050
reduce, but they would do it sort of in 
fairly ad hoc ways. 

243
00:15:59,050 --> 00:16:02,795
Okay. 
And so we'll come back to that and talk 

244
00:16:02,795 --> 00:16:09,540
about it but, you know, analytics, low 
latency micro-updates. 

245
00:16:09,540 --> 00:16:13,415
So those two big classes of systems. 
And then analytics with iteration is 

246
00:16:13,415 --> 00:16:18,649
perhaps a third class of system that we, 
that we'll talk about. 

