1
23:59:59,650 --> 00:00:05,671
[MUSIC]. 

2
00:00:05,671 --> 00:00:08,682
Okay. 
So, the next few segments I want to spend 

3
00:00:08,682 --> 00:00:11,950
talking about NoSqL. 
And so, these systems are typically 

4
00:00:11,950 --> 00:00:16,110
associated with building very large 
scalable web applications as opposed to 

5
00:00:16,110 --> 00:00:19,549
analyzing data. 
Which is really the focus of this course. 

6
00:00:19,549 --> 00:00:23,370
However, I think it's important to cover 
this topic for a few reasons. 

7
00:00:23,370 --> 00:00:26,725
One is, you know, as a data scientist, 
you'll be manipulating data that is 

8
00:00:26,725 --> 00:00:31,029
increasingly found in one of these 
variants of a NoSQL system. 

9
00:00:31,029 --> 00:00:34,149
but, but also the, the systems and the 
terminology in this space are really 

10
00:00:34,149 --> 00:00:38,182
influencing peoples thinking about how to 
deal with large scale data. 

11
00:00:38,182 --> 00:00:41,494
And so, being cognizant of the landscape 
here, being cognizant of the major trends 

12
00:00:41,494 --> 00:00:44,999
in history allow you to make informed 
decisions about this. 

13
00:00:44,999 --> 00:00:48,331
You know, as a data scientist, you may be 
asked to make recommendations about what 

14
00:00:48,331 --> 00:00:52,691
kind of platform to go with, to do your 
kind of ana, to do your analysis. 

15
00:00:52,691 --> 00:00:55,715
And so, understanding the pros and cons 
of some of these systems and how they 

16
00:00:55,715 --> 00:00:58,830
work can be pretty important, okay? 
You know? 

17
00:00:58,830 --> 00:01:02,592
And then maybe third, the same concepts 
that we've been discussing in other 

18
00:01:02,592 --> 00:01:06,156
segments, you know, sort of relational 
algebra. 

19
00:01:06,156 --> 00:01:10,367
logical data dependence, you know, simple 
analytics going a long way. 

20
00:01:10,367 --> 00:01:15,040
Simple scalable analytics indexing these 
ideas come up in this space as well. 

21
00:01:15,040 --> 00:01:18,010
And so, it's yet another application of 
these concepts. 

22
00:01:18,010 --> 00:01:21,358
And then maybe, finally the data 
scientists may be called upon to actually 

23
00:01:21,358 --> 00:01:25,300
build some of the large scale systems as 
part of their work. 

24
00:01:25,300 --> 00:01:28,638
It's not unheard of. 
When I mention in the first few segments 

25
00:01:28,638 --> 00:01:32,630
that building data products is important 
part of data science. 

26
00:01:32,630 --> 00:01:36,599
And so, when manifestation of data 
products is some of these large scale 

27
00:01:36,599 --> 00:01:41,365
well application, so NoSQUl may or may 
not be a part of that. 

28
00:01:41,365 --> 00:01:45,259
Fine, so let's get started. 
Alright, so where we are so far is we 

29
00:01:45,259 --> 00:01:48,899
mentioned that data science may be has 
these three steps, data preparation at 

30
00:01:48,899 --> 00:01:53,139
large scale, alright? 
And we talked about manipulation and 

31
00:01:53,139 --> 00:01:57,780
munging, and data jujitsu, and so on. 
And that's all these step one. 

32
00:01:57,780 --> 00:02:00,510
And then step two is analytics or 
actually running the model. 

33
00:02:00,510 --> 00:02:03,380
And we're going to get there at, in the 
next week. 

34
00:02:03,380 --> 00:02:05,492
And then third is communicating the 
results, interpreting and communicating 

35
00:02:05,492 --> 00:02:07,398
the results. 
And this will be visualization will be a 

36
00:02:07,398 --> 00:02:09,530
big part of the way we talk about 
communication, okay? 

37
00:02:09,530 --> 00:02:13,100
So, we're still in the, this data 
preparation phase. 

38
00:02:13,100 --> 00:02:15,284
And we're spending a little bit longer on 
that. 

39
00:02:15,284 --> 00:02:18,769
Then, you know, say a third of the 
course, more than a third of the course. 

40
00:02:18,769 --> 00:02:21,851
For the reason that I gave in the first 
few segments, that, really, this is the 

41
00:02:21,851 --> 00:02:25,917
part that keeps people up at night. 
So, I want to make sure that you're armed 

42
00:02:25,917 --> 00:02:28,840
with how to use databases to do this data 
munging. 

43
00:02:28,840 --> 00:02:31,190
How to use MapReduce to do this data 
munging. 

44
00:02:31,190 --> 00:02:35,066
also make the point that a lot of times 
even analytics itself can be pushed into 

45
00:02:35,066 --> 00:02:38,714
these systems, which you saw in, in 
hopefully in, in assignment two in the 

46
00:02:38,714 --> 00:02:43,358
database assignment, okay? 
So then, some of the key ideas from 

47
00:02:43,358 --> 00:02:46,394
databases that we took away where this 
concept of the relation algebra comes up 

48
00:02:46,394 --> 00:02:51,935
even extern, even outside of databases. 
It's not only found, you know, with SQL 

49
00:02:51,935 --> 00:02:55,130
systems. 
And we'll see that again. 

50
00:02:55,130 --> 00:03:01,763
And this notion of physical logical data 
independence comes up over and over 

51
00:03:01,763 --> 00:03:07,299
again, maybe indexing. 
And then we talked about MapReduce, and 

52
00:03:07,299 --> 00:03:11,520
we gave a lot of examples of some basic 
operations of MapReduce. 

53
00:03:11,520 --> 00:03:15,510
And we started an assignment involving 
writing your own MapReduce programs, at 

54
00:03:15,510 --> 00:03:20,201
least at the programming model level. 
And so, here we saw that part of the 

55
00:03:20,201 --> 00:03:24,553
advantages of this system itself first at 
Google, and then in the implementation 

56
00:03:24,553 --> 00:03:29,828
Hadoop was volt tolerance at scale. 
Another one was you didn't have to load 

57
00:03:29,828 --> 00:03:33,298
the data. 
You could just work with it as is unlike 

58
00:03:33,298 --> 00:03:36,238
data bases where you have to actually 
impose some sort of a scheme on it in 

59
00:03:36,238 --> 00:03:39,620
order to even touch it for the first 
time. 

60
00:03:39,620 --> 00:03:42,550
And so, that first step can be a dozy, 
okay? 

61
00:03:42,550 --> 00:03:44,780
Extending this point, you're direct 
programming on in situ data. 

62
00:03:44,780 --> 00:03:46,985
So, anything that comes at you, you can 
sort of, if you can write a program and 

63
00:03:46,985 --> 00:03:49,050
process it. 
You can probably write a MapReduce 

64
00:03:49,050 --> 00:03:52,568
program that can process it at scale. 
And that's a very powerful thing. 

65
00:03:52,568 --> 00:03:59,400
I'll give you another way to put this is 
sort of single developer, right? 

66
00:03:59,400 --> 00:04:03,018
You're kind of up and running within the 
hour, with MapReduce, and that was never 

67
00:04:03,018 --> 00:04:06,627
really a property. 
The databases had, it was always sort of 

68
00:04:06,627 --> 00:04:10,353
a difficult, it was a significant project 
to get a database installed and running, 

69
00:04:10,353 --> 00:04:14,918
fine. 
So that the background, but we haven't 

70
00:04:14,918 --> 00:04:19,890
talked about is this NoSQL system. 
And so, I want to use this table as a way 

71
00:04:19,890 --> 00:04:24,450
to sort of organize the road map for this 
discussion. 

72
00:04:24,450 --> 00:04:28,403
And so, what I've done here is tried to 
list out by feature, a bunch of relevant 

73
00:04:28,403 --> 00:04:32,912
systems in this space. 
and right now, they're sort of sorted by 

74
00:04:32,912 --> 00:04:35,638
time. 
And so, these features are admittedly 

75
00:04:35,638 --> 00:04:38,785
selected by me for what the important 
ones are. 

76
00:04:38,785 --> 00:04:42,325
But I don't think they'd be too 
controversial, and I don't think there's 

77
00:04:42,325 --> 00:04:49,270
anything obviously missing here either. 
Okay, so going through these briefly. 

78
00:04:49,270 --> 00:04:53,630
And a major one here is that it needs to 
scale to sort of thousands of machines. 

79
00:04:53,630 --> 00:04:56,030
Then, there are some need to look up by a 
primary index. 

80
00:04:56,030 --> 00:04:58,217
What I mean by that is by some key value, 
right? 

81
00:04:58,217 --> 00:05:00,492
You can look up by some record ident, 
identifier. 

82
00:05:00,492 --> 00:05:04,320
And then another feature that they may or 
may not have is the look at my secondary 

83
00:05:04,320 --> 00:05:07,090
indexes. 
So, I mean here you can look up by some 

84
00:05:07,090 --> 00:05:10,267
attribute that is not that key, okay? 
So, databases have this. 

85
00:05:10,267 --> 00:05:13,127
For example, you can build an index on 
any attribute you want and the optimizer 

86
00:05:13,127 --> 00:05:17,235
will take advantage of it. 
the third one, and this is the one we'll 

87
00:05:17,235 --> 00:05:21,330
spend a lot of time on, it's sort of 
pretty fundamental to the motivation for 

88
00:05:21,330 --> 00:05:26,578
y nodes equal systems. 
Sort of earn their own moniker or, or are 

89
00:05:26,578 --> 00:05:31,608
different or, trans, transactions, okay? 
And you can also see that there's some 

90
00:05:31,608 --> 00:05:34,330
complexity here that I'll try to explain 
as we go through this. 

91
00:05:34,330 --> 00:05:39,460
So, it's not just a yes or no. 
It's kind of been a case by case, okay? 

92
00:05:39,460 --> 00:05:43,461
And then, this field is whether or not 
these system support, essentially join. 

93
00:05:43,461 --> 00:05:46,725
But I generalize that to analytic, and 
I'm perhaps guilty of making these two 

94
00:05:46,725 --> 00:05:50,325
things almost synonymous. 
If you, if you can do some, if you can do 

95
00:05:50,325 --> 00:05:53,640
joins, then there's a whole space of 
different kind of analytic things you can 

96
00:05:53,640 --> 00:05:56,634
do. 
And if you can't do join, then you're 

97
00:05:56,634 --> 00:06:00,900
leaving all that work up to the client 
and all you can do is retrieve data. 

98
00:06:00,900 --> 00:06:06,048
And so, that's really cannot do join a 
key indicator of how much computation you 

99
00:06:06,048 --> 00:06:10,320
can push into the system. 
And how much you have to sort of bring 

100
00:06:10,320 --> 00:06:14,968
the data back out to, to the client. 
Okay, and then integrity constraints. 

101
00:06:14,968 --> 00:06:19,429
I debate about calling these schemas of 
integrity constraints. 

102
00:06:19,429 --> 00:06:23,707
But as I think, we'll see arguably some 
of these systems do indeed have a schema, 

103
00:06:23,707 --> 00:06:27,848
but may or may not actually enforce that 
schema. 

104
00:06:27,848 --> 00:06:32,937
And so, it's just to avoid the confusion, 
I'm going to call this integrity, okay? 

105
00:06:32,937 --> 00:06:36,204
So, this is kind of hard schemas, if you 
will, okay? 

106
00:06:36,204 --> 00:06:40,296
And in views, which were if you remember 
is I'm a declaring that to be synonymous 

107
00:06:40,296 --> 00:06:45,846
with logical data and dependence. 
And then finally, is there some sort of 

108
00:06:45,846 --> 00:06:49,014
decorative language or algebraic way of 
programming this thing or is it really 

109
00:06:49,014 --> 00:06:52,360
just sort of a low level simple operation 
API? 

110
00:06:52,360 --> 00:06:57,040
Okay. 
All right, so a couple of caveats here. 

111
00:06:57,040 --> 00:07:00,640
One is I haven't included any parallel 
data bases on this list at all. 

112
00:07:00,640 --> 00:07:04,200
Although there's absolutely no reason why 
you couldn't include them here. 

113
00:07:04,200 --> 00:07:06,477
They'd tend to have a lot of check marks 
across this row but since the focus is 

114
00:07:06,477 --> 00:07:10,690
NoSQL systems, I am leaving those out. 
And then the other category I made is 

115
00:07:10,690 --> 00:07:15,402
that individuals sales in this table may 
be debatable depending on how you 

116
00:07:15,402 --> 00:07:21,189
interpret the columns. 
So, this isn't necessarily meant to be 

117
00:07:21,189 --> 00:07:25,789
Hard and fast rules. 
But again, I don't think they'll be 

118
00:07:25,789 --> 00:07:30,988
wildly controversial either, okay? 
So, one of the first stories you can tell 

119
00:07:30,988 --> 00:07:34,770
by staring at this table is that 
relational databases, you know, have been 

120
00:07:34,770 --> 00:07:39,661
around for quite a while. 
And have check marks sort of everywhere 

121
00:07:39,661 --> 00:07:43,607
and all these new features. 
Except they weren't really ever shown to 

122
00:07:43,607 --> 00:07:46,820
scale to lots and lots and lots of 
computers, right? 

123
00:07:46,820 --> 00:07:50,960
Everything was sort of in the order of 
tens of machines, okay? 

124
00:07:50,960 --> 00:07:55,248
So, why don't they scale? 
Well, we talked in the MapReduce segment 

125
00:07:55,248 --> 00:08:00,330
about sort of re-performance and related, 
and sort of analyzing experimental 

126
00:08:00,330 --> 00:08:06,293
results from a paper in 2009. 
Comparing sort of the benefits and 

127
00:08:06,293 --> 00:08:12,350
strengths of MapReduce versus databases. 
So, maybe on re-performance there's an 

128
00:08:12,350 --> 00:08:17,477
argument to be made that they did scale. 
But one, one area where they certainly 

129
00:08:17,477 --> 00:08:21,702
didn't or least certainly weren't shown 
to scale it to, to this level was in 

130
00:08:21,702 --> 00:08:25,082
updates. 
So not just the read, not just the 

131
00:08:25,082 --> 00:08:28,220
analytic workload, but the transaction 
processing workloads. 

132
00:08:29,290 --> 00:08:29,560
Alright. 

