1
00:00:00,001 --> 00:00:06,098
[MUSIC]. 

2
00:00:06,098 --> 00:00:08,739
Welcome back to introduction to data 
science. 

3
00:00:08,739 --> 00:00:12,683
So in this segment I want to talk about 
these four dimensions that I introduced 

4
00:00:12,683 --> 00:00:16,975
last time and I want to justify the first 
three of them. 

5
00:00:16,975 --> 00:00:19,887
And we'll talk about this one next time, 
okay. 

6
00:00:19,887 --> 00:00:23,251
And by justifying what I mean is I 
want to explain why I, I've positioned 

7
00:00:23,251 --> 00:00:28,955
the needles the the the way I you know, 
at the locations I have for this course. 

8
00:00:28,955 --> 00:00:31,546
Okay. 
So the first point here is is this 

9
00:00:31,546 --> 00:00:36,358
dimension of tools versus abstractions. 
And this may seem sort of obvious that we 

10
00:00:36,358 --> 00:00:40,290
want to focus on fundamental concepts as 
opposed to specific tools. 

11
00:00:40,290 --> 00:00:42,738
But I can appreciate the people that are 
taking this course and many other courses 

12
00:00:42,738 --> 00:00:44,880
really want a sort of hands-on 
experience. 

13
00:00:44,880 --> 00:00:47,670
And we're definitely going to try to 
strike a balance, but let me try motivate 

14
00:00:47,670 --> 00:00:50,980
why I think its important to sort of 
focus on this angle. 

15
00:00:50,980 --> 00:00:55,010
and to do this, let me tell one, one 
particular story that you see happen 

16
00:00:55,010 --> 00:00:59,766
through the time and time again. 
So in this case we're talking about sort 

17
00:00:59,766 --> 00:01:03,675
of databases and what is currently going 
on in the no SQL systems. 

18
00:01:03,675 --> 00:01:08,030
Alright so, before 2004, you had, you 
know, the big three relational database 

19
00:01:08,030 --> 00:01:13,670
vendors, plus some open source solutions 
like MySQL and Postgres SQL. 

20
00:01:13,670 --> 00:01:18,320
And then, arguably a big event in 2004 
was when Jeff Dean and His colleagues 

21
00:01:18,320 --> 00:01:22,670
published his paper on MapReduce at 
Google. 

22
00:01:22,670 --> 00:01:26,310
And if you haven't heard about MapReduce, 
we'll talk about it at length and if you 

23
00:01:26,310 --> 00:01:31,090
have, bear with me. 
So, this was great. 

24
00:01:31,090 --> 00:01:35,445
What this will allow you to do is process 
very very large data sets and it sort of 

25
00:01:35,445 --> 00:01:40,390
rebooted the database feature set. 
So it really stuck on. 

26
00:01:40,390 --> 00:01:43,501
It really focused on just scale out 
parallelism and that's it, none of the 

27
00:01:43,501 --> 00:01:47,386
other features of databases. 
And this was exciting to a lot of people 

28
00:01:47,386 --> 00:01:50,937
because they didn't have to sort of deal 
with the extra features that they didn't 

29
00:01:50,937 --> 00:01:54,976
need in databases. 
Nor pay for the exorbitant license fees 

30
00:01:54,976 --> 00:01:59,190
associated with databases. 
So this seemed like, boy, this is the 

31
00:01:59,190 --> 00:02:01,530
right solution. 
Okay. 

32
00:02:01,530 --> 00:02:05,007
And you know, a few, it took a few years, 
but a few years later you had an open 

33
00:02:05,007 --> 00:02:08,598
source implementation of the ideas in 
this paper called Hadoop, led by some 

34
00:02:08,598 --> 00:02:13,570
folks at Yahoo. 
Now, even in the same year, one of the 

35
00:02:13,570 --> 00:02:18,178
earliest and most successful projects 
within the Hadoop ecosystem was this 

36
00:02:18,178 --> 00:02:23,580
system called PIG? 
And what PIG essentially was, was a 

37
00:02:23,580 --> 00:02:28,550
relational algebra programming 
environment or a dupe. 

38
00:02:28,550 --> 00:02:32,246
And if you haven't heard a relation 
algebra don't worry, we'll talk about it. 

39
00:02:32,246 --> 00:02:35,282
But let me convince you that rela-, 
notice our relational, relational algebra 

40
00:02:35,282 --> 00:02:38,175
is the secret sauce within relational 
databases. 

41
00:02:38,175 --> 00:02:41,892
And, so very really on project their that 
was deemed necessary in the Hadoop 

42
00:02:41,892 --> 00:02:45,262
community. 
And was wildly successful, was to have 

43
00:02:45,262 --> 00:02:49,657
relational style programming on top of 
this non-relational system. 

44
00:02:49,657 --> 00:02:54,193
Okay, moreover you had other competitive, 
competing products like DryadLINQ, which 

45
00:02:54,193 --> 00:02:59,109
well, well Dryad and then DryadLINQ which 
is an interface of Dryad. 

46
00:02:59,109 --> 00:03:03,456
which also provided a relational 
algebra-oriented programming environment 

47
00:03:03,456 --> 00:03:07,460
for large scale Parallel data processing 
applications. 

48
00:03:07,460 --> 00:03:10,815
Then you had people literally put the 
language SQL on top of Hadoop. 

49
00:03:10,815 --> 00:03:13,461
So instead of just the underlying 
formulas, it literally had the 

50
00:03:13,461 --> 00:03:18,572
programming language you needed. 
Then a bit later you had indexing for 

51
00:03:18,572 --> 00:03:26,680
Hadoop, which is another feature that 
databases have that we'll talk about. 

52
00:03:26,680 --> 00:03:31,153
You had people talk about schemes and 
more sophisticated types of indexing 

53
00:03:31,153 --> 00:03:35,258
which is two other things the databases 
have. 

54
00:03:35,258 --> 00:03:39,093
And then you start now to see transaction 
processing being a very hot, very 

55
00:03:39,093 --> 00:03:42,928
important topic in new SQL systems is how 
to support concurrent access at very 

56
00:03:42,928 --> 00:03:48,052
large scale transparently. 
And this slide is perhaps a little bit 

57
00:03:48,052 --> 00:03:51,168
old, it's not 2013 at the of this 
recording. 

58
00:03:51,168 --> 00:03:55,003
And the Spanner system from Google is an 
important system to look at that we'll 

59
00:03:55,003 --> 00:03:58,704
talk a little bit about later. 
Okay so, now this isn't to say that 

60
00:03:58,704 --> 00:04:00,979
MapReduce was useless, we're going to 
talk about at length, and for a very good 

61
00:04:00,979 --> 00:04:04,902
reason. 
It actually has some pretty important 

62
00:04:04,902 --> 00:04:10,600
permanent contributions three of which I 
mentioned here. 

63
00:04:10,600 --> 00:04:14,380
One is, you know, it was the first system 
to really emphasize fault tolerance. 

64
00:04:14,380 --> 00:04:18,120
And the idea here, in a nutshell, is that 
when you're working with a 1000 computers 

65
00:04:18,120 --> 00:04:22,905
at a time, for any length of time at all, 
a few minutes, a few hours. 

66
00:04:22,905 --> 00:04:26,760
The odds of one of them failing in some 
way is extremely high. 

67
00:04:26,760 --> 00:04:29,322
And, so, databases didn't typically have 
to worry about this because the 

68
00:04:29,322 --> 00:04:31,968
assumption, first of all they weren't 
running on thousands of computers at 

69
00:04:31,968 --> 00:04:34,713
once. 
And second of all they were sort of under 

70
00:04:34,713 --> 00:04:38,160
the assumption that your queries would 
typically be pretty fast, okay? 

71
00:04:38,160 --> 00:04:41,301
and, so, fault tolerance, during query 
processing. 

72
00:04:41,301 --> 00:04:45,006
So you don't lose all the work you, you, 
you've started on when you're running a 

73
00:04:45,006 --> 00:04:48,711
query, was something that the Hadoop and 
the MapReduce produced paper really 

74
00:04:48,711 --> 00:04:52,913
emphasized. 
And has now been sort of accepted by the 

75
00:04:52,913 --> 00:04:56,812
large community. 
The other notion which is a little more 

76
00:04:56,812 --> 00:05:01,091
subtle is this idea of schema-on-reed. 
And what I mean by that is you know the 

77
00:05:01,091 --> 00:05:04,985
way databases worked in the past and 
largely still work is. 

78
00:05:04,985 --> 00:05:08,625
You know we design the thing called the 
schema which is a particular structure 

79
00:05:08,625 --> 00:05:12,013
for new data. 
And then you job is to fix your data into 

80
00:05:12,013 --> 00:05:16,470
that schema and until you do so we don't 
really want to talk to you. 

81
00:05:16,470 --> 00:05:19,154
Database has nothing to offer you until 
you are able to fit in some sort of a 

82
00:05:19,154 --> 00:05:21,728
schema. 
Okay and the observation was that well, 

83
00:05:21,728 --> 00:05:24,977
look a lot of data doesn't had come pre- 
equipped with the schema. 

84
00:05:24,977 --> 00:05:27,749
We don't have a schemer just sort of 
lying around, we had to do something with 

85
00:05:27,749 --> 00:05:30,008
it. 
You know and it's huge, you know, It's 

86
00:05:30,008 --> 00:05:33,360
many hundreds of terabytes or something. 
So, what do we do? 

87
00:05:33,360 --> 00:05:36,350
Well, you know, one answer is, you can 
use these map reduce bases, bases z-, you 

88
00:05:36,350 --> 00:05:39,420
know, Hadoop for this. 
But having to say that, before you're 

89
00:05:39,420 --> 00:05:42,020
allowed to touch your data, you must load 
into a database, that was kind of a non 

90
00:05:42,020 --> 00:05:45,080
starter for a lot of applications. 
Okay. 

91
00:05:45,080 --> 00:05:48,344
And then finally, you know, this idea of 
user defined functions is something that 

92
00:05:48,344 --> 00:05:52,990
all database, or most databases support. 
and it's the idea that you might want to 

93
00:05:52,990 --> 00:05:56,840
do things outside of what you can do in a 
normal SQL query. 

94
00:05:56,840 --> 00:05:59,579
You might want to write your own code and 
push it into the database. 

95
00:05:59,579 --> 00:06:02,099
But the experience of having to sort of 
write and maintain and manage and use 

96
00:06:02,099 --> 00:06:05,064
these things is not great. 
And that's why a lot of people put their 

97
00:06:05,064 --> 00:06:07,624
logic inside the application as opposed 
to pushing it down into the database 

98
00:06:07,624 --> 00:06:13,180
where arguably it could do more good. 
for reasons we'll, we'll talk about. 

99
00:06:13,180 --> 00:06:16,084
Okay, and so I think, you know, MapReduce 
would have argued, you know, look, you 

100
00:06:16,084 --> 00:06:18,812
can actually have, you can give the Java 
programmers what they want, a Java 

101
00:06:18,812 --> 00:06:22,474
programming environment. 
And let them write scalable systems 

102
00:06:22,474 --> 00:06:25,246
without forcing them to kind of use this 
crazy user, user-defined function 

103
00:06:25,246 --> 00:06:27,650
interface that databases offer. 
Okay? 

104
00:06:28,670 --> 00:06:30,050
So fine. 
So what's my point over all this? 

105
00:06:30,050 --> 00:06:33,108
Well. 
You know, if we focus too much on tools, 

106
00:06:33,108 --> 00:06:37,754
what you would get is a snapshot in time 
of what tools are important. 

107
00:06:37,754 --> 00:06:41,707
As opposed to seeing that some of these 
features around databases are, you know, 

108
00:06:41,707 --> 00:06:45,297
they sort of ebb and flow in their 
popularity. 

109
00:06:45,297 --> 00:06:49,368
But their all, their sort of a permanent 
there a permanent value when you're 

110
00:06:49,368 --> 00:06:54,476
reasoning about large-scale systems. 
Similarly you might lose track of what's 

111
00:06:54,476 --> 00:06:58,362
actually novel and what's actually new in 
the midst of the conversation of about 

112
00:06:58,362 --> 00:07:03,289
original databases versus SQL systems. 
Okay so I want to focus on these 

113
00:07:03,289 --> 00:07:07,182
abstractions throughout the course and 
then we can. 

114
00:07:07,182 --> 00:07:11,915
Great, now. 
Okay so fine, we're going to focus on the 

115
00:07:11,915 --> 00:07:14,205
attractions, what are the abstractions of 
data science? 

116
00:07:14,205 --> 00:07:16,607
Well, it's not clear that people really 
know yet. 

117
00:07:16,607 --> 00:07:18,992
And I'll give my case for this is the 
next slide. 

118
00:07:18,992 --> 00:07:22,442
but you know, the reason I don't think we 
really know yet is if you see these words 

119
00:07:22,442 --> 00:07:26,727
being used like, Data Jujitsu and Data 
Wrangling and Data Munging. 

120
00:07:26,727 --> 00:07:30,318
And you know, this is the real skill of 
data scientists, so they had to be able 

121
00:07:30,318 --> 00:07:33,792
to wrangle data, well, what does that 
mean? 

122
00:07:33,792 --> 00:07:36,012
Okay. 
So, my translation of this is, we don't 

123
00:07:36,012 --> 00:07:41,366
really know what we're talking about yet. 
With that said, there's probably a few 

124
00:07:41,366 --> 00:07:47,090
candidates we can consider here. 
So, maybe everything's a matrix and 

125
00:07:47,090 --> 00:07:51,110
everything we want to do with data can be 
expressed in linear algebra. 

126
00:07:51,110 --> 00:07:54,182
If you're a database person, maybe 
everything is a relation, and everything 

127
00:07:54,182 --> 00:07:57,630
you want to is expressed in relational 
algebra. 

128
00:07:57,630 --> 00:08:00,075
If you're more of an object oriented 
programmer. 

129
00:08:00,075 --> 00:08:03,025
Everything is an object and we 
communicate between objects by sending 

130
00:08:03,025 --> 00:08:06,450
messages back and forth through by 
calling methods. 

131
00:08:06,450 --> 00:08:09,440
If you're more of a sysadmin type then, 
you know, everything's a file and we, we 

132
00:08:09,440 --> 00:08:13,822
write bash scripts to process it. 
And if you're an R programmer, then maybe 

133
00:08:13,822 --> 00:08:17,780
everything's a data frame that we call 
functions in this library. 

134
00:08:17,780 --> 00:08:20,952
And Matlab, similarly with Matlab, 
everything's an array or a matrix or a 

135
00:08:20,952 --> 00:08:25,640
vector I guess and their parlance and 
everything's a function on that, okay. 

136
00:08:25,640 --> 00:08:29,736
So, I think of all these possibilities, 
there are two that stand out as likely 

137
00:08:29,736 --> 00:08:34,436
candidates, as fundamental abstractions 
for data science. 

138
00:08:34,436 --> 00:08:38,709
And those are the first two here. 
And the reason is, is that we see these 

139
00:08:38,709 --> 00:08:44,690
abstractions appear over and over again, 
independent of particular tools. 

140
00:08:44,690 --> 00:08:48,218
Now, relations and relational algebra are 
closely associated with databases, but as 

141
00:08:48,218 --> 00:08:52,170
I argued a few slides ago and as you'll 
see throughout the course. 

142
00:08:52,170 --> 00:08:55,070
You'll see this come up time and time 
again, and we even see it in say you 

143
00:08:55,070 --> 00:08:59,850
know, object or in languages, or you see 
it in, R and so on, okay. 

144
00:08:59,850 --> 00:09:07,115
So these are the two that remain the 
focus on, in this course. 

145
00:09:07,115 --> 00:09:13,834
So now I'm want to motivate desktop scale 
versus cloud scale. 

146
00:09:13,834 --> 00:09:17,047
and, you know the argument here for 
desktop scale is that, well, you know 

147
00:09:17,047 --> 00:09:21,063
data science is really about the 
functions and the statistics. 

148
00:09:21,063 --> 00:09:24,494
And the manipulation of techniques. 
So therefore we can sort of push large 

149
00:09:24,494 --> 00:09:27,978
scale data into a separate course, or a 
separate category, and really just focus 

150
00:09:27,978 --> 00:09:32,910
on the, the math and the functions. 
And I think this is a bit of a mistake. 

151
00:09:32,910 --> 00:09:36,485
For a data science course and the reason 
is that you know, this is a fundamental 

152
00:09:36,485 --> 00:09:39,977
limitation of a whole category of 
technologies. 

153
00:09:39,977 --> 00:09:43,399
And R itself is included in that, 
although there's a lot of great work on 

154
00:09:43,399 --> 00:09:47,768
how to sort of scale R up. 
But as it,you know in its basic usage 

155
00:09:47,768 --> 00:09:51,548
what you do with or read a file load the 
whole thing into main memory on one 

156
00:09:51,548 --> 00:09:58,250
machine and then call functions on that. 
And your data doesn't fit in the main 

157
00:09:58,250 --> 00:10:01,892
memory on one machine, you kind of out of 
lock. 

158
00:10:01,892 --> 00:10:05,984
Now, you can be clever and start use 
indices a kind of limit the data you need 

159
00:10:05,984 --> 00:10:09,277
to access. 
And you can start to try to be parallel 

160
00:10:09,277 --> 00:10:11,645
to take advantage of the fact that 
there's now, you know four, and six, and 

161
00:10:11,645 --> 00:10:15,415
eight, and 12 cores in your machine. 
In your computers you'll buy nowadays. 

162
00:10:15,415 --> 00:10:19,087
But trying to be clever and doing that 
yourself overlooks the fact that a lot of 

163
00:10:19,087 --> 00:10:21,963
this. 
A lot of these techniques are pretty 

164
00:10:21,963 --> 00:10:25,489
well-understood and already implemented 
in other systems. 

165
00:10:25,489 --> 00:10:28,675
Okay, so being able to be cognizant of 
what other systems can do and take 

166
00:10:28,675 --> 00:10:32,826
advantages of those flexibly. 
And you know write your application in 

167
00:10:32,826 --> 00:10:36,498
terms of these other systems that already 
do scale out use a critical skill in data 

168
00:10:36,498 --> 00:10:39,852
science. 
And so the point we made in this slide 

169
00:10:39,852 --> 00:10:44,004
that is somewhat out of date although you 
can get the idea is that. 

170
00:10:44,004 --> 00:10:47,298
Simple, I, simple tools that are 
available on every machine such as GREP, 

171
00:10:47,298 --> 00:10:51,356
which if you haven't heard of GREP and 
you're a Windows user. 

172
00:10:51,356 --> 00:10:54,542
And don't use GREP too often then this is 
essentially search a file for a 

173
00:10:54,542 --> 00:10:58,415
particular pattern. 
but it searches it linearly, right and 

174
00:10:58,415 --> 00:11:02,680
looks at every single line of the file 
and checks for the pattern. 

175
00:11:02,680 --> 00:11:06,136
And so you can do a linear scan of a 
megabyte in maybe a second, and a 

176
00:11:06,136 --> 00:11:09,590
gigabyte in a minute. 
And so on. 

177
00:11:09,590 --> 00:11:13,250
And so at a very large scale data sets 
you can't do this linear scan anymore. 

178
00:11:13,250 --> 00:11:15,352
You have to search in a more, in a 
smarter way. 

179
00:11:15,352 --> 00:11:17,348
Okay? 
And you sort of have some cost over here 

180
00:11:17,348 --> 00:11:19,784
that are probably hor, horribly out of 
date now. 

181
00:11:19,784 --> 00:11:24,483
Alright? 
fine so, the point is large scale data is 

182
00:11:24,483 --> 00:11:29,253
not just bigger it's different. 
It requires a different way of thinking 

183
00:11:29,253 --> 00:11:31,514
about techniques. 
And it requires a different stack in 

184
00:11:31,514 --> 00:11:34,056
technologies and to ignore, it's a 
mistake to ignore that in a data science 

185
00:11:34,056 --> 00:11:37,009
course. 
Okay, and then this final dimension of 

186
00:11:37,009 --> 00:11:40,786
sort of hackers versus analysts. 
And again what I mean here is you know 

187
00:11:40,786 --> 00:11:44,324
doing am I going to require sort of deep 
programming proficiency in order to 

188
00:11:44,324 --> 00:11:49,448
participate with this data science class. 
And in general data science kind of 

189
00:11:49,448 --> 00:11:51,885
activities. 
And the answer to that is I don't think 

190
00:11:51,885 --> 00:11:55,408
so, I think, I think we need. 
At least two types of people and really 

191
00:11:55,408 --> 00:11:58,680
sort of a broad spectrum of people. 
And I, and this isn't really my idea, 

192
00:11:58,680 --> 00:12:02,682
this, this often quoted report from the 
Mckinsey Global Institute you'll see this 

193
00:12:02,682 --> 00:12:06,852
quote time and again. 
but they, but the people that use this 

194
00:12:06,852 --> 00:12:09,780
quote tend to focus on this first part 
that talks about 140,000 to 190,000 

195
00:12:09,780 --> 00:12:15,036
people with deep analytical skills. 
But the second part of the quote is well 

196
00:12:15,036 --> 00:12:18,864
you also need 1.5 million managers and 
analysts who know how to use the analysis 

197
00:12:18,864 --> 00:12:22,380
to make effective decisions. 
Okay? 

198
00:12:22,380 --> 00:12:26,528
so this means that it won't be just the 
programmers who are working in this 

199
00:12:26,528 --> 00:12:30,732
space. 
And I wanted to think about how to design 

200
00:12:30,732 --> 00:12:35,666
a course that could Appeal and inform 
both categories of people. 

201
00:12:35,666 --> 00:12:39,040
Alright, and this is my last slide of the 
segment. 

202
00:12:39,040 --> 00:12:42,274
The other reason why I think hackers vs 
analysts is that the line between them is 

203
00:12:42,274 --> 00:12:45,823
kind of blurry nowadays. 
And technology can actually help you, 

204
00:12:45,823 --> 00:12:48,100
right? 
It doesn't require a PhD in computer 

205
00:12:48,100 --> 00:12:51,292
science or a bachelor's degree in 
computer science in some cases, to 

206
00:12:51,292 --> 00:12:55,648
manipulate large data sets. 
And for in order to back up this claim 

207
00:12:55,648 --> 00:12:58,796
you need an example from some of my work 
where. 

208
00:12:58,796 --> 00:13:02,044
we have done some work to try to make 
databases easier to use for say 

209
00:13:02,044 --> 00:13:05,383
biologists. 
And this really nasty looking SQL query 

210
00:13:05,383 --> 00:13:08,683
that if you squint closely you can see 
that it's actually doing Interval 

211
00:13:08,683 --> 00:13:13,640
arithmetic over genetic sequences. 
Right this, this is a pretty tough query 

212
00:13:13,640 --> 00:13:17,735
to understand for even experts. 
This was written by somebody who doesn't 

213
00:13:17,735 --> 00:13:21,550
do any programming whatsoever. 
She doesn't write a line of Python she 

214
00:13:21,550 --> 00:13:25,238
doesn't write a line of pearl. 
She doesn't write a line R. 

215
00:13:25,238 --> 00:13:29,097
And she's able to these SQL queries to 
process very large data sets. 

216
00:13:29,097 --> 00:13:32,625
Okay so the fact that you know if you 
understand what's going on and if you can 

217
00:13:32,625 --> 00:13:36,173
think in terms of some of these 
abstractions. 

218
00:13:36,173 --> 00:13:39,638
and you understand your problem well 
enough, you can participate in the 

219
00:13:39,638 --> 00:13:44,678
activity of manipulating our data sets. 
And doing data science even without a a 

220
00:13:44,678 --> 00:13:49,146
deep background in software engineering. 
Okay and that's why I want to push this 

221
00:13:49,146 --> 00:13:53,055
needle, somewhere over this way. 
I'll probably put this in the middle, I 

222
00:13:53,055 --> 00:13:54,831
suppose. 
I'm not so much trying to focus on only 

223
00:13:54,831 --> 00:13:57,094
the analysts, I just trying to make sure 
that they're included. 

224
00:13:57,094 --> 00:14:00,511
Okay. 
Next time, we'll pick up at the last 

225
00:14:00,511 --> 00:14:02,230
dimension. 

