1
00:00:00,103 --> 00:00:05,059
[MUSIC]. 

2
00:00:05,059 --> 00:00:07,527
Welcome back. 
Last time we talked about three out of 

3
00:00:07,527 --> 00:00:10,565
four of these dimensions in describing 
how we designed this course in data 

4
00:00:10,565 --> 00:00:13,959
science. 
And so, in this segment, I want to talk 

5
00:00:13,959 --> 00:00:18,249
about this last dimension of what I, what 
I call structs versus stats and so this 

6
00:00:18,249 --> 00:00:21,653
is. 
The relative importance of data 

7
00:00:21,653 --> 00:00:25,803
manipulation versus deeper mathematics. 
And you can see that I've sort of put 

8
00:00:25,803 --> 00:00:30,158
the, the dial here a little bit to the 
left, and I'll try to motivate that in, 

9
00:00:30,158 --> 00:00:36,336
in the next few minutes, alright? 
So, we already saw one example of this in 

10
00:00:36,336 --> 00:00:40,305
the first segment, where I gave some 
examples of data science from you know, 

11
00:00:40,305 --> 00:00:44,589
recent history and one of these was Nate 
Silver's prediction of the electoral 

12
00:00:44,589 --> 00:00:50,581
college votes for the 2012 US 
Presidential Election. 

13
00:00:50,581 --> 00:00:54,424
And if you recall you know, this was, 
this prediction was accomplished by 

14
00:00:54,424 --> 00:00:59,645
essentially taking the average of the 
state polls for each state. 

15
00:00:59,645 --> 00:01:03,485
Okay, so it really didn't require a 
sophisticated statistical model and yet 

16
00:01:03,485 --> 00:01:07,332
it had massive impact. 
Okay, so you know a quote that I think 

17
00:01:07,332 --> 00:01:10,812
sums this up a little bit, comes from 
Aaron Kimball at a company called 

18
00:01:10,812 --> 00:01:14,532
WibiData, he, he says, you know 80% of 
the analytics is really just sums and 

19
00:01:14,532 --> 00:01:20,455
averages and so if you can get these. 
What he means by this is if you can get 

20
00:01:20,455 --> 00:01:23,399
these sums and averages right, if you can 
do it at any scale on any data that you 

21
00:01:23,399 --> 00:01:26,343
might see, then you can always sort of 
build up more and more more advanced 

22
00:01:26,343 --> 00:01:30,030
techniques. 
Everything sort of boiled down to just 

23
00:01:30,030 --> 00:01:32,788
sums and averages. 
Okay, so I think this is a motivation for 

24
00:01:32,788 --> 00:01:36,652
why focusing on data manipulation, which 
can, which typically is associated with 

25
00:01:36,652 --> 00:01:40,684
being able to express sums and averages, 
for example you know what you can do in a 

26
00:01:40,684 --> 00:01:46,722
database query which we talk about in the 
next couple of lectures. 

27
00:01:46,722 --> 00:01:50,699
That gets you a pretty long way, alright? 
It is the 80% of the problem, so another 

28
00:01:50,699 --> 00:01:54,263
way of looking at it is there is three 
main tasks involved with you know a data 

29
00:01:54,263 --> 00:01:59,270
science project. 
There's preparing to run the model. 

30
00:01:59,270 --> 00:02:01,832
Running the actual surgical model and 
then interpreting the results and 

31
00:02:01,832 --> 00:02:04,144
communicating it. 
I got the animation out of border here 

32
00:02:04,144 --> 00:02:07,090
you can ignore that red. 
But the point here is you know, again 

33
00:02:07,090 --> 00:02:10,172
Aaron Kimball from a conversation[LAUGH] 
with him where he got this, was the you 

34
00:02:10,172 --> 00:02:13,070
know 80% of the work is really in this 
first step where you are gathering data 

35
00:02:13,070 --> 00:02:16,244
and cleaning it and integrating it and 
restructuring it transforming it, loading 

36
00:02:16,244 --> 00:02:21,482
it and so on. 
So, verifying all these verbs you see 

37
00:02:21,482 --> 00:02:26,765
here, this is the hard part, right? 
And so, actually running the model or 

38
00:02:26,765 --> 00:02:31,451
even choosing the model and then running 
it doesn't tend to keep people up at 

39
00:02:31,451 --> 00:02:36,236
night in practice, okay? 
So, and then the joke is perhaps that the 

40
00:02:36,236 --> 00:02:38,924
other 80% of the work, you know, implying 
that there's sort of 160% of a normal 

41
00:02:38,924 --> 00:02:42,510
task is in data science is in this 
interpreting the results. 

42
00:02:42,510 --> 00:02:46,074
So, this is the visualization and the 
communication and the explanation of the 

43
00:02:46,074 --> 00:02:49,105
results. 
Okay, so this is another reason why I 

44
00:02:49,105 --> 00:02:52,825
want to focus in this course on data 
manipulation tasks that are associated 

45
00:02:52,825 --> 00:02:58,052
with this first number one task. 
Okay, you know another way of looking at 

46
00:02:58,052 --> 00:03:01,150
this is a quote that is now really old, 
right? 

47
00:03:01,150 --> 00:03:05,615
So, this is 12 years old or so at, at the 
time of this recording from Doug Laney. 

48
00:03:05,615 --> 00:03:09,071
And this is the document that first 
coined this notion of big data of being 

49
00:03:09,071 --> 00:03:12,689
the 3D's of volume, velocity, and variety 
and we'll talk about that in a couple of 

50
00:03:12,689 --> 00:03:16,776
segments. 
But he has this quote, you know, no 

51
00:03:16,776 --> 00:03:19,784
greater barrier to effective data 
management will exist than the variety of 

52
00:03:19,784 --> 00:03:23,027
incompatible data formats, non-aligned 
data structures, and inconsistent data 

53
00:03:23,027 --> 00:03:26,329
semantics. 
So this is what the database community, 

54
00:03:26,329 --> 00:03:29,038
you know, my community calls the data 
integration problem, this is the hard 

55
00:03:29,038 --> 00:03:32,197
part. 
And so, he was saying this back in 2001 

56
00:03:32,197 --> 00:03:35,890
and I would argue that it's still true 
today. 

57
00:03:35,890 --> 00:03:38,962
This is the greatest barrier, and in the 
context of this, he was talking about 

58
00:03:38,962 --> 00:03:42,178
this notion that variety being harder 
than volume or velocity, and I'll explain 

59
00:03:42,178 --> 00:03:46,158
more about what those Vs mean in a couple 
of segments. 

60
00:03:46,158 --> 00:03:50,740
Alright, so another been yet here, is 
something that we like to ask the 

61
00:03:50,740 --> 00:03:54,978
scientists we work with. 
So these are, you know, astronomers and 

62
00:03:54,978 --> 00:03:57,825
oceanographers and biologists. 
We ask them sort of informally, how much 

63
00:03:57,825 --> 00:04:01,605
time they spend quote, handling data, as 
opposed to, quote, doing science? 

64
00:04:01,605 --> 00:04:04,805
Now, you know, we let them interpret 
these quotes however they want, but what 

65
00:04:04,805 --> 00:04:07,755
we mean by doing science, you know, 
choosing the statistical method or 

66
00:04:07,755 --> 00:04:12,065
designing a statistical model. 
They absolutely consider it part of their 

67
00:04:12,065 --> 00:04:14,946
science and so, we mean by handling data 
is altogether crap, you know, the format 

68
00:04:14,946 --> 00:04:18,677
conversions and so on. 
And so, what do you think the most common 

69
00:04:18,677 --> 00:04:21,541
answer is here? 
Or you can guess to yourself for a second 

70
00:04:21,541 --> 00:04:25,246
but they, they don't even blink and they 
say things like 90% and so, this number 

71
00:04:25,246 --> 00:04:30,573
should you know give pause, right? 
This is taxpayer money that goes to 

72
00:04:30,573 --> 00:04:35,330
federal funding agencies, to come back to 
pay some post-doctoral fellow to spend 

73
00:04:35,330 --> 00:04:41,705
90% of her time, doing something that she 
doesn't even consider science. 

74
00:04:41,705 --> 00:04:44,991
And so, this is the why it's really 
important as a data scientist to focus on 

75
00:04:44,991 --> 00:04:47,160
this problem. 
Now, you might say, well that's just 

76
00:04:47,160 --> 00:04:49,617
science, what about business? 
But as I will try to make the point 

77
00:04:49,617 --> 00:04:53,721
throughout the course there's a an 
increasing alignment between what's going 

78
00:04:53,721 --> 00:04:57,757
on in business and what's going on in 
science. 

79
00:04:57,757 --> 00:05:01,617
Okay, and we, we will talk about that 
more in a couple of segments. 

80
00:05:01,617 --> 00:05:04,704
Alright, so if 90%of the problem is 
handling data, you know boy we have spent 

81
00:05:04,704 --> 00:05:08,570
a lot of tension on that. 
Alright, so you know another argument 

82
00:05:08,570 --> 00:05:12,332
that sort of follows on the first side I 
gave, is that struck's you know, the, the 

83
00:05:12,332 --> 00:05:16,379
data manipulation platforms and databases 
in particular, actually go a pretty long 

84
00:05:16,379 --> 00:05:19,628
way to be in able to express more 
advanced things and this isn't just a 

85
00:05:19,628 --> 00:05:26,510
matter way you can express anything you 
hear sums and averages. 

86
00:05:26,510 --> 00:05:32,208
It's also even fairly advanced techniques 
you can, there's increasing amount of 

87
00:05:32,208 --> 00:05:38,155
interest in figuring out how to get this 
stuff into the database. 

88
00:05:38,155 --> 00:05:42,565
Okay, so this is the a side of thinking 
from Christian Grant, where he argues 

89
00:05:42,565 --> 00:05:47,045
that look you know, if you consider 
databases versus statistical packages, 

90
00:05:47,045 --> 00:05:52,987
such as SAS or Matlab/R or SPSS. 
you know, this is what they are doing 

91
00:05:52,987 --> 00:05:58,430
now, their taking, they are downloading 
data to use in their[INAUDIBLE] package. 

92
00:05:58,430 --> 00:06:01,390
frequently under the assumption that well 
of course I have to, right of course the 

93
00:06:01,390 --> 00:06:05,555
only thing we can possible express this. 
Well most of these sand pack, here, just 

94
00:06:05,555 --> 00:06:09,020
the first thing you will do is, read the 
data off the disc and load into memory 

95
00:06:09,020 --> 00:06:15,118
and then start calling functions on it. 
Well if it, it increasingly data sets 

96
00:06:15,118 --> 00:06:20,340
simply don't fit in memory on a single 
machine certainly not on your laptop. 

97
00:06:20,340 --> 00:06:23,658
I'll take a couple of choices here. 
Either you shift into some kind of, fancy 

98
00:06:23,658 --> 00:06:27,812
cluster diversion of the tool for which 
they exist for things like SPSS and 

99
00:06:27,812 --> 00:06:34,560
Matlab although they are quite expensive. 
Or you sample the data, so that you all 

100
00:06:34,560 --> 00:06:38,404
may ever, only can work with the subset 
that actually does fit in memory and 

101
00:06:38,404 --> 00:06:45,062
you'll see this to be very, very common. 
It's that it's just powerfully course to, 

102
00:06:45,062 --> 00:06:48,844
to, to take the sample of the, the data 
in order to be able to work with it 

103
00:06:48,844 --> 00:06:53,731
efficiently, right? 
But the point here is that this isn't 

104
00:06:53,731 --> 00:06:57,310
really required if you use different 
packaging. 

105
00:06:57,310 --> 00:07:00,170
In this case, if you, if the argument 
here is that if you use databases. 

106
00:07:00,170 --> 00:07:04,202
If you can figure out how to, how to 
perform your task in the database, you'll 

107
00:07:04,202 --> 00:07:08,886
get the scalability for free. 
Moreover, you know, these tool kits don't 

108
00:07:08,886 --> 00:07:12,840
have, don't necessarily have any kind of 
notion of parallelism, right? 

109
00:07:12,840 --> 00:07:16,220
So even if it does fit in memory, every 
machine you buy nowadays has you know, at 

110
00:07:16,220 --> 00:07:21,116
least four cores in it and probably more 
like eight and soon to be 12 and 16. 

111
00:07:21,116 --> 00:07:24,584
So to take advantage of all those cores 
on your problem is, is something you're 

112
00:07:24,584 --> 00:07:28,540
going to be looking for in a package. 
And this something the databases can do 

113
00:07:28,540 --> 00:07:32,218
automatically, most databases not all. 
In fact, the, the ones you may be 

114
00:07:32,218 --> 00:07:36,104
familiar with my SQL and sparse matrix do 
not, but other databases will and we'll 

115
00:07:36,104 --> 00:07:40,526
talk more about this. 
Okay, so you get parallelism for free if 

116
00:07:40,526 --> 00:07:43,958
you can use a database and you get 
scalability beyond the size of main 

117
00:07:43,958 --> 00:07:47,234
memory for free if you can use a database 
and that perhaps is a big if and we'll 

118
00:07:47,234 --> 00:07:51,825
talk about it. 
Okay, the leading example here, and 

119
00:07:51,825 --> 00:07:55,049
you'll actually do this as part of a 
homework assignment, but you know, can 

120
00:07:55,049 --> 00:07:59,971
you express matrix multiplication in SQL? 
and if you can, then I'd argue, well, 

121
00:07:59,971 --> 00:08:03,058
hey, any formula that you can express 
using matrix multiplication, you can 

122
00:08:03,058 --> 00:08:07,069
perhaps express in SQL, by doing this 
over and over again. 

123
00:08:07,069 --> 00:08:10,363
Okay, and the answer is yes, and in fact 
the simplest version of this is, is 

124
00:08:10,363 --> 00:08:14,295
pretty straightforward. 
so if you haven't ever seen SQL before, 

125
00:08:14,295 --> 00:08:16,610
don't worry, we will talk a little more 
about this. 

126
00:08:16,610 --> 00:08:20,258
But if you have, you know, bear with me, 
imagine you have two matrices, A and B, 

127
00:08:20,258 --> 00:08:28,904
oops, I'm using the wrong device here. 
Two matrices using A and B and what you 

128
00:08:28,904 --> 00:08:38,686
want to do is find all of the you know 
and the, the the representation of this 

129
00:08:38,686 --> 00:08:52,367
matrix here is a row excuse me, row ID, 
column ID, and value. 

130
00:08:52,367 --> 00:08:58,856
Alright, so that's a relation. 
Now, this is a very inefficient relation 

131
00:08:58,856 --> 00:09:02,898
if your matrix is dense. 
And I'll let you think about what, well 

132
00:09:02,898 --> 00:09:05,490
I'll tell you why and you think about it 
a little bit more as well. 

133
00:09:05,490 --> 00:09:09,376
Is that, you know, an implicit 
representation of this only has, let's 

134
00:09:09,376 --> 00:09:13,932
say you have in rows, let's say you have 
five rows and six columns, then you only 

135
00:09:13,932 --> 00:09:21,094
need the 30 values, 5 times 6. 
But here you're doing, you have to do 30 

136
00:09:21,094 --> 00:09:26,290
row ID's plus 30 column ID's plus 30 
values. 

137
00:09:26,290 --> 00:09:29,881
So, you sort of triple the size of your 
data relative to, you know, efficient 

138
00:09:29,881 --> 00:09:33,190
name memory representations. 
So why would you do that? 

139
00:09:33,190 --> 00:09:36,134
Well, it turns out that a lot of matrix's 
in practice are sparse, and I put that 

140
00:09:36,134 --> 00:09:39,340
word right up here on top. 
In a sparse matrix, not all the cells 

141
00:09:39,340 --> 00:09:42,630
actually have a value and so you don't 
actually need to store them. 

142
00:09:42,630 --> 00:09:45,712
And so this representation in terms of 
you know, explicitly having a row ID, a 

143
00:09:45,712 --> 00:09:48,872
column ID, and value turns out be pretty 
efficient. 

144
00:09:48,872 --> 00:09:53,356
Okay and in fact, sparse matrix solvers, 
this is exactly the kind of 

145
00:09:53,356 --> 00:09:58,977
representation they use internally. 
Alright, so if you have a sparse matrix 

146
00:09:58,977 --> 00:10:02,689
and if you encode it in, if you represent 
in a database then expressing matrix 

147
00:10:02,689 --> 00:10:07,800
multiply is not too bad. 
You, what you want to do is find all the 

148
00:10:07,800 --> 00:10:12,225
columns, you know, for each column number 
in, in the, matrix A, find the 

149
00:10:12,225 --> 00:10:18,517
corresponding row number, in column B. 
and then add up all the all the 

150
00:10:18,517 --> 00:10:22,910
contributions to the new value, and also 
a diagram of this after. 

151
00:10:22,910 --> 00:10:25,367
In fact, you know, let me skip going into 
too much detail about this right now 

152
00:10:25,367 --> 00:10:27,824
because I'm going to talk about this in 
detail in preparation for the homework 

153
00:10:27,824 --> 00:10:32,224
while you'll do this. 
So right now, I guess take away, what I 

154
00:10:32,224 --> 00:10:35,552
want you to take away is that 
representing matrices inside of a 

155
00:10:35,552 --> 00:10:40,584
database sounds very unusual. 
It's actually not the world's worst idea 

156
00:10:40,584 --> 00:10:44,103
and in part of the readings from this mad 
skills paper, you'll try to, you'll see 

157
00:10:44,103 --> 00:10:46,710
why. 
So, right now, I just want you to take 

158
00:10:46,710 --> 00:10:49,420
away that it can be done and it's not 
necessarily a terrible idea. 

