1
00:00:00,617 --> 00:00:05,821
[MUSIC]. 

2
00:00:05,821 --> 00:00:08,286
Welcome back. 
So I want to talk a little bit about how 

3
00:00:08,286 --> 00:00:11,635
the term "data science" relates to other 
fields of science. 

4
00:00:11,635 --> 00:00:14,690
And in particular, I want to introduce 
this term eScience, which through first 

5
00:00:14,690 --> 00:00:18,180
approximation you can think of as 
equivalent to Data Science. 

6
00:00:18,180 --> 00:00:21,308
So while the eScience is associated with 
Astronomy and Oceanography and Biology 

7
00:00:21,308 --> 00:00:24,030
data science has been adopted more in 
business. 

8
00:00:24,030 --> 00:00:25,960
But they involve a lot of the same 
concepts. 

9
00:00:25,960 --> 00:00:28,050
So let me tell you what's going on in 
Science. 

10
00:00:28,050 --> 00:00:31,314
So for thousands of years, you know, 
scientific inquiry has been empirical, 

11
00:00:31,314 --> 00:00:33,461
right. 
You observe the natural world, or maybe 

12
00:00:33,461 --> 00:00:35,755
in some cases replicate the natural world 
in a controlled environment in a 

13
00:00:35,755 --> 00:00:39,173
laboratory. 
And make observations about that. 

14
00:00:39,173 --> 00:00:43,528
In the last few hundred years science has 
accepted theoretical models as a valid 

15
00:00:43,528 --> 00:00:47,346
method of inquiry. 
One that is reinforcing empirical methods 

16
00:00:47,346 --> 00:00:50,428
so you know new theories suggests new 
experiments and the theories help explain 

17
00:00:50,428 --> 00:00:53,720
the observed data that you get from the 
experiments. 

18
00:00:54,870 --> 00:00:58,170
In the last 50 years or so high speed 
computation has emitted an entirely new 

19
00:00:58,170 --> 00:01:02,114
method of scientific inquiry. 
Soon you can simulate in the computer 

20
00:01:02,114 --> 00:01:05,678
phenomena on that otherwise couldn't be 
re, you couldn't, you can't observe 

21
00:01:05,678 --> 00:01:08,935
directly. 
And you can't reproduce in the lab and 

22
00:01:08,935 --> 00:01:12,769
even the theoretical models become too 
complex to solve analytically using, you 

23
00:01:12,769 --> 00:01:17,612
know essentially paper and pencil. 
Right, but you can actually start from 

24
00:01:17,612 --> 00:01:21,280
initial conditions and run the simulation 
to get a result. 

25
00:01:21,280 --> 00:01:24,513
So this is maybe what goes on in the 
interior stars or the shift of tectonic 

26
00:01:24,513 --> 00:01:27,852
plates or the evolution of the universe 
or the effects of the ecology on some 

27
00:01:27,852 --> 00:01:32,940
species dying out, and so on. 
So that's fine. 

28
00:01:32,940 --> 00:01:35,864
That's three methods of inquiry. 
But in the last ten years or so, there's 

29
00:01:35,864 --> 00:01:39,341
been arguably a fourth method of 
scientific Inquiry, which is to acquire 

30
00:01:39,341 --> 00:01:43,684
massive data sets from instruments or 
from simulations. 

31
00:01:43,684 --> 00:01:48,320
And then explore these data sets using 
new algorithms and infrastructure. 

32
00:01:48,320 --> 00:01:51,115
And so eScience is really about massive 
and complex data, data large enough to 

33
00:01:51,115 --> 00:01:54,170
require you know, automated or 
semi-automated analysis. 

34
00:01:54,170 --> 00:01:57,380
You can't look at it, you can't inspect 
it directly, okay. 

35
00:01:57,380 --> 00:02:00,565
And so the relevant tools here are the 
same as those for data science, you know, 

36
00:02:00,565 --> 00:02:03,760
databases, visualization scale out 
computing. 

37
00:02:03,760 --> 00:02:07,720
Maybe the new sequel systems. 
Machine learning techniques. 

38
00:02:07,720 --> 00:02:12,620
Web services and so on. 
Okay. 

39
00:02:12,620 --> 00:02:15,884
So the way, that, this, this, idea of the 
fourth paradigm, there's a book that's in 

40
00:02:15,884 --> 00:02:18,840
the reading list that you can refer to 
here. 

41
00:02:18,840 --> 00:02:22,206
And there's a lot of, just some other 
articles in your reading list you can 

42
00:02:22,206 --> 00:02:25,060
also refer to. 
The story has been told lots of ways. 

43
00:02:25,060 --> 00:02:27,912
The way I like to talk about this story 
is that science has always been about 

44
00:02:27,912 --> 00:02:30,902
asking questions but conventionally it 
was really about querying the world, 

45
00:02:30,902 --> 00:02:34,290
right? 
You would, sort of, have data acquisition 

46
00:02:34,290 --> 00:02:37,910
or activities, experiments or field 
studies. 

47
00:02:37,910 --> 00:02:41,186
They will couple to very specific 
hypothesis where you have the question in 

48
00:02:41,186 --> 00:02:45,210
mind first and we click the data. 
But eScience has really shifted a bit 

49
00:02:45,210 --> 00:02:48,694
where now you're kind of downloading data 
on mass, you're downloading the world 

50
00:02:48,694 --> 00:02:52,990
first putting some sort of representation 
in the computer. 

51
00:02:52,990 --> 00:02:56,420
And then creating that database to test 
your hypothesis and so it's the, the data 

52
00:02:56,420 --> 00:03:00,695
can be acquired independent of any 
specific hypothesis in some case. 

53
00:03:00,695 --> 00:03:03,437
Okay. 
And this is due in part to the cost of 

54
00:03:03,437 --> 00:03:08,430
data acquisition dropping precipitously, 
thanks for advances in technology, right. 

55
00:03:08,430 --> 00:03:11,745
So the telescopes you can build now, that 
we'll talk about in the next couple of 

56
00:03:11,745 --> 00:03:16,320
slides, can acquire at enormous amounts 
of data at very high resolution. 

57
00:03:17,520 --> 00:03:20,103
And in the life sciences, you have sort 
of laboratory automation and sort of 

58
00:03:20,103 --> 00:03:23,238
high-throughput sequencing. 
In oceanography the sensors are getting 

59
00:03:23,238 --> 00:03:25,801
cheaper. 
The models thanks to advances and thanks 

60
00:03:25,801 --> 00:03:28,380
to more's laws. 
And thanks to advances in computing. 

61
00:03:28,380 --> 00:03:30,975
The simulations you can run are getting 
bigger and higher resolution. 

62
00:03:30,975 --> 00:03:33,490
And producing, and therefore producing 
larger, and larger amounts of data. 

63
00:03:33,490 --> 00:03:36,540
And so on, and so you know, the rate at 
which data can be produced has far out 

64
00:03:36,540 --> 00:03:39,790
paced the rate that we can analyze it or 
come up with the questions we need to ask 

65
00:03:39,790 --> 00:03:42,240
about it. 
Okay. 

66
00:03:42,240 --> 00:03:44,690
And this suggested a new approach to 
science. 

67
00:03:45,960 --> 00:03:50,316
So let me give some examples, so we said 
that eScience is driven by data more than 

68
00:03:50,316 --> 00:03:54,705
by the computation. 
Alright so some examples on the size of 

69
00:03:54,705 --> 00:03:58,629
the data is that's coming on. 
The Apache point telescope that was the 

70
00:03:58,629 --> 00:04:01,915
primary instrument for the Sloan Digital 
Sky Survey that we might refer to 

71
00:04:01,915 --> 00:04:06,562
multiple times in this course. 
Produced 80 terabytes of raw new data 

72
00:04:06,562 --> 00:04:10,560
over, a seven year period. 
You know, at the time, this is a pretty 

73
00:04:10,560 --> 00:04:15,060
significant data size, and even by many 
standards is still today. 

74
00:04:15,060 --> 00:04:18,960
The next generation of this the next 
generation project that's in the same 

75
00:04:18,960 --> 00:04:22,500
sort of spirit as Sloan Digital Sky 
Survey is the large synoptic survey 

76
00:04:22,500 --> 00:04:26,902
telescope. 
So this guy can produce 40 terabytes per 

77
00:04:26,902 --> 00:04:31,580
day. 
it will do so for over a ten year period. 

78
00:04:31,580 --> 00:04:35,176
So in total 100 plus petabytes and 
producing a single amount of data that's 

79
00:04:35,176 --> 00:04:39,790
soon the skies will be produced in over a 
total entire lifetime. 

80
00:04:39,790 --> 00:04:41,910
It can be produce that over every two 
days. 

81
00:04:41,910 --> 00:04:44,565
Okay. 
And so this is a pretty staggering amount 

82
00:04:44,565 --> 00:04:48,050
of data, and requires a, a pretty 
different approach. 

83
00:04:48,050 --> 00:04:51,039
One thing I want to mention maybe about 
Sloan Digital Sky Survey, what they 

84
00:04:51,039 --> 00:04:53,979
actually did here, was to take the 
images, cook them, right, extract the 

85
00:04:53,979 --> 00:04:57,115
relevant objects from it, put all those 
objects into a, an off-the-shelf or a 

86
00:04:57,115 --> 00:05:02,172
relational database. 
In fact, it was Microsoft SQL server, and 

87
00:05:02,172 --> 00:05:05,826
critically host this database online and 
serve it out over the web, and this 

88
00:05:05,826 --> 00:05:10,620
require a pretty significant investment 
in infrastructure. 

89
00:05:10,620 --> 00:05:14,585
But as a result f doing this, of making 
all the data public and queryable, it 

90
00:05:14,585 --> 00:05:19,130
became the most productive astronomy 
facility in history. 

91
00:05:19,130 --> 00:05:21,860
Right, so the number of papers that have 
been produced on this data is on the 

92
00:05:21,860 --> 00:05:25,490
order thousands. 
In the original, you know, PIs of the 

93
00:05:25,490 --> 00:05:28,890
project, the principal investigators of 
the project had, sort of, maybe only 100 

94
00:05:28,890 --> 00:05:33,034
papers in mind for the data. 
And the other 4900 papers that have been 

95
00:05:33,034 --> 00:05:37,370
written all came from external partners 
writing queries against this database. 

96
00:05:37,370 --> 00:05:41,610
So it's just a wild, wild success. 
Now, the problem is, is that the same 

97
00:05:41,610 --> 00:05:44,945
technology stack in, into some extent 
even the same approach. 

98
00:05:44,945 --> 00:05:49,250
It is difficult to apply in this case of 
large synoptic survey telescope. 

99
00:05:49,250 --> 00:05:52,706
The reason why this guy is producing so 
much more data is not just because it's 

100
00:05:52,706 --> 00:05:57,662
much higher resolution and it can 
perceive a much deeper field in the sky. 

101
00:05:57,662 --> 00:06:01,237
But also because it's returning to the 
same point in the sky frequently, every 

102
00:06:01,237 --> 00:06:05,859
three days, and so this allows you to 
look at things change over time. 

103
00:06:05,859 --> 00:06:09,845
So Asteroids, comets, you might catch, 
super novas and so forth. 

104
00:06:09,845 --> 00:06:14,815
Okay, and by comparing these images in 
the time series, there, there's all sorts 

105
00:06:14,815 --> 00:06:19,080
of new questions you can ask. 
Okay. 

106
00:06:19,080 --> 00:06:22,050
So, both because of the science that 
they're going to do. 

107
00:06:22,050 --> 00:06:25,314
And because of the sheer scale. 
and because of some of the complexity of 

108
00:06:25,314 --> 00:06:31,506
the, details of how the data is acquired. 
The Existing, the previous solution won't 

109
00:06:31,506 --> 00:06:34,246
work. 
And so, this is motivated in a whole new 

110
00:06:34,246 --> 00:06:37,479
area of research to study, data 
management techniques and data analysis 

111
00:06:37,479 --> 00:06:43,430
techniques to support this project. 
So in Life Sciences, these high 

112
00:06:43,430 --> 00:06:46,355
throughput sequencers are capable of 
producing, you know, terabytes per day 

113
00:06:46,355 --> 00:06:50,030
when run continuously. 
And you know, major labs that do this 

114
00:06:50,030 --> 00:06:53,294
work, such as the Joint Genome Institute, 
have 25 to 100 of these machines running 

115
00:06:53,294 --> 00:06:57,117
all the time, alright. 
So this is spitting out an enormous 

116
00:06:57,117 --> 00:07:01,136
amount of data for, well, I was going to 
say a variety of samples. 

117
00:07:01,136 --> 00:07:04,422
So it, so it may be individual organisms 
or it could even be samples from the 

118
00:07:04,422 --> 00:07:08,026
environment where there's no particular 
one organism in there but there's an 

119
00:07:08,026 --> 00:07:12,246
entire population. 
Alright, so for a variety of uses these 

120
00:07:12,246 --> 00:07:19,312
guys are are able to spit out the data. 
In oceanography, the regional scale nodes 

121
00:07:19,312 --> 00:07:23,110
of the NSF Ocean Observatories Initiative 
is a project led here at UW. 

122
00:07:23,110 --> 00:07:26,640
Ocean Observatories Initiative is a 
multi-institutional partnership. 

123
00:07:26,640 --> 00:07:30,380
the regional scale nodes part is run at 
the University of Washington. 

124
00:07:30,380 --> 00:07:34,412
So these, this project does involve 
laying, you know, a thousand kilometers 

125
00:07:34,412 --> 00:07:39,697
of fiber optic cable on the seafloor. 
Connecting thousands of instruments in 

126
00:07:39,697 --> 00:07:44,233
chemical, physical and biological 
thousands of chemical, biological and 

127
00:07:44,233 --> 00:07:49,107
physical sensors. 
Including live video from the seafloor to 

128
00:07:49,107 --> 00:07:55,122
measure to monitor volcanic activity. 
Okay, so again, the database the, the if 

129
00:07:55,122 --> 00:07:58,140
not a relational database. 
[LAUGH]. 

130
00:07:58,140 --> 00:08:02,471
The data sets and data infrastructure 
required to support this effort is 

131
00:08:02,471 --> 00:08:07,112
significant and has motivated a lot of 
new research. 

132
00:08:07,112 --> 00:08:10,214
Alright. 
in the information space there is a lot 

133
00:08:10,214 --> 00:08:15,048
of science to be done on the web itself. 
And so, just the web, a single computer 

134
00:08:15,048 --> 00:08:19,144
can read 30 to 35 megabytes per second 
from one disk, and so it would take about 

135
00:08:19,144 --> 00:08:30,785
four months just to read the entire web. 
So new clusters of machines. 

136
00:08:30,785 --> 00:08:39,430
So summing up a little bit eScience is 
about the analysis of data. 

137
00:08:39,430 --> 00:08:41,910
So the automated or semi-automated 
extraction of knowledge from massive 

138
00:08:41,910 --> 00:08:44,716
volumes of data. 
And so your main instrument for looking 

139
00:08:44,716 --> 00:08:47,992
for answers is the, or the algorithms and 
the technology as oppose to direct 

140
00:08:47,992 --> 00:08:50,700
inspection. 
There's just too much of it to look at it 

141
00:08:50,700 --> 00:08:53,234
yourself. 
But it's not just a matter of volume as 

142
00:08:53,234 --> 00:08:57,499
we'll talk about in the next segment. 
this is another link back to what's going 

143
00:08:57,499 --> 00:09:00,030
in business. 
Right, there's these, there's this 

144
00:09:00,030 --> 00:09:02,175
concept of big data and there's the three 
V's of big data that we'll talk about a 

145
00:09:02,175 --> 00:09:05,250
little bit more next time. 
But let me just mention them here. 

146
00:09:05,250 --> 00:09:08,476
Where volume, sort of the three V's are 
volume, variety, and velocity. 

147
00:09:08,476 --> 00:09:11,356
They will deal about and I will give you 
the, where this stuffs came from in the 

148
00:09:11,356 --> 00:09:14,193
next segment. 
Involving first just into the number of 

149
00:09:14,193 --> 00:09:16,760
rows and the number of bytes at the 
serious scale. 

150
00:09:16,760 --> 00:09:20,792
Variety is perhaps a number of columns 
and dimensions but you know in science 

151
00:09:20,792 --> 00:09:25,076
for example lot of in allied sciences 
particularly you will have experiments 

152
00:09:25,076 --> 00:09:31,610
that involve accessing multiple databases 
as well as multiple sensors. 

153
00:09:31,610 --> 00:09:34,874
you're own data that you've collected and 
that of your colleagues and the 

154
00:09:34,874 --> 00:09:38,087
integration task of putting all these 
data together is a significant model 

155
00:09:38,087 --> 00:09:41,416
like. 
Even when the actual scale on the data is 

156
00:09:41,416 --> 00:09:45,020
not necessarily all that bad, okay? 
So those are the complexity of the data. 

157
00:09:45,020 --> 00:09:48,970
And then the velocity, you know, we saw 
what the large survey telescope that. 

158
00:09:50,800 --> 00:09:53,790
You know although the, the, the scale 
itself is enormous, that fact that 40, 40 

159
00:09:53,790 --> 00:09:56,964
terabytes are being collected every two 
days means that the infrastructure needs 

160
00:09:56,964 --> 00:10:01,494
to keep up with that pace. 
And just transferring that data from the 

161
00:10:01,494 --> 00:10:06,892
telescope facility to the data analysis 
facility is an engineering challenge. 

162
00:10:06,892 --> 00:10:08,761
Okay? 
And you'll also see other V's here, with 

163
00:10:08,761 --> 00:10:11,040
things like veracity. 
You know, can we actually trust this 

164
00:10:11,040 --> 00:10:15,830
data? 
Okay, so a bit more of that next time. 

165
00:10:15,830 --> 00:10:18,866
So to summarize here, science is in the 
midst of a generational shift from a data 

166
00:10:18,866 --> 00:10:23,020
poor enterprise, where you can never, you 
know there's never enough data. 

167
00:10:23,020 --> 00:10:25,225
To a data rich enterprise where there's 
so much of it you don't know what to do 

168
00:10:25,225 --> 00:10:27,714
with it. 
And as a result, you know data analysis 

169
00:10:27,714 --> 00:10:31,184
has replaced data acquisition as the new 
bottleneck to discovery. 

170
00:10:31,184 --> 00:10:32,474
Right? 
So it's not the cost of going out and 

171
00:10:32,474 --> 00:10:34,514
getting the data, it's the cost of 
actually analyzing the data you might 

172
00:10:34,514 --> 00:10:38,474
already have. 
So this is fine, but what does this have 

173
00:10:38,474 --> 00:10:41,372
to do with business, which is probably 
where a lot of you are coming from, and 

174
00:10:41,372 --> 00:10:45,660
where your interests lie. 
Well, what we see is that business is 

175
00:10:45,660 --> 00:10:50,240
beginning to look a lot like what's 
always been happening in science. 

176
00:10:50,240 --> 00:10:52,634
So though, you know businesses of 
requiring data aggressively and keeping 

177
00:10:52,634 --> 00:10:55,420
it around indefinitely, in case it 
becomes useful. 

178
00:10:55,420 --> 00:10:58,800
They're beginning to hire people that 
have training and skill sets that look a 

179
00:10:58,800 --> 00:11:01,868
lot like whats been important in science 
for a long time, especially in 

180
00:11:01,868 --> 00:11:06,152
mathematical depth. 
And they're beginning to make decisions 

181
00:11:06,152 --> 00:11:09,650
with this data that are very empirical, 
so we're always wanting to make up every 

182
00:11:09,650 --> 00:11:13,570
decision with a a clear case based on 
data. 

183
00:11:13,570 --> 00:11:16,276
And so for these reasons, I think you can 
take the lessons learned in science and 

184
00:11:16,276 --> 00:11:19,330
apply them in business, and actually vice 
versa as well. 

185
00:11:19,330 --> 00:11:21,850
The one thing where science is lagging 
behind business is in the adoption of 

186
00:11:21,850 --> 00:11:24,069
technology. 
There's been, there's been 

187
00:11:24,069 --> 00:11:27,620
proportionately a lot less spent on IT 
infrastructure in science than there has 

188
00:11:27,620 --> 00:11:31,010
in business. 
And so there's, it's a great time for 

189
00:11:31,010 --> 00:11:34,888
this, crosspollination of ideas between 
both fields. 

190
00:11:34,888 --> 00:11:38,248
Okay, and so you know going back to this 
first slide I gave, e-science and 

191
00:11:38,248 --> 00:11:42,029
data-science have essentially everything 
in common. 

192
00:11:42,029 --> 00:11:45,285
So we might use examples interchangeably 
between the two. 

193
00:11:45,285 --> 00:11:46,481
Okay. 

