[MUSIC]. Welcome back. So I want to talk a little bit about how the term "data science" relates to other fields of science. And in particular, I want to introduce this term eScience, which through first approximation you can think of as equivalent to Data Science. So while the eScience is associated with Astronomy and Oceanography and Biology data science has been adopted more in business. But they involve a lot of the same concepts. So let me tell you what's going on in Science. So for thousands of years, you know, scientific inquiry has been empirical, right. You observe the natural world, or maybe in some cases replicate the natural world in a controlled environment in a laboratory. And make observations about that. In the last few hundred years science has accepted theoretical models as a valid method of inquiry. One that is reinforcing empirical methods so you know new theories suggests new experiments and the theories help explain the observed data that you get from the experiments. In the last 50 years or so high speed computation has emitted an entirely new method of scientific inquiry. Soon you can simulate in the computer phenomena on that otherwise couldn't be re, you couldn't, you can't observe directly. And you can't reproduce in the lab and even the theoretical models become too complex to solve analytically using, you know essentially paper and pencil. Right, but you can actually start from initial conditions and run the simulation to get a result. So this is maybe what goes on in the interior stars or the shift of tectonic plates or the evolution of the universe or the effects of the ecology on some species dying out, and so on. So that's fine. That's three methods of inquiry. But in the last ten years or so, there's been arguably a fourth method of scientific Inquiry, which is to acquire massive data sets from instruments or from simulations. And then explore these data sets using new algorithms and infrastructure. And so eScience is really about massive and complex data, data large enough to require you know, automated or semi-automated analysis. You can't look at it, you can't inspect it directly, okay. And so the relevant tools here are the same as those for data science, you know, databases, visualization scale out computing. Maybe the new sequel systems. Machine learning techniques. Web services and so on. Okay. So the way, that, this, this, idea of the fourth paradigm, there's a book that's in the reading list that you can refer to here. And there's a lot of, just some other articles in your reading list you can also refer to. The story has been told lots of ways. The way I like to talk about this story is that science has always been about asking questions but conventionally it was really about querying the world, right? You would, sort of, have data acquisition or activities, experiments or field studies. They will couple to very specific hypothesis where you have the question in mind first and we click the data. But eScience has really shifted a bit where now you're kind of downloading data on mass, you're downloading the world first putting some sort of representation in the computer. And then creating that database to test your hypothesis and so it's the, the data can be acquired independent of any specific hypothesis in some case. Okay. And this is due in part to the cost of data acquisition dropping precipitously, thanks for advances in technology, right. So the telescopes you can build now, that we'll talk about in the next couple of slides, can acquire at enormous amounts of data at very high resolution. And in the life sciences, you have sort of laboratory automation and sort of high-throughput sequencing. In oceanography the sensors are getting cheaper. The models thanks to advances and thanks to more's laws. And thanks to advances in computing. The simulations you can run are getting bigger and higher resolution. And producing, and therefore producing larger, and larger amounts of data. And so on, and so you know, the rate at which data can be produced has far out paced the rate that we can analyze it or come up with the questions we need to ask about it. Okay. And this suggested a new approach to science. So let me give some examples, so we said that eScience is driven by data more than by the computation. Alright so some examples on the size of the data is that's coming on. The Apache point telescope that was the primary instrument for the Sloan Digital Sky Survey that we might refer to multiple times in this course. Produced 80 terabytes of raw new data over, a seven year period. You know, at the time, this is a pretty significant data size, and even by many standards is still today. The next generation of this the next generation project that's in the same sort of spirit as Sloan Digital Sky Survey is the large synoptic survey telescope. So this guy can produce 40 terabytes per day. it will do so for over a ten year period. So in total 100 plus petabytes and producing a single amount of data that's soon the skies will be produced in over a total entire lifetime. It can be produce that over every two days. Okay. And so this is a pretty staggering amount of data, and requires a, a pretty different approach. One thing I want to mention maybe about Sloan Digital Sky Survey, what they actually did here, was to take the images, cook them, right, extract the relevant objects from it, put all those objects into a, an off-the-shelf or a relational database. In fact, it was Microsoft SQL server, and critically host this database online and serve it out over the web, and this require a pretty significant investment in infrastructure. But as a result f doing this, of making all the data public and queryable, it became the most productive astronomy facility in history. Right, so the number of papers that have been produced on this data is on the order thousands. In the original, you know, PIs of the project, the principal investigators of the project had, sort of, maybe only 100 papers in mind for the data. And the other 4900 papers that have been written all came from external partners writing queries against this database. So it's just a wild, wild success. Now, the problem is, is that the same technology stack in, into some extent even the same approach. It is difficult to apply in this case of large synoptic survey telescope. The reason why this guy is producing so much more data is not just because it's much higher resolution and it can perceive a much deeper field in the sky. But also because it's returning to the same point in the sky frequently, every three days, and so this allows you to look at things change over time. So Asteroids, comets, you might catch, super novas and so forth. Okay, and by comparing these images in the time series, there, there's all sorts of new questions you can ask. Okay. So, both because of the science that they're going to do. And because of the sheer scale. and because of some of the complexity of the, details of how the data is acquired. The Existing, the previous solution won't work. And so, this is motivated in a whole new area of research to study, data management techniques and data analysis techniques to support this project. So in Life Sciences, these high throughput sequencers are capable of producing, you know, terabytes per day when run continuously. And you know, major labs that do this work, such as the Joint Genome Institute, have 25 to 100 of these machines running all the time, alright. So this is spitting out an enormous amount of data for, well, I was going to say a variety of samples. So it, so it may be individual organisms or it could even be samples from the environment where there's no particular one organism in there but there's an entire population. Alright, so for a variety of uses these guys are are able to spit out the data. In oceanography, the regional scale nodes of the NSF Ocean Observatories Initiative is a project led here at UW. Ocean Observatories Initiative is a multi-institutional partnership. the regional scale nodes part is run at the University of Washington. So these, this project does involve laying, you know, a thousand kilometers of fiber optic cable on the seafloor. Connecting thousands of instruments in chemical, physical and biological thousands of chemical, biological and physical sensors. Including live video from the seafloor to measure to monitor volcanic activity. Okay, so again, the database the, the if not a relational database. [LAUGH]. The data sets and data infrastructure required to support this effort is significant and has motivated a lot of new research. Alright. in the information space there is a lot of science to be done on the web itself. And so, just the web, a single computer can read 30 to 35 megabytes per second from one disk, and so it would take about four months just to read the entire web. So new clusters of machines. So summing up a little bit eScience is about the analysis of data. So the automated or semi-automated extraction of knowledge from massive volumes of data. And so your main instrument for looking for answers is the, or the algorithms and the technology as oppose to direct inspection. There's just too much of it to look at it yourself. But it's not just a matter of volume as we'll talk about in the next segment. this is another link back to what's going in business. Right, there's these, there's this concept of big data and there's the three V's of big data that we'll talk about a little bit more next time. But let me just mention them here. Where volume, sort of the three V's are volume, variety, and velocity. They will deal about and I will give you the, where this stuffs came from in the next segment. Involving first just into the number of rows and the number of bytes at the serious scale. Variety is perhaps a number of columns and dimensions but you know in science for example lot of in allied sciences particularly you will have experiments that involve accessing multiple databases as well as multiple sensors. you're own data that you've collected and that of your colleagues and the integration task of putting all these data together is a significant model like. Even when the actual scale on the data is not necessarily all that bad, okay? So those are the complexity of the data. And then the velocity, you know, we saw what the large survey telescope that. You know although the, the, the scale itself is enormous, that fact that 40, 40 terabytes are being collected every two days means that the infrastructure needs to keep up with that pace. And just transferring that data from the telescope facility to the data analysis facility is an engineering challenge. Okay? And you'll also see other V's here, with things like veracity. You know, can we actually trust this data? Okay, so a bit more of that next time. So to summarize here, science is in the midst of a generational shift from a data poor enterprise, where you can never, you know there's never enough data. To a data rich enterprise where there's so much of it you don't know what to do with it. And as a result, you know data analysis has replaced data acquisition as the new bottleneck to discovery. Right? So it's not the cost of going out and getting the data, it's the cost of actually analyzing the data you might already have. So this is fine, but what does this have to do with business, which is probably where a lot of you are coming from, and where your interests lie. Well, what we see is that business is beginning to look a lot like what's always been happening in science. So though, you know businesses of requiring data aggressively and keeping it around indefinitely, in case it becomes useful. They're beginning to hire people that have training and skill sets that look a lot like whats been important in science for a long time, especially in mathematical depth. And they're beginning to make decisions with this data that are very empirical, so we're always wanting to make up every decision with a a clear case based on data. And so for these reasons, I think you can take the lessons learned in science and apply them in business, and actually vice versa as well. The one thing where science is lagging behind business is in the adoption of technology. There's been, there's been proportionately a lot less spent on IT infrastructure in science than there has in business. And so there's, it's a great time for this, crosspollination of ideas between both fields. Okay, and so you know going back to this first slide I gave, e-science and data-science have essentially everything in common. So we might use examples interchangeably between the two. Okay.