[MUSIC]. Welcome back to introduction to data science. So in this segment I want to talk about these four dimensions that I introduced last time and I want to justify the first three of them. And we'll talk about this one next time, okay. And by justifying what I mean is I want to explain why I, I've positioned the needles the the the way I you know, at the locations I have for this course. Okay. So the first point here is is this dimension of tools versus abstractions. And this may seem sort of obvious that we want to focus on fundamental concepts as opposed to specific tools. But I can appreciate the people that are taking this course and many other courses really want a sort of hands-on experience. And we're definitely going to try to strike a balance, but let me try motivate why I think its important to sort of focus on this angle. and to do this, let me tell one, one particular story that you see happen through the time and time again. So in this case we're talking about sort of databases and what is currently going on in the no SQL systems. Alright so, before 2004, you had, you know, the big three relational database vendors, plus some open source solutions like MySQL and Postgres SQL. And then, arguably a big event in 2004 was when Jeff Dean and His colleagues published his paper on MapReduce at Google. And if you haven't heard about MapReduce, we'll talk about it at length and if you have, bear with me. So, this was great. What this will allow you to do is process very very large data sets and it sort of rebooted the database feature set. So it really stuck on. It really focused on just scale out parallelism and that's it, none of the other features of databases. And this was exciting to a lot of people because they didn't have to sort of deal with the extra features that they didn't need in databases. Nor pay for the exorbitant license fees associated with databases. So this seemed like, boy, this is the right solution. Okay. And you know, a few, it took a few years, but a few years later you had an open source implementation of the ideas in this paper called Hadoop, led by some folks at Yahoo. Now, even in the same year, one of the earliest and most successful projects within the Hadoop ecosystem was this system called PIG? And what PIG essentially was, was a relational algebra programming environment or a dupe. And if you haven't heard a relation algebra don't worry, we'll talk about it. But let me convince you that rela-, notice our relational, relational algebra is the secret sauce within relational databases. And, so very really on project their that was deemed necessary in the Hadoop community. And was wildly successful, was to have relational style programming on top of this non-relational system. Okay, moreover you had other competitive, competing products like DryadLINQ, which well, well Dryad and then DryadLINQ which is an interface of Dryad. which also provided a relational algebra-oriented programming environment for large scale Parallel data processing applications. Then you had people literally put the language SQL on top of Hadoop. So instead of just the underlying formulas, it literally had the programming language you needed. Then a bit later you had indexing for Hadoop, which is another feature that databases have that we'll talk about. You had people talk about schemes and more sophisticated types of indexing which is two other things the databases have. And then you start now to see transaction processing being a very hot, very important topic in new SQL systems is how to support concurrent access at very large scale transparently. And this slide is perhaps a little bit old, it's not 2013 at the of this recording. And the Spanner system from Google is an important system to look at that we'll talk a little bit about later. Okay so, now this isn't to say that MapReduce was useless, we're going to talk about at length, and for a very good reason. It actually has some pretty important permanent contributions three of which I mentioned here. One is, you know, it was the first system to really emphasize fault tolerance. And the idea here, in a nutshell, is that when you're working with a 1000 computers at a time, for any length of time at all, a few minutes, a few hours. The odds of one of them failing in some way is extremely high. And, so, databases didn't typically have to worry about this because the assumption, first of all they weren't running on thousands of computers at once. And second of all they were sort of under the assumption that your queries would typically be pretty fast, okay? and, so, fault tolerance, during query processing. So you don't lose all the work you, you, you've started on when you're running a query, was something that the Hadoop and the MapReduce produced paper really emphasized. And has now been sort of accepted by the large community. The other notion which is a little more subtle is this idea of schema-on-reed. And what I mean by that is you know the way databases worked in the past and largely still work is. You know we design the thing called the schema which is a particular structure for new data. And then you job is to fix your data into that schema and until you do so we don't really want to talk to you. Database has nothing to offer you until you are able to fit in some sort of a schema. Okay and the observation was that well, look a lot of data doesn't had come pre- equipped with the schema. We don't have a schemer just sort of lying around, we had to do something with it. You know and it's huge, you know, It's many hundreds of terabytes or something. So, what do we do? Well, you know, one answer is, you can use these map reduce bases, bases z-, you know, Hadoop for this. But having to say that, before you're allowed to touch your data, you must load into a database, that was kind of a non starter for a lot of applications. Okay. And then finally, you know, this idea of user defined functions is something that all database, or most databases support. and it's the idea that you might want to do things outside of what you can do in a normal SQL query. You might want to write your own code and push it into the database. But the experience of having to sort of write and maintain and manage and use these things is not great. And that's why a lot of people put their logic inside the application as opposed to pushing it down into the database where arguably it could do more good. for reasons we'll, we'll talk about. Okay, and so I think, you know, MapReduce would have argued, you know, look, you can actually have, you can give the Java programmers what they want, a Java programming environment. And let them write scalable systems without forcing them to kind of use this crazy user, user-defined function interface that databases offer. Okay? So fine. So what's my point over all this? Well. You know, if we focus too much on tools, what you would get is a snapshot in time of what tools are important. As opposed to seeing that some of these features around databases are, you know, they sort of ebb and flow in their popularity. But their all, their sort of a permanent there a permanent value when you're reasoning about large-scale systems. Similarly you might lose track of what's actually novel and what's actually new in the midst of the conversation of about original databases versus SQL systems. Okay so I want to focus on these abstractions throughout the course and then we can. Great, now. Okay so fine, we're going to focus on the attractions, what are the abstractions of data science? Well, it's not clear that people really know yet. And I'll give my case for this is the next slide. but you know, the reason I don't think we really know yet is if you see these words being used like, Data Jujitsu and Data Wrangling and Data Munging. And you know, this is the real skill of data scientists, so they had to be able to wrangle data, well, what does that mean? Okay. So, my translation of this is, we don't really know what we're talking about yet. With that said, there's probably a few candidates we can consider here. So, maybe everything's a matrix and everything we want to do with data can be expressed in linear algebra. If you're a database person, maybe everything is a relation, and everything you want to is expressed in relational algebra. If you're more of an object oriented programmer. Everything is an object and we communicate between objects by sending messages back and forth through by calling methods. If you're more of a sysadmin type then, you know, everything's a file and we, we write bash scripts to process it. And if you're an R programmer, then maybe everything's a data frame that we call functions in this library. And Matlab, similarly with Matlab, everything's an array or a matrix or a vector I guess and their parlance and everything's a function on that, okay. So, I think of all these possibilities, there are two that stand out as likely candidates, as fundamental abstractions for data science. And those are the first two here. And the reason is, is that we see these abstractions appear over and over again, independent of particular tools. Now, relations and relational algebra are closely associated with databases, but as I argued a few slides ago and as you'll see throughout the course. You'll see this come up time and time again, and we even see it in say you know, object or in languages, or you see it in, R and so on, okay. So these are the two that remain the focus on, in this course. So now I'm want to motivate desktop scale versus cloud scale. and, you know the argument here for desktop scale is that, well, you know data science is really about the functions and the statistics. And the manipulation of techniques. So therefore we can sort of push large scale data into a separate course, or a separate category, and really just focus on the, the math and the functions. And I think this is a bit of a mistake. For a data science course and the reason is that you know, this is a fundamental limitation of a whole category of technologies. And R itself is included in that, although there's a lot of great work on how to sort of scale R up. But as it,you know in its basic usage what you do with or read a file load the whole thing into main memory on one machine and then call functions on that. And your data doesn't fit in the main memory on one machine, you kind of out of lock. Now, you can be clever and start use indices a kind of limit the data you need to access. And you can start to try to be parallel to take advantage of the fact that there's now, you know four, and six, and eight, and 12 cores in your machine. In your computers you'll buy nowadays. But trying to be clever and doing that yourself overlooks the fact that a lot of this. A lot of these techniques are pretty well-understood and already implemented in other systems. Okay, so being able to be cognizant of what other systems can do and take advantages of those flexibly. And you know write your application in terms of these other systems that already do scale out use a critical skill in data science. And so the point we made in this slide that is somewhat out of date although you can get the idea is that. Simple, I, simple tools that are available on every machine such as GREP, which if you haven't heard of GREP and you're a Windows user. And don't use GREP too often then this is essentially search a file for a particular pattern. but it searches it linearly, right and looks at every single line of the file and checks for the pattern. And so you can do a linear scan of a megabyte in maybe a second, and a gigabyte in a minute. And so on. And so at a very large scale data sets you can't do this linear scan anymore. You have to search in a more, in a smarter way. Okay? And you sort of have some cost over here that are probably hor, horribly out of date now. Alright? fine so, the point is large scale data is not just bigger it's different. It requires a different way of thinking about techniques. And it requires a different stack in technologies and to ignore, it's a mistake to ignore that in a data science course. Okay, and then this final dimension of sort of hackers versus analysts. And again what I mean here is you know doing am I going to require sort of deep programming proficiency in order to participate with this data science class. And in general data science kind of activities. And the answer to that is I don't think so, I think, I think we need. At least two types of people and really sort of a broad spectrum of people. And I, and this isn't really my idea, this, this often quoted report from the Mckinsey Global Institute you'll see this quote time and again. but they, but the people that use this quote tend to focus on this first part that talks about 140,000 to 190,000 people with deep analytical skills. But the second part of the quote is well you also need 1.5 million managers and analysts who know how to use the analysis to make effective decisions. Okay? so this means that it won't be just the programmers who are working in this space. And I wanted to think about how to design a course that could Appeal and inform both categories of people. Alright, and this is my last slide of the segment. The other reason why I think hackers vs analysts is that the line between them is kind of blurry nowadays. And technology can actually help you, right? It doesn't require a PhD in computer science or a bachelor's degree in computer science in some cases, to manipulate large data sets. And for in order to back up this claim you need an example from some of my work where. we have done some work to try to make databases easier to use for say biologists. And this really nasty looking SQL query that if you squint closely you can see that it's actually doing Interval arithmetic over genetic sequences. Right this, this is a pretty tough query to understand for even experts. This was written by somebody who doesn't do any programming whatsoever. She doesn't write a line of Python she doesn't write a line of pearl. She doesn't write a line R. And she's able to these SQL queries to process very large data sets. Okay so the fact that you know if you understand what's going on and if you can think in terms of some of these abstractions. and you understand your problem well enough, you can participate in the activity of manipulating our data sets. And doing data science even without a a deep background in software engineering. Okay and that's why I want to push this needle, somewhere over this way. I'll probably put this in the middle, I suppose. I'm not so much trying to focus on only the analysts, I just trying to make sure that they're included. Okay. Next time, we'll pick up at the last dimension.