[MUSIC]. Welcome back. Last time we talked about three out of four of these dimensions in describing how we designed this course in data science. And so, in this segment, I want to talk about this last dimension of what I, what I call structs versus stats and so this is. The relative importance of data manipulation versus deeper mathematics. And you can see that I've sort of put the, the dial here a little bit to the left, and I'll try to motivate that in, in the next few minutes, alright? So, we already saw one example of this in the first segment, where I gave some examples of data science from you know, recent history and one of these was Nate Silver's prediction of the electoral college votes for the 2012 US Presidential Election. And if you recall you know, this was, this prediction was accomplished by essentially taking the average of the state polls for each state. Okay, so it really didn't require a sophisticated statistical model and yet it had massive impact. Okay, so you know a quote that I think sums this up a little bit, comes from Aaron Kimball at a company called WibiData, he, he says, you know 80% of the analytics is really just sums and averages and so if you can get these. What he means by this is if you can get these sums and averages right, if you can do it at any scale on any data that you might see, then you can always sort of build up more and more more advanced techniques. Everything sort of boiled down to just sums and averages. Okay, so I think this is a motivation for why focusing on data manipulation, which can, which typically is associated with being able to express sums and averages, for example you know what you can do in a database query which we talk about in the next couple of lectures. That gets you a pretty long way, alright? It is the 80% of the problem, so another way of looking at it is there is three main tasks involved with you know a data science project. There's preparing to run the model. Running the actual surgical model and then interpreting the results and communicating it. I got the animation out of border here you can ignore that red. But the point here is you know, again Aaron Kimball from a conversation[LAUGH] with him where he got this, was the you know 80% of the work is really in this first step where you are gathering data and cleaning it and integrating it and restructuring it transforming it, loading it and so on. So, verifying all these verbs you see here, this is the hard part, right? And so, actually running the model or even choosing the model and then running it doesn't tend to keep people up at night in practice, okay? So, and then the joke is perhaps that the other 80% of the work, you know, implying that there's sort of 160% of a normal task is in data science is in this interpreting the results. So, this is the visualization and the communication and the explanation of the results. Okay, so this is another reason why I want to focus in this course on data manipulation tasks that are associated with this first number one task. Okay, you know another way of looking at this is a quote that is now really old, right? So, this is 12 years old or so at, at the time of this recording from Doug Laney. And this is the document that first coined this notion of big data of being the 3D's of volume, velocity, and variety and we'll talk about that in a couple of segments. But he has this quote, you know, no greater barrier to effective data management will exist than the variety of incompatible data formats, non-aligned data structures, and inconsistent data semantics. So this is what the database community, you know, my community calls the data integration problem, this is the hard part. And so, he was saying this back in 2001 and I would argue that it's still true today. This is the greatest barrier, and in the context of this, he was talking about this notion that variety being harder than volume or velocity, and I'll explain more about what those Vs mean in a couple of segments. Alright, so another been yet here, is something that we like to ask the scientists we work with. So these are, you know, astronomers and oceanographers and biologists. We ask them sort of informally, how much time they spend quote, handling data, as opposed to, quote, doing science? Now, you know, we let them interpret these quotes however they want, but what we mean by doing science, you know, choosing the statistical method or designing a statistical model. They absolutely consider it part of their science and so, we mean by handling data is altogether crap, you know, the format conversions and so on. And so, what do you think the most common answer is here? Or you can guess to yourself for a second but they, they don't even blink and they say things like 90% and so, this number should you know give pause, right? This is taxpayer money that goes to federal funding agencies, to come back to pay some post-doctoral fellow to spend 90% of her time, doing something that she doesn't even consider science. And so, this is the why it's really important as a data scientist to focus on this problem. Now, you might say, well that's just science, what about business? But as I will try to make the point throughout the course there's a an increasing alignment between what's going on in business and what's going on in science. Okay, and we, we will talk about that more in a couple of segments. Alright, so if 90%of the problem is handling data, you know boy we have spent a lot of tension on that. Alright, so you know another argument that sort of follows on the first side I gave, is that struck's you know, the, the data manipulation platforms and databases in particular, actually go a pretty long way to be in able to express more advanced things and this isn't just a matter way you can express anything you hear sums and averages. It's also even fairly advanced techniques you can, there's increasing amount of interest in figuring out how to get this stuff into the database. Okay, so this is the a side of thinking from Christian Grant, where he argues that look you know, if you consider databases versus statistical packages, such as SAS or Matlab/R or SPSS. you know, this is what they are doing now, their taking, they are downloading data to use in their[INAUDIBLE] package. frequently under the assumption that well of course I have to, right of course the only thing we can possible express this. Well most of these sand pack, here, just the first thing you will do is, read the data off the disc and load into memory and then start calling functions on it. Well if it, it increasingly data sets simply don't fit in memory on a single machine certainly not on your laptop. I'll take a couple of choices here. Either you shift into some kind of, fancy cluster diversion of the tool for which they exist for things like SPSS and Matlab although they are quite expensive. Or you sample the data, so that you all may ever, only can work with the subset that actually does fit in memory and you'll see this to be very, very common. It's that it's just powerfully course to, to, to take the sample of the, the data in order to be able to work with it efficiently, right? But the point here is that this isn't really required if you use different packaging. In this case, if you, if the argument here is that if you use databases. If you can figure out how to, how to perform your task in the database, you'll get the scalability for free. Moreover, you know, these tool kits don't have, don't necessarily have any kind of notion of parallelism, right? So even if it does fit in memory, every machine you buy nowadays has you know, at least four cores in it and probably more like eight and soon to be 12 and 16. So to take advantage of all those cores on your problem is, is something you're going to be looking for in a package. And this something the databases can do automatically, most databases not all. In fact, the, the ones you may be familiar with my SQL and sparse matrix do not, but other databases will and we'll talk more about this. Okay, so you get parallelism for free if you can use a database and you get scalability beyond the size of main memory for free if you can use a database and that perhaps is a big if and we'll talk about it. Okay, the leading example here, and you'll actually do this as part of a homework assignment, but you know, can you express matrix multiplication in SQL? and if you can, then I'd argue, well, hey, any formula that you can express using matrix multiplication, you can perhaps express in SQL, by doing this over and over again. Okay, and the answer is yes, and in fact the simplest version of this is, is pretty straightforward. so if you haven't ever seen SQL before, don't worry, we will talk a little more about this. But if you have, you know, bear with me, imagine you have two matrices, A and B, oops, I'm using the wrong device here. Two matrices using A and B and what you want to do is find all of the you know and the, the the representation of this matrix here is a row excuse me, row ID, column ID, and value. Alright, so that's a relation. Now, this is a very inefficient relation if your matrix is dense. And I'll let you think about what, well I'll tell you why and you think about it a little bit more as well. Is that, you know, an implicit representation of this only has, let's say you have in rows, let's say you have five rows and six columns, then you only need the 30 values, 5 times 6. But here you're doing, you have to do 30 row ID's plus 30 column ID's plus 30 values. So, you sort of triple the size of your data relative to, you know, efficient name memory representations. So why would you do that? Well, it turns out that a lot of matrix's in practice are sparse, and I put that word right up here on top. In a sparse matrix, not all the cells actually have a value and so you don't actually need to store them. And so this representation in terms of you know, explicitly having a row ID, a column ID, and value turns out be pretty efficient. Okay and in fact, sparse matrix solvers, this is exactly the kind of representation they use internally. Alright, so if you have a sparse matrix and if you encode it in, if you represent in a database then expressing matrix multiply is not too bad. You, what you want to do is find all the columns, you know, for each column number in, in the, matrix A, find the corresponding row number, in column B. and then add up all the all the contributions to the new value, and also a diagram of this after. In fact, you know, let me skip going into too much detail about this right now because I'm going to talk about this in detail in preparation for the homework while you'll do this. So right now, I guess take away, what I want you to take away is that representing matrices inside of a database sounds very unusual. It's actually not the world's worst idea and in part of the readings from this mad skills paper, you'll try to, you'll see why. So, right now, I just want you to take away that it can be done and it's not necessarily a terrible idea.