[MUSIC] So let's talk a little bit about what distinguishes the term data science from other related fields. So one related field is business intelligence. Business intelligence systems are associated with a couple of concepts. One is a data warehouse, and the other is a set of dashboards or reports that consume data from the data warehouse, and are used to answer particular questions. So both of these components require a lot of upfront effort to design and build, and are therefore, not too adaptable when requirements change. Okay, and so therefore a software stack designed for business intelligence may or may not be appropriate for any particular data science problems, where changing requirements are considered the norm. And so it sort of warrants a new term, is that business intelligence became associated with a particular approach to a particular set of problems, and a data science in some sense broader. Okay. The other point I'd like to talk about business intelligence is that, the BI engineers are not typically expected to consume their own data products, and perform their own analysis, and make, and make the business decisions themselves. Usually they're building tools for others to make decisions with. Okay, as a data scientist, you'll be doing both. So, what about statistics? Well, statistical methods are at the heart of what a data scientist does, day to day. But a statistician will typically be comfortable with the, with assuming that any data set they encounter will fit in main memory on a single machine. And this makes sense, because the whole field was born out of the need to extract the most information possible from a very sparse, very expensive to collect and, and to typically therefpre, very small data set. Okay, so if you only have 20 patients in the world with a particular disease, you can't just go find 20 more cheaply. So therefore, you need to come up with new mathematics to squeeze as much information as you can out of the 20 you already have. But that's not always the problem anymore, right? So as we shift from a data-poor regime to a data-rich regime, the set of challenges move from the need for new mathematics to squeeze information out of a data set to new engineering to even handle or process very, very large data sets. Okay. However, some of the methods, some of the models you'll build are, are the same in both cases. So database experts, database programmers and, and administrators bring a lot of skills to the table that make them appropriate for data science tasks, but there's a, there's a focus on a particular data model. Which is usually the relational data model. So this is rows and columns. So we have data coming from sources that are video, or audio, or even text, or to some extent even graph graphs, nodes and edges which we'll talk about. a relational database may or may not be the right tool and even the concepts that transcend any particular database system may or not be appropriate. And we'll sort of explore when and where it isn't appropriate as we get into the course. Okay, so visualization experts also bring a lot of skills to the table. But like statisticians, are historically, less concerned with massive scale, data that spans many hundreds of machines. and then, finally, machine learning, is perhaps the closest to data science. But here and we'll try to make a, more of a point about this later. The, the, the, as a proportion of the time you'll spend on a data science problem, actually choosing the, the right model or algorithm, machine learning technique, and applying it and running it is a fairly small fraction. What you'll be spending much more time on is the preparation of the data, the manipulation of the data, the cleaning of the data the wrangling of the data, as some have been saying. And, for this, machine learning techniques are, are not particularly relevant. And so, this falls back more towards the the database managers. Managers, the database experts and database programmers, okay? So, there's a lot of courses that could be considered data science courses. Some of these use data science in the name, some of the newer ones. others have been around for a long time but are, but are obviously in this, in this same space. And so, I want to spend a little bit of time describing the dimensions by which you could describe these courses, and then choose a particular point in this design space that we'd use to motivate this course. Okay. So as a preface, let me show you this quote from Aaron Kimball, who's a CTO at the CTO at, at Wibidata. And so he said to me that he worries that the data scientist role is perhaps like the mythical webmaster of the 90's. That they were expected to do everything. Alright, the web manual, companies knew they needed to get on the Internet in the mid-90s but they didn't know how. And so they said, well, you know, we'll hire a webmaster, problem solved. Alright, the webmaster will write all of the content for the website, they'll do the design and manage the user experience, they'll write the code that will wire the website to the order fulfillment system in the back-end. they'll actually structure the pages and do the navigation they'll do the logging required to make sure that the site stays up all the time and, and has reasonably high availability. they'll design the schema to hold the data that will be served out through the website and so on. And so it wasn't really feasible that you were going to get this in a single person. And so instead the, the Internet strategy became a, a broader team. Similarly, that might be what we see happening with, with data science. But here's what it means to me. The term data science tells me, that if you're a data-based administrator and your skills are solely about relational databases, the current trend is you will need to learn more about unstructured data and statistical modelling. If you're a statistician, you will need to learn to deal with data that does not fit in memory. If you're a software engineer whose used to sort of building systems and working with files directly, you'll learn, you'll need to learn some statistical modelling and how to communicate your results to your managers. You'll need to work with these data sets and actually use them to, to make decisions. And if you're a business analyst who is trained to make your decisions based on data, you're going to need to start understanding a little bit more about the algorithms and tradeoffs, especially at scale. And for a couple of reasons. One is the cost changed dramatically based on the technology you're picking. What's happening with cloud computing that we'll talk a little bit about, and what's happening with these algorithms is that, you know, you can, you might be able to get an answer, but it may cost more or less than, than, than it did five years ago. The other reason is that, you know, as we do more fly-by-wire business, meaning, you know, we, we trust algorithms more and more to make some decisions for us they become these opaque black boxes. And, if you don't understand what's going inside, going on inside that black box, you're bound to misinterpret the results. And so, it's not, it, it's no longer safe to just, sort of, throw your trust over the wall to some algorithm. Or to, or to your staff that's running these algorithms. You may need to sort of understand, internalize the trade offs and choosing one model versus another yourself. Okay. So here are the dimensions by which I like to describe these different courses. The first one is breadth. and so I divide breadth into tools versus abstractions. And so every sophisticated course would prefer to cheat towards abstractions, right? You want, you want concepts that transcend any particular implementation. However, what students are interested in is, hands on experience using tools that they can use, you know, tomorrow at a job. And so you always have this tension between these two. Alright. And so some examples here are, you know, Hadoop, which we'll talk about, is an implementation of an abstraction called MapReduce. And then, the, the MapReduce abstraction certainly transcends its particular implementation in, in, in Hadoop. Okay. And so here, as we'll, as I'll mention in, in the next segment I want to cheat towards abstraction whenever possible to make sure that there are assignments that give you the hands-on skills that people are interested in. Alright, the next dimension here is depth, and so by depth, I intend the distinction between structural manipulation of data and statistical manipulation of data. Okay, and so here you can think about the elation to algebra as a structural formalism, a formalism for man manipulating data structurally while the linear algebra is perhaps a formalism of manipulating data statistically. Okay and so here, try to strike a balance, but I actually mean towards structure, and I'm going to defend that position in the next segment. The next dimension you can think about is scale. And so here is you know, on one end is, is yes, its a, a main, main member on a single machine versus what I'll call cloud, meaning the, or it might require hundreds of machines to, to work on it. And here I cheat towards cloud and the reasons that I've already sort of described are that you know, it's no longer safe to assume the data fits in main memory, and to train people to work only with data of that size, you know, the, the whole world changes when you start to move even two machines, let alone 100. and to not have, give you some exposure to that change, would no, would not equip you, to be, to be an effective data scientist. Okay. And then the final dimension I use is sort of the target audience, right? Is this for hackers or more for analysts? And by hacker I mean, you know, you already have significant programming experience and you're looking to sort of round out your skills in some of the mathematics. or are you more of a technology decision maker, who is trying to bring a little bit of technical depth? And here I like to actually strike a balance. I don't want this course to be solely assuming that you, that you are a a seasoned developer. but nor, nor can we sort of ignore all programming whatsoever. So we're going to try to strike a balance between these two. Alright, so here's the choices we, we sort of made in this course. So one is we cheat towards abstractions, we cheat towards structs. we definitely like large scale. And then we, I, I say we'll strike a balance. But we actually cheat towards the analyst side. We favor, we favor the fact that there are going to be analysts in the room who don't necessarily have significant programming experience. And I've already gotten a lot of questions from folks over email who say, hey look, you know I haven't done, been doing programming day to day, am I going to be able to take something away in this course? And I think the answer is, is yes. Although there will be some programming, so, so be ready.