[MUSIC]. Okay, in this segment I want to talk about logistics of the course. So how we've organized this course is a guided tour of important trends, along with a deep dive into specific topics. And then there's a set of hands-on assignments that are intended to deliver specific skills and experiences. And that's perhaps the most important part. Okay. And so overall the course is not, you know, the challenge here was to design a course that would be broad enough to cover the topics that we want. And also inclusive enough that we didn't sort of have to dial it in for a very specific cohort. But the challenge then is that it's going to be very difficult for some people and others may find it, some aspects of it certainly routine. I'd be surprised if anybody finds the whole thing routine. If so, then I'd be surprised they took this course. [LAUGH] Okay, so the prerequisites here are pretty light, since we are trying to cast such a wide net. So some prior programming experience in some language is going to be, really critical. then, you know, we're going to use terminology from the basic college statistics or the advanced high school statistics. So when I talk about linear aggression, you should know what that means. You should also be able to sort of, look at some visualization of data and be able to understand what it's telling you. Okay. And then perhaps the toughest one has statistics. Perhaps the number one is to have some exposure to databases and databases concepts. And you know, if you're just starting out in college, that's not always an easy proficiency to have gained or an experience to have gained. But you know, it's not, the, the, we're going to couch a lot of the discussion in terms of databases. And in the relationships to databases, and so some idea of what that means, what they are, is going to be helpful. Okay. So, to that end, one assignment will involving writing SQL, and if you've never written SQL before but you understand databases a little bit. You will probably be able to power through the assignment. If you're an expert in SQL, there are some parts of it that might still be interesting to you. And two assignments will be required, or will involve writing Python. One optional, sorry, one optional assignment will involve sort of processing big data using Amazon Web Services. And here, you know, one of the reasons it's optional is that because of the varying skill sets, but another reason is that you'll have to pay out of pocket for the cloud resources. And the reason for that is there's, you know, 60,000 students who signed up for the course and we can't sort of pay for all of them. The good news is it will cost sort of less than $10 or so. Okay. And is optional, so if you don't feel comfortable with that, you don't have to do it. Alright. Then another assignment will involve, all right, in competing in a kaggle .com project, of a kaggle, participating in a kaggle.com competition using whatever you want. And so, this may or may not involve any programming. You a lot of valid assignments, you know, you can, you can certainly compete by using Excel and other kinds of Gooey tools. Okay. This last bullet probably isn't true so lets just ignore that actualyl. So learning objective here is i really want people to come out of this course being able to talk intelligently about the landscape of data science concept tools algorithms technologies. And this will be sort of a spring board to dive deeper into particular areas. So, for example machine learning. This is not a machine learning course, but you can dive deeper into machine learning by taking this course. This is not a database course, but you can dive deeper by taking this course, and so on. Okay? And I also wanted to deliver some hands on experience manipulating data. I don't know levels of people that don't have any programming experience and provides some specific experiences for those of you that do have some programming experiences. For example, the first Python assignment will involve competing some Cinnamon analyses using some twitter data. So, if you already know Python, the learning Python won't be much of a contribution to that assignment. But Perhaps this is the first time you've been able to work with the live twitter stream, okay? And so the end result of this is that we hope you'll be sort of an advanced beginner in a variety of data science topics. And as I've said, you know, the tough, the tough part here is sort of how to do something more than just superfiicial access given that data science encompasses such a broad area of, as we've discussed. As we think we put together pretty good program but you'll, you'll have to be there, to judge that. Okay. Alright, so, the risk of belaboring this, of the course the velocity here is been that the skills needed by the data scientist span a variety of areas, statistics programming, databases distributed systems, visualization. But the traditional organization of these topics is sort of vertical and is not ideal for becoming sort of introductory in Data Science, right. So in order to get introductory level knowledge in all these areas what you end up having to do is take an introductory course in seven different areas or something. So a lot of different courses. Okay, and so our goal is to try to expose and simplify the links between these different areas. Okay. As opposed to sort of narrowing our attention on what makes them unique, okay? Right. Alright, so you know, after taking this course you will not be an expert in, statistics. You will not be an expert of machine learning certainly, you will not emerge an expert in databases and or even NoSQL. nor will you sort of have programing preferences in all of these language. However, you will use all these tools, you will understand the basic concepts of all these tools and you will have applied. Not many of these tools, okay. The assignments well there, there is a will have on-line short quizzes during the lectures of which you've already seen some these finger exercises quizzes they will be a set of the full length offline assignments as I mentioned. And some of these assignments will be graded by some of the programming assignments will be graded automatically. some of the assignments that don't lend themselves to autograding will be assessed using the peer assessment tools. So an example of that is you're going to write up a description of your Kaggle solution in addition to submitting your score for the Kaggle competition. And other students are going to sort of grade whether, whether it's comprehensive or not. Okay. So, here's my background in one slide. So, I have a Bachelor's degree in Industrial and Systems Engineering from Georgia Tech. But, you know, all of the problems in Industrial Engineering tended to be about optimization and automation, which seem to require software. So, I sort of got more interested in Computer science so I, spent a couple years consulting with some big firms. Somerget does oil feed services, oil, oil field services, and Siebel does customer relationship management software. And you, probably have heard of Microsoft and Verizon and Deloitte as a managing consulting firm. And then I sent back to grad school, and got a PhD in Computer Science from working with oceanographers on query systems for large scale oceangraphic models. And then I spent a couple years working directly with oceanographers as kind of a data architect. And before coming to the University of Washington where now I lead a group in Scalable Data Analytics for the University of Washington eScience Institute. And also I'm an affiliate assistant professor position in computer science engineering. And so, there's a bit of a mix of very practical, kind of applied work, as well as my research agenda. And so I think that this data science trend that's occurring is sort of, strikes close to home with me. I think it's a, I think it's a great time for it.