[MUSIC] Welcome to Introduction to Data Science. My name is Bill Howe, and I'm the director of research for scalable data analytics at the University of Washington eScience institute. And an affiliate assistant professor in computer science and enginerreing, also at the University of Washington. So in this first segment, what I want to do is go through some examples of data science activities and projects in the recent past that I found interesting. And use them to sort of whet your appetites for the concepts that we are going to learn in this course okay. So, the first one I want to mention here is the presidential election from 2012. And I know you're probably sick of hearing about this if you're, if you live in the United States. And even if you don't you may be sick of hearing about it, but bear with me. So, this is a map of the electoral college, and each state is colored by the for the candidate that took the electoral votes. And the numbers represent how many electoral votes each state has. And so if you recall, what was interesting about this map at the time was that it led to a pretty significant discussion in the media about Data Science. Because Nate Silver of the 538 blog was able to predict this map perfectly before the election. Alright, and you know that discussion in the media, had, talked a lot about you know, what a genius Nate Silver was. And mentioned the sophisticated mathematics he was using, and how you know, he's sort of a wiz with these things. But, what I thought was interesting about this was that Nate Silver would be the first one to tell you that. The methods she was employing to make this prediction were actually pretty simple. Right? And so he says here in a series of quotes from blog posts around that time. This first one from October 26 was. The intuition behind this ought to be very simple. Mr Obama is maintaining leads in the polls in Ohio and other states that are sufficient for him to win 270 electoral votes. And what was funny was a few days later he becomes, sort of, more blunt. The argument we're making is exceedingly simple. Here it is, Obama is ahead in Ohio, right? It's not a magic trick. And so then after the election on November 10th, and he was shown to be right and got this, sort of, flawless prediction. The blog post that this last quote is taken from. Whoops, excuse me. That the last quote was taken from was describing why he started the 538 blog in the first place. And he says, look, you know, the bar set by the competition was invitingly low. Someone could look like a genius simply by doing some fairly basic research into what really has predicted power in a political campaign. And so what really had predictive power in this case was the state polls themselves aggregated right. So, historically the state polls aggregated did a pretty good job of predicting the outcome of the general election and so that's was he did. Now there wasn't In quantifying the uncertainty and certainly in presenting these results. There's a lot of beautiful interactive visualizations that he created in order to sort of convey these ideas to the public. And that's one of the points I want to make about this is that getting the answer in some cases is the easy part. It's in interpreting the results and in convincing others of the result. By presenting them usually through visualization, that can be the hard part. This is one of the themes that we'll come back to throughout this course, okay. so, so just to summarize it though, I'm not sure I said this. The simple methods plus enough good data. You know wins, that trumps more sophisticated methods in many cases. That's another theme we'll come back to. Alright. So something else related to the campaign before we move on to other topics. Was the system that the Obama campaign used for their data driven ground game to so speak. The Build you sort of target, direct, you know, targets to specific categories of users. And so what they did was they built and maintained a really significantly sized massive voter database. And used it to design these highly tailored messages to very, very specific groups. So, you know the mother of two in a small town in Ohio who tweeted about the environment and mentioned organic vegetables on her Facebook page. And you know, who had voted in 2008, and had registered on Obama's website but had never donated. Okay, you know, she would get a message from Michelle Obama that highlighted Barack Obama's environmental policies. Okay. And so in order to design these message, what you had to do was, kind of do add hawk hypothesis testing about what might work and what didn't. You kind of had to slice and dice this da, data at, kind of interactive speeds. And this is another theme that we'll return to is the need for these kind of adhoc interactive analysis. And the systems they use for this are pretty interesting too, you know, this was a SQL database, a very fast one called Vertica. And we'll talk a little bit about what makes Vertica special, I hope, toward the end of the course. but it is a SQL database and so SQL sometimes gets a bad name in data science context. And sort of the old guard, that can't be possibly used for analytics and doesn't really make sense in today's era, but, you know, don't believe it. Right? It has a role to play in many cases. And so here, you know, they did use Hadoop, right, to do the aggregate generations of anything not real-time, he says here. But for the speed-of-thought queries about the data they used, this Vertica database. Okay, and so we'll come back to systems, in, in a, several segments from now. Okay. So moving on this was around the same time, you know when Hurricane Sandy made landfall, one of the things that struck me was. The fact that visualizations of available data were starting to emerge in real time in response to the storm. And there's some very nice examples of people who have used Twitter data to analyze or to produce a map of where the power was going out. And in this even simpler case Josef Fruehwald got public data from local weather stations. And just from two different weather stations and just simply plotted them. And so this is the barometric pressure over the course of this, you know, few day period. Two days I guess in two locations, Atlantic City and Philadelphia. And so you can see this enormous dip is the storm passing through. You can also see the time lag between Atlantic City and Philadelphia. And you can also see the intensity is probably a little bit higher given that the barometric drop is more significant in Atlantic City. Okay. So couple things here. One is pulling data down from the web and re purposing it sort of in real time. Or at least in, in short time, not real time, to produce visualizations and then publishing those back onto the web. I think this is very much the characters of data science activities. In this particular example, there's not necessarily a large data set involved. But re-purposing data that was collected for a different purpose is, is a theme that we'll come back to, and again, we see the ad hoc nature of this as well. Fine, so another plot here is wind speeds and they sort of peak out at 40 at Atlantic City which is green. And you can see that Atlantic City is indeed more intense here, and the gray here is arrow bars. so again, another variance on the same data. Okay, so changing gears a little bit, this was a study, the title here is called The Expression of Emotions in 20th Century Books. And so what they were interested in is had the words that we choose to use in our collective literature changed over time. And eventually does that tell us something about, sort of, culture or civilization? I find the scientific inquiry sort of compelling, but what I think is most striking about this. And why I wanted to include this example, is that the methodology that they used here is pretty straight forward. You could do this yourself with not a significant background in either technology or in statistics. Or even if in you know, linguistics, or anything. And so this is what they did, right? So the first step is kind of a doozy. This is take all the books written in the 20th century and digitize them. Well that would be a non starter, none of us could do that but that's okay. Google's already done it for us and they've made the data available at this URL and so you can go check that out. What they've done is digitize the books, done the character recognition on it and produced these Ingram data sets. So these are tables of data where each row has an ingram and followed by the year and the counts of the number of times that the ingram occurred. This has already been broken down and processed into a form that's digestible, okay. And so what's an ingram. Well, it's pretty simple. A one gram is just a single word like yesterday. A five gram, an example here is, the phrase, analysis is often described as. Okay, and so in this study, they just ignored everything but the one grams. and then they took some subset of, some subset of those one grams, and assigned them a mood score. So, how did they do this? Well you can imagine that certain words are, are charged with a particular mood or associated with joy or sadness or fear and so on. And you can also imagine that synonyms of those words might also be associated with that mood. And so this analysis sounds non-trivial, and it is, but once again, that's already been done for you. There's a, a resource on the web called WordNet, where they've done this kind of effect analysis. And so the authors of this paper were able to take the digitized books from Google. or already broken down into Ingram, and the [INAUDIBLE] scores from WordNet. And then do this calculation, which, you know, may sort of looking intimidating if you're not used to staring at these mathematical expressions. But it's actually pretty simple. This is the count of a particular, word in the set, the set being the set of word net word, which is not as all the words. Only some words are able to be scored as mood. And then you normalize by the count of the number of occurrences of the word the. So why do they do that? Well you need to normalize over something in order to account for the fact that perhaps we just write more books in 2005 than we did in 1937. Or we've been able to digitize more books. So we need to normalize by that total. Well, why don't I just normalize by the total number of words? Well, the reason is that the, the word the is a better indicator of prose than the total number of, of words. And this is because we've also, apparently we've also started to produce more sort of captions in figures, and more sort of technical language. And more sort of formula and more expressions, more non-prose, utterances in these, in these, books. And therefore we can sort of skew the results. We really want to capture, in our language, when we actually write full complete sentences, however these words being used. Okay. And then you add those up and you divide by the total number of words in the set. And then there's one more transformation here that should look familiar to you if you sort of recall your high school statistics. And so you subtract the mean and divide the standard deviation. Okay. So this is normalizing with respect to a normal distribution. Okay. But that's about it. You know, there's a count, and there's a division. and then there's two data sets that you can pull from the web. And they're big but they're not exceedingly big. They fit in memory on most of your laptops nowadays. So it's a, it's a significant computational task, but nothing that requires Hadoop. You can do this in a weekend if you, had thought about it. Okay, so I find that pretty compelling. Fine. So, these are the results, this is joy words minus sadness words so this is the z-score for joy and sadness. And you can see that there's sort of a big dip after World War two. And that's one of the points they make in the, in the paper. And then you can see this sort of thing start to increase in the late 90s, okay. I won't try to analyze this for the scientific value, I'll just present the results. What I think is maybe more interesting is this one. So this is now emotion words total, minus random words total and there's a sort of prominent downward slope over time. So what is this, what is this, what's going on here. Well, apparently you can make the argument that we're using fewer emotion words over time. Okay. That said, there's a bit of an uptick in this red line. And so what does that represent? Well, that's fear words. And, y'know, you can imagine some of the reasons why there might be an increase in fear words, since the, since the 1980s. Okay. So this is pretty fun though this is a significant analysis that can be done just by taking these, these data sets that they didn't have to prepare themselves. All right. And then the other point I want to make about this. This is just a copy and paste of a segment of the papers that this paper cites. And I just was struck by the titles here, you know. Quantitative Analysis of Culture Using Millions of Digitized Books. Quantifying the Evolutionary Dynamics of Language. Frequency of word use predicts rates of lexical evolution through Indo European history. song lyrics, I mean linguistic markers. I mean what strikes me about this is that you know linguistics anthropology history, culture. These studies are becoming hard sciences by the virtue of data driven methods, right. So all science is becoming data science. Right? And therefore data scientists have a lot of power in this regime, it's a great time to be you know, a data geek okay. You know, there's being a journalism as well. I mean one point, I promised you for the slide in here about this. But you know when the Wiki leaks material came out, you know, what, you're not going to pour yourself a pot of coffee and pore over these materials. You know, print them all out and sort of go through them, one by one, you're going to write algorithms that do this kind of an analysis. You know, word-use analysis, look for email chains and dialogues, these sort of computational methods, in order to analyze that material. That's on journalism itself is a computational enterprise, is a data science problem, not a or at least amenable to data science technique. So, you know, as a data scientist the world is your oyster. Alright. So, let me pause there. And we'll pick up with a couple more examples before moving on in the next segment.