[MUSIC] So, let's go through a few more examples. So, if you were asked to decide how important a particular scientific paper was relative to other papers, how might you go about doing that? One way to decide between say these two papers here, I'll mark one with a blue circle and one with a red circle is to wait until other papers start to site these, and count up the number of citations. So, here you know, the paper marked with the blue circle has had four other papers bother, you know authors have bothered to read this paper. And some therefore must have had more impact inside of the community than the ones marked with a red circle. But if we wait any longer and that's indicated by the one in the darker blue color, if we wait even longer you might get even more people citing this intermediate paper. And so maybe we can conclude that well, over time, ultimately this one had more impact because it, this, this paper was influenced by this paper and all of these papers were influenced by this one. So, therefore perhaps we change our answer and that this one is more important. So, how do we decide between these two interpretations? Well this problem looks a lot like the problem of judging the relative importance of pages on the web. And so one thing you can do is say, well look, you know, a particular website is important, if a lot of other important websites point to it. And so here you can say a scientific paper is important, if a lot of other important papers point to it. And this method that Google proposed and implemented and ultimately led to a pretty significant part of their success, was this algorithm page rank. And so page rank did exactly this. You add up all the weights of your neighbors. And give them to yourself, and then, pass on that weight to everybody that you link to. And you keep going with this, until you've reached some convergence condition, and you end up with the relative importance of these these pages. Okay, and so, this is a pre, this is a method that comes up quite a bit. Whenever you have a graph, it makes sense to potentially run paging algorithm on it. Even though it had nothing to do with, ranking things on the web, which, for which it was originally designed. Okay. Using the same data set, you can, Carl Bergstrom and Martin Rosvall created this visualization. So, ignore the importance question, just think about the graph of the citation network and doing analytics on it. One kind of analysis you can do is judging importance. Another kind of analysis you can do is this, and so what this is, is over time, run a clustering algorithm, which I haven't described what that is. But it groups, clusters of similar documents to, and then, map those clusters into the rel, these fields. And so, you can determine that with some probab, with, with some confidence, that this cluster represents medicine. Because it has the lancet and has other kinds of medical journals in it. And you can conclude that this cluster has, is, represents molecular and cell biology. But what's pretty striking is that if you lay this out on this timeline like this, you can see that some fraction of the molecular and cell biology community. And the n, neurology community, started to combine to form a brand new field of science called Neuroscience. And so just by doing this kind of analytics on this graph, right, to do this data science on this graph, you can uncover the emergence of new fields of science. And so I find that pretty striking. And they've gone on to do many other kinds of analysis on this a same data set. And the overall field of this kind of meta field of studying the scientific literature in order to draw inferences is called bibliometrics. This pen is, aq little wonky, bibliometrics, sorry for the bad b there. Okay, so what can't, what is not ameanble to data science? Well, you know, you might think food, and this is a paper in a fairly respectable journal that apply, applied some date-driven techniques to analyzing food paring. Okay, so what they did here is pretty interesting, right? So, induce a graph on the ingredients by saying that if two ingredients appear together in a recipe, you draw an edge between them. You're graphing vertices and edges. And if you've never really heard, if you're not familiar with graphs, you will be by the end of the course. But bear with me right now. so connect two ingredients if they appear together in some recipe okay? Build that big graph and now you can analyze it in ways that are similar to what we just talked about. You can look at the community structure. You can find the clusters within this graph and see if these clusters correspond to the well-known methods of, of food pairings, okay? And in some cases they do, and in some cases they don't. And then, the authors sort of show that they've uncovered things that were not necessarily known, but appear to be there in the data, right? And so, this data-driven approach, you know, the fact that, that there's web sites full of recipes online. Has allowed us to put things on, on a more quantitative basis that were previously simply, you know old wives' tales, essentially, okay. So, I thought that was kind of a fun example of, of an, an unusual use of data science, again, in a respectable journal. All right? So another example, just from the Last FM blog, they looked at the tags associated with the songs in Last F, Last FM. And used it to do some simple analysis of the emergence of genres over time, but you know, based on the popularity of the tags for songs coming from that time. And so you can see things like well post punk in red here came after punk in purple, which you'd hope the graph shows you'd expect. He also uses a kind of rise in here. Rock and roll over time, and them maybe a bit of a, a, a, a dip more recently. And so this is kind of interesting. But the, the other theme here we have is, that we mentioned in the previous segment is repurposing data, right. So, this data was collected simply to help with search, you know find similar music, and it's not being reused to sort of draw inferences about the emergence of, of entire genres. Okay. So, another example a few of you may be familiar with. Google was able to show that, by analyzing the search logs, the frequency of search terms, it could do a better job predicting the severity and the, scope of flu outbreaks. Than the Centers for Disease Control, and by ,better, here we mean essentially earlier. Right? I was able to, give more a head start to the, the, health community. Okay. And also, it was, it was sort of more accurate. So, how do they do this? Well, you know when you're getting the flu, it turns out that you want to search for flu symptoms, terms associated with flu symptoms more often. And by watching that uptick, you can predict that there's, that the flu outbreak is coming. Okay. So, that's great and they put in a this work, and they sort of published a paper about it and then they put up this interactive visualization. A lot of people just sort of analyzed [UNKNOWN] going forward. But just this year, you know some folks showed that he didn't do a very good job in this last year. So, scientific hindsight shows that the Google flu trends far overstated this years flu season. And the reason they think explains this is that there was lots of media attention associated with this year's flu season, because it was a, was a bit of an uptick, and so it got amplified. And so this caused people to search more flu, flu-related terms more often, perhaps because their worried about experiencing the symptoms perhaps. Because they are trying to understand more about the flu outbreak, perhaps because they are searching for articles. Perhaps because they are worried about their kids more. But it was a second order effect, based on the media attention on the problem which lead to skewed results and ultimately wrong answer. And so the point here is this is, great when their re-purposing data from the search engine to try and make predictions about something else. But it's, it's biased its [INAUDIBLE] it is biased data so you have to be careful with, what you conclude from it. Okay? So there are limits here. All right. So, another example also, also, analyzing web search traffic. That was done with perhaps a little more of scientific rigor, was done by some folks with the Microsoft research. And so, here what you are looking for is side effects associated with particular drugs. And so these results are pretty striking. So, what this graph is showing that is over time, a set users was around about a million that had, that had a permission to sort of monitor there there web search traffic. That when you searched for this drug, in green, what percentage of the time did you also search for terms associated with hyperglycemia symptoms? And the answer is somewhere around 5%. For this other drug it was somewhere around 4%. For the, in the background, for the average case it was pretty close to 0%. If you searched for both of these drugs, the odds that you also searched for hyperglycemia symptoms went up to 10%. Okay, so what's striking about this, is that hyperglycemia is not a known side effect of these drugs. But it seems impossible to ignore from the web search data. Alright, there's just no reason to believe that this could be explained by coincidence. And they develop this argument more in the, in the paper then I, then I just have there. But, so fine. So, re-purposing data, this is another example, a large data sets, that had to use web search, And those are probably the two points I want to make about that, but a pretty fun one. Okay, so the last example I'll give is a different take. It's more about prediction then data. But this from last October, If you recall there were six Italian seismologists who were convicted of manslaughter, for failing to predict a magnitude 6.0 earthquake in April 2009. And while the locals were concerned about the seismic activity, the researchers were deemed to be just too reassuring about the verdict. And so the point I want to make here is that, there is liability I mean, so this, this, this, this scientific community was completely aghast that this happened. And I'm completely aghast this happened, and pretty much everybody is. That you can imagine to hold researchers responsible for failing to predict something that is demonstrably and known to be impossible to predict, right? So, there's no seismologist on the planet, that would argue the earthquakes are maybe even remotely predictable. and yet the courts sort of decided that somehow they, you know, because they got the wrong answer, it's bad. But it does sort of bring up the issue that when you make a prediction, there's a certain amount of weight you're, you're going to put behind it. Whether, whether intentionally or not. And so understanding how confident you are about that prediction, is sort of an important part of the game plan. Okay. So, that was the last example. The, the, the themes that we saw come out here. We gave a couple examples of graph analytics. We say that databases were sort of useful in the Obama grounding case. So, a lot of examples of visualization and communicating these results, interpreting these results. we saw some examples of using very large data sets, other examples of using very small data sets, and not everything is about big data. A couple of bullets that aren't on here: We talked about, y'know, ad hoc interactive analysis. It's sort of, not just faster but different. Also supporting that is important. And then we talked about re-purposing data. Right. So, data collected by, perhaps by someone else for some other purpose, reusing that as a drawing for something else, that's a pretty common theme here. Okay. In the next couple of segments, we'll talk about how we organized this course and some of the design decisions we made in, in creating the material.