1
00:00:06,380 --> 00:00:07,750
[MUSIC] So, let's go through a few more 
examples. 

2
00:00:07,750 --> 00:00:12,727
So, if you were asked to decide how 
important a particular scientific paper 

3
00:00:12,727 --> 00:00:18,909
was relative to other papers, how might 
you go about doing that? 

4
00:00:21,310 --> 00:00:25,720
One way to decide between say these two 
papers here, I'll mark one with a blue 

5
00:00:25,720 --> 00:00:29,990
circle and one with a red circle is to 
wait until other papers start to site 

6
00:00:29,990 --> 00:00:35,420
these, and count up the number of 
citations. 

7
00:00:35,420 --> 00:00:39,775
So, here you know, the paper marked with 
the blue circle has had four other papers 

8
00:00:39,775 --> 00:00:44,619
bother, you know authors have bothered to 
read this paper. 

9
00:00:44,619 --> 00:00:49,035
And some therefore must have had more 
impact inside of the community than the 

10
00:00:49,035 --> 00:00:53,957
ones marked with a red circle. 
But if we wait any longer and that's 

11
00:00:53,957 --> 00:00:58,117
indicated by the one in the darker blue 
color, if we wait even longer you might 

12
00:00:58,117 --> 00:01:03,121
get even more people citing this 
intermediate paper. 

13
00:01:03,121 --> 00:01:06,956
And so maybe we can conclude that well, 
over time, ultimately this one had more 

14
00:01:06,956 --> 00:01:10,909
impact because it, this, this paper was 
influenced by this paper and all of these 

15
00:01:10,909 --> 00:01:17,298
papers were influenced by this one. 
So, therefore perhaps we change our 

16
00:01:17,298 --> 00:01:21,300
answer and that this one is more 
important. 

17
00:01:21,300 --> 00:01:22,998
So, how do we decide between these two 
interpretations? 

18
00:01:22,998 --> 00:01:27,812
Well this problem looks a lot like the 
problem of judging the relative 

19
00:01:27,812 --> 00:01:34,699
importance of pages on the web. 
And so one thing you can do is say, well 

20
00:01:34,699 --> 00:01:38,534
look, you know, a particular website is 
important, if a lot of other important 

21
00:01:38,534 --> 00:01:43,978
websites point to it. 
And so here you can say a scientific 

22
00:01:43,978 --> 00:01:49,399
paper is important, if a lot of other 
important papers point to it. 

23
00:01:49,399 --> 00:01:53,943
And this method that Google proposed and 
implemented and ultimately led to a 

24
00:01:53,943 --> 00:02:00,040
pretty significant part of their success, 
was this algorithm page rank. 

25
00:02:00,040 --> 00:02:03,448
And so page rank did exactly this. 
You add up all the weights of your 

26
00:02:03,448 --> 00:02:06,724
neighbors. 
And give them to yourself, and then, pass 

27
00:02:06,724 --> 00:02:08,900
on that weight to everybody that you link 
to. 

28
00:02:08,900 --> 00:02:12,092
And you keep going with this, until 
you've reached some convergence 

29
00:02:12,092 --> 00:02:17,648
condition, and you end up with the 
relative importance of these these pages. 

30
00:02:17,648 --> 00:02:25,260
Okay, and so, this is a pre, this is a 
method that comes up quite a bit. 

31
00:02:25,260 --> 00:02:28,342
Whenever you have a graph, it makes sense 
to potentially run paging algorithm on 

32
00:02:28,342 --> 00:02:30,846
it. 
Even though it had nothing to do with, 

33
00:02:30,846 --> 00:02:34,625
ranking things on the web, which, for 
which it was originally designed. 

34
00:02:34,625 --> 00:02:40,926
Okay. 
Using the same data set, you can, Carl 

35
00:02:40,926 --> 00:02:48,704
Bergstrom and Martin Rosvall created this 
visualization. 

36
00:02:48,704 --> 00:02:52,400
So, ignore the importance question, just 
think about the graph of the citation 

37
00:02:52,400 --> 00:02:56,004
network and doing analytics on it. 
One kind of analysis you can do is 

38
00:02:56,004 --> 00:02:59,214
judging importance. 
Another kind of analysis you can do is 

39
00:02:59,214 --> 00:03:03,802
this, and so what this is, is over time, 
run a clustering algorithm, which I 

40
00:03:03,802 --> 00:03:09,954
haven't described what that is. 
But it groups, clusters of similar 

41
00:03:09,954 --> 00:03:15,660
documents to, and then, map those 
clusters into the rel, these fields. 

42
00:03:15,660 --> 00:03:19,944
And so, you can determine that with some 
probab, with, with some confidence, that 

43
00:03:19,944 --> 00:03:25,013
this cluster represents medicine. 
Because it has the lancet and has other 

44
00:03:25,013 --> 00:03:29,446
kinds of medical journals in it. 
And you can conclude that this cluster 

45
00:03:29,446 --> 00:03:32,655
has, is, represents molecular and cell 
biology. 

46
00:03:32,655 --> 00:03:37,278
But what's pretty striking is that if you 
lay this out on this timeline like this, 

47
00:03:37,278 --> 00:03:43,667
you can see that some fraction of the 
molecular and cell biology community. 

48
00:03:43,667 --> 00:03:48,359
And the n, neurology community, started 
to combine to form a brand new field of 

49
00:03:48,359 --> 00:03:53,076
science called Neuroscience. 
And so just by doing this kind of 

50
00:03:53,076 --> 00:03:57,172
analytics on this graph, right, to do 
this data science on this graph, you can 

51
00:03:57,172 --> 00:04:01,815
uncover the emergence of new fields of 
science. 

52
00:04:01,815 --> 00:04:06,731
And so I find that pretty striking. 
And they've gone on to do many other 

53
00:04:06,731 --> 00:04:09,419
kinds of analysis on this a same data 
set. 

54
00:04:09,419 --> 00:04:12,457
And the overall field of this kind of 
meta field of studying the scientific 

55
00:04:12,457 --> 00:04:16,330
literature in order to draw inferences is 
called bibliometrics. 

56
00:04:16,330 --> 00:04:29,030
This pen is, aq little wonky, 
bibliometrics, sorry for the bad b there. 

57
00:04:29,030 --> 00:04:34,420
Okay, so what can't, what is not ameanble 
to data science? 

58
00:04:34,420 --> 00:04:38,776
Well, you know, you might think food, and 
this is a paper in a fairly respectable 

59
00:04:38,776 --> 00:04:43,198
journal that apply, applied some 
date-driven techniques to analyzing food 

60
00:04:43,198 --> 00:04:49,330
paring. 
Okay, so what they did here is pretty 

61
00:04:49,330 --> 00:04:54,701
interesting, right? 
So, induce a graph on the ingredients by 

62
00:04:54,701 --> 00:04:58,292
saying that if two ingredients appear 
together in a recipe, you draw an edge 

63
00:04:58,292 --> 00:05:03,130
between them. 
You're graphing vertices and edges. 

64
00:05:03,130 --> 00:05:06,120
And if you've never really heard, if 
you're not familiar with graphs, you will 

65
00:05:06,120 --> 00:05:09,008
be by the end of the course. 
But bear with me right now. 

66
00:05:09,008 --> 00:05:16,190
so connect two ingredients if they appear 
together in some recipe okay? 

67
00:05:16,190 --> 00:05:19,070
Build that big graph and now you can 
analyze it in ways that are similar to 

68
00:05:19,070 --> 00:05:23,090
what we just talked about. 
You can look at the community structure. 

69
00:05:23,090 --> 00:05:29,189
You can find the clusters within this 
graph and see if these clusters 

70
00:05:29,189 --> 00:05:37,952
correspond to the well-known methods of, 
of food pairings, okay? 

71
00:05:37,952 --> 00:05:40,200
And in some cases they do, and in some 
cases they don't. 

72
00:05:40,200 --> 00:05:42,988
And then, the authors sort of show that 
they've uncovered things that were not 

73
00:05:42,988 --> 00:05:46,230
necessarily known, but appear to be there 
in the data, right? 

74
00:05:46,230 --> 00:05:50,582
And so, this data-driven approach, you 
know, the fact that, that there's web 

75
00:05:50,582 --> 00:05:55,603
sites full of recipes online. 
Has allowed us to put things on, on a 

76
00:05:55,603 --> 00:06:00,129
more quantitative basis that were 
previously simply, you know old wives' 

77
00:06:00,129 --> 00:06:05,076
tales, essentially, okay. 
So, I thought that was kind of a fun 

78
00:06:05,076 --> 00:06:10,422
example of, of an, an unusual use of data 
science, again, in a respectable journal. 

79
00:06:10,422 --> 00:06:15,176
All right? 
So another example, just from the Last FM 

80
00:06:15,176 --> 00:06:20,524
blog, they looked at the tags associated 
with the songs in Last F, Last FM. 

81
00:06:20,524 --> 00:06:24,684
And used it to do some simple analysis of 
the emergence of genres over time, but 

82
00:06:24,684 --> 00:06:30,750
you know, based on the popularity of the 
tags for songs coming from that time. 

83
00:06:30,750 --> 00:06:34,890
And so you can see things like well post 
punk in red here came after punk in 

84
00:06:34,890 --> 00:06:40,180
purple, which you'd hope the graph shows 
you'd expect. 

85
00:06:40,180 --> 00:06:45,284
He also uses a kind of rise in here. 
Rock and roll over time, and them maybe a 

86
00:06:45,284 --> 00:06:49,390
bit of a, a, a, a dip more recently. 
And so this is kind of interesting. 

87
00:06:49,390 --> 00:06:53,650
But the, the other theme here we have is, 
that we mentioned in the previous segment 

88
00:06:53,650 --> 00:06:58,324
is repurposing data, right. 
So, this data was collected simply to 

89
00:06:58,324 --> 00:07:02,548
help with search, you know find similar 
music, and it's not being reused to sort 

90
00:07:02,548 --> 00:07:08,165
of draw inferences about the emergence 
of, of entire genres. 

91
00:07:08,165 --> 00:07:12,776
Okay. 
So, another example a few of you may be 

92
00:07:12,776 --> 00:07:16,105
familiar with. 
Google was able to show that, by 

93
00:07:16,105 --> 00:07:21,885
analyzing the search logs, the frequency 
of search terms, it could do a better job 

94
00:07:21,885 --> 00:07:29,410
predicting the severity and the, scope of 
flu outbreaks. 

95
00:07:29,410 --> 00:07:32,512
Than the Centers for Disease Control, and 
by ,better, here we mean essentially 

96
00:07:32,512 --> 00:07:33,850
earlier. 
Right? 

97
00:07:33,850 --> 00:07:44,215
I was able to, give more a head start to 
the, the, health community. 

98
00:07:44,215 --> 00:07:45,810
Okay. 
And also, it was, it was sort of more 

99
00:07:45,810 --> 00:07:47,930
accurate. 
So, how do they do this? 

100
00:07:47,930 --> 00:07:50,855
Well, you know when you're getting the 
flu, it turns out that you want to search 

101
00:07:50,855 --> 00:07:54,460
for flu symptoms, terms associated with 
flu symptoms more often. 

102
00:07:54,460 --> 00:07:59,523
And by watching that uptick, you can 
predict that there's, that the flu 

103
00:07:59,523 --> 00:08:04,160
outbreak is coming. 
Okay. 

104
00:08:04,160 --> 00:08:08,768
So, that's great and they put in a this 
work, and they sort of published a paper 

105
00:08:08,768 --> 00:08:14,510
about it and then they put up this 
interactive visualization. 

106
00:08:14,510 --> 00:08:17,270
A lot of people just sort of analyzed 
[UNKNOWN] going forward. 

107
00:08:17,270 --> 00:08:22,710
But just this year, you know some folks 
showed that he didn't do a very good job 

108
00:08:22,710 --> 00:08:27,842
in this last year. 
So, scientific hindsight shows that the 

109
00:08:27,842 --> 00:08:31,740
Google flu trends far overstated this 
years flu season. 

110
00:08:31,740 --> 00:08:36,150
And the reason they think explains this 
is that there was lots of media attention 

111
00:08:36,150 --> 00:08:40,308
associated with this year's flu season, 
because it was a, was a bit of an uptick, 

112
00:08:40,308 --> 00:08:45,543
and so it got amplified. 
And so this caused people to search more 

113
00:08:45,543 --> 00:08:48,807
flu, flu-related terms more often, 
perhaps because their worried about 

114
00:08:48,807 --> 00:08:53,221
experiencing the symptoms perhaps. 
Because they are trying to understand 

115
00:08:53,221 --> 00:08:58,120
more about the flu outbreak, perhaps 
because they are searching for articles. 

116
00:08:58,120 --> 00:09:00,250
Perhaps because they are worried about 
their kids more. 

117
00:09:00,250 --> 00:09:04,534
But it was a second order effect, based 
on the media attention on the problem 

118
00:09:04,534 --> 00:09:09,530
which lead to skewed results and 
ultimately wrong answer. 

119
00:09:09,530 --> 00:09:12,617
And so the point here is this is, great 
when their re-purposing data from the 

120
00:09:12,617 --> 00:09:16,513
search engine to try and make predictions 
about something else. 

121
00:09:16,513 --> 00:09:19,747
But it's, it's biased its [INAUDIBLE] it 
is biased data so you have to be careful 

122
00:09:19,747 --> 00:09:22,500
with, what you conclude from it. 
Okay? 

123
00:09:22,500 --> 00:09:26,990
So there are limits here. 
All right. 

124
00:09:26,990 --> 00:09:34,420
So, another example also, also, analyzing 
web search traffic. 

125
00:09:34,420 --> 00:09:39,348
That was done with perhaps a little more 
of scientific rigor, was done by some 

126
00:09:39,348 --> 00:09:46,458
folks with the Microsoft research. 
And so, here what you are looking for is 

127
00:09:46,458 --> 00:09:52,690
side effects associated with particular 
drugs. 

128
00:09:52,690 --> 00:09:57,087
And so these results are pretty striking. 
So, what this graph is showing that is 

129
00:09:57,087 --> 00:10:01,911
over time, a set users was around about a 
million that had, that had a permission 

130
00:10:01,911 --> 00:10:07,404
to sort of monitor there there web search 
traffic. 

131
00:10:07,404 --> 00:10:13,514
That when you searched for this drug, in 
green, what percentage of the time did 

132
00:10:13,514 --> 00:10:21,190
you also search for terms associated with 
hyperglycemia symptoms? 

133
00:10:21,190 --> 00:10:26,690
And the answer is somewhere around 5%. 
For this other drug it was somewhere 

134
00:10:26,690 --> 00:10:30,036
around 4%. 
For the, in the background, for the 

135
00:10:30,036 --> 00:10:34,935
average case it was pretty close to 0%. 
If you searched for both of these drugs, 

136
00:10:34,935 --> 00:10:40,655
the odds that you also searched for 
hyperglycemia symptoms went up to 10%. 

137
00:10:40,655 --> 00:10:46,050
Okay, so what's striking about this, is 
that hyperglycemia is not a known side 

138
00:10:46,050 --> 00:10:52,068
effect of these drugs. 
But it seems impossible to ignore from 

139
00:10:52,068 --> 00:10:56,986
the web search data. 
Alright, there's just no reason to 

140
00:10:56,986 --> 00:11:02,250
believe that this could be explained by 
coincidence. 

141
00:11:02,250 --> 00:11:05,117
And they develop this argument more in 
the, in the paper then I, then I just 

142
00:11:05,117 --> 00:11:07,519
have there. 
But, so fine. 

143
00:11:07,519 --> 00:11:11,667
So, re-purposing data, this is another 
example, a large data sets, that had to 

144
00:11:11,667 --> 00:11:16,237
use web search, 
And those are probably the two points I 

145
00:11:16,237 --> 00:11:19,577
want to make about that, but a pretty fun 
one. 

146
00:11:19,577 --> 00:11:25,090
Okay, so the last example I'll give is a 
different take. 

147
00:11:25,090 --> 00:11:30,754
It's more about prediction then data. 
But this from last October, If you recall 

148
00:11:30,754 --> 00:11:35,110
there were six Italian seismologists who 
were convicted of manslaughter, for 

149
00:11:35,110 --> 00:11:40,510
failing to predict a magnitude 6.0 
earthquake in April 2009. 

150
00:11:40,510 --> 00:11:43,434
And while the locals were concerned about 
the seismic activity, the researchers 

151
00:11:43,434 --> 00:11:47,090
were deemed to be just too reassuring 
about the verdict. 

152
00:11:47,090 --> 00:11:49,928
And so the point I want to make here is 
that, there is liability I mean, so this, 

153
00:11:49,928 --> 00:11:54,170
this, this, this scientific community was 
completely aghast that this happened. 

154
00:11:54,170 --> 00:11:57,210
And I'm completely aghast this happened, 
and pretty much everybody is. 

155
00:11:57,210 --> 00:12:01,306
That you can imagine to hold researchers 
responsible for failing to predict 

156
00:12:01,306 --> 00:12:06,950
something that is demonstrably and known 
to be impossible to predict, right? 

157
00:12:06,950 --> 00:12:10,214
So, there's no seismologist on the 
planet, that would argue the earthquakes 

158
00:12:10,214 --> 00:12:15,174
are maybe even remotely predictable. 
and yet the courts sort of decided that 

159
00:12:15,174 --> 00:12:20,570
somehow they, you know, because they got 
the wrong answer, it's bad. 

160
00:12:20,570 --> 00:12:23,898
But it does sort of bring up the issue 
that when you make a prediction, there's 

161
00:12:23,898 --> 00:12:28,020
a certain amount of weight you're, you're 
going to put behind it. 

162
00:12:28,020 --> 00:12:32,439
Whether, whether intentionally or not. 
And so understanding how confident you 

163
00:12:32,439 --> 00:12:36,730
are about that prediction, is sort of an 
important part of the game plan. 

164
00:12:36,730 --> 00:12:40,410
Okay. 
So, that was the last example. 

165
00:12:40,410 --> 00:12:42,100
The, the, the themes that we saw come out 
here. 

166
00:12:42,100 --> 00:12:43,870
We gave a couple examples of graph 
analytics. 

167
00:12:43,870 --> 00:12:47,630
We say that databases were sort of useful 
in the Obama grounding case. 

168
00:12:47,630 --> 00:12:50,070
So, a lot of examples of visualization 
and communicating these results, 

169
00:12:50,070 --> 00:12:53,841
interpreting these results. 
we saw some examples of using very large 

170
00:12:53,841 --> 00:12:57,481
data sets, other examples of using very 
small data sets, and not everything is 

171
00:12:57,481 --> 00:13:03,214
about big data. 
A couple of bullets that aren't on here: 

172
00:13:03,214 --> 00:13:09,855
We talked about, y'know, ad hoc 
interactive analysis. 

173
00:13:09,855 --> 00:13:14,950
It's sort of, not just faster but 
different. 

174
00:13:14,950 --> 00:13:20,982
Also supporting that is important. 
And then we talked about re-purposing 

175
00:13:20,982 --> 00:13:23,116
data. 
Right. 

176
00:13:23,116 --> 00:13:28,296
So, data collected by, perhaps by someone 
else for some other purpose, reusing that 

177
00:13:28,296 --> 00:13:34,375
as a drawing for something else, that's a 
pretty common theme here. 

178
00:13:34,375 --> 00:13:38,096
Okay. 
In the next couple of segments, we'll 

179
00:13:38,096 --> 00:13:41,270
talk about how we organized this course 
and some of the design decisions we made 

180
00:13:41,270 --> 00:13:43,740
in, in creating the material. 

