1
00:00:05,910 --> 00:00:07,920
[MUSIC] Welcome to Introduction to Data 
Science. 

2
00:00:07,920 --> 00:00:10,693
My name is Bill Howe, and I'm the 
director of research for scalable data 

3
00:00:10,693 --> 00:00:14,310
analytics at the University of Washington 
eScience institute. 

4
00:00:14,310 --> 00:00:17,208
And an affiliate assistant professor in 
computer science and enginerreing, also 

5
00:00:17,208 --> 00:00:20,662
at the University of Washington. 
So in this first segment, what I want to 

6
00:00:20,662 --> 00:00:23,990
do is go through some examples of data 
science activities and projects in the 

7
00:00:23,990 --> 00:00:28,462
recent past that I found interesting. 
And use them to sort of whet your 

8
00:00:28,462 --> 00:00:32,774
appetites for the concepts that we are 
going to learn in this course okay. 

9
00:00:32,774 --> 00:00:38,530
So, the first one I want to mention here 
is the presidential election from 2012. 

10
00:00:38,530 --> 00:00:41,568
And I know you're probably sick of 
hearing about this if you're, if you live 

11
00:00:41,568 --> 00:00:44,915
in the United States. 
And even if you don't you may be sick of 

12
00:00:44,915 --> 00:00:48,763
hearing about it, but bear with me. 
So, this is a map of the electoral 

13
00:00:48,763 --> 00:00:53,011
college, and each state is colored by the 
for the candidate that took the electoral 

14
00:00:53,011 --> 00:00:56,568
votes. 
And the numbers represent how many 

15
00:00:56,568 --> 00:01:00,872
electoral votes each state has. 
And so if you recall, what was 

16
00:01:00,872 --> 00:01:05,632
interesting about this map at the time 
was that it led to a pretty significant 

17
00:01:05,632 --> 00:01:10,294
discussion in the media about Data 
Science. 

18
00:01:10,294 --> 00:01:14,716
Because Nate Silver of the 538 blog was 
able to predict this map perfectly before 

19
00:01:14,716 --> 00:01:18,884
the election. 
Alright, and you know that discussion in 

20
00:01:18,884 --> 00:01:24,924
the media, had, talked a lot about you 
know, what a genius Nate Silver was. 

21
00:01:24,924 --> 00:01:28,877
And mentioned the sophisticated 
mathematics he was using, and how you 

22
00:01:28,877 --> 00:01:32,540
know, he's sort of a wiz with these 
things. 

23
00:01:32,540 --> 00:01:35,642
But, what I thought was interesting about 
this was that Nate Silver would be the 

24
00:01:35,642 --> 00:01:39,107
first one to tell you that. 
The methods she was employing to make 

25
00:01:39,107 --> 00:01:41,530
this prediction were actually pretty 
simple. 

26
00:01:41,530 --> 00:01:43,100
Right? 
And so he says here in a series of quotes 

27
00:01:43,100 --> 00:01:46,983
from blog posts around that time. 
This first one from October 26 was. 

28
00:01:46,983 --> 00:01:49,775
The intuition behind this ought to be 
very simple. 

29
00:01:49,775 --> 00:01:53,185
Mr Obama is maintaining leads in the 
polls in Ohio and other states that are 

30
00:01:53,185 --> 00:01:56,695
sufficient for him to win 270 electoral 
votes. 

31
00:01:56,695 --> 00:02:01,280
And what was funny was a few days later 
he becomes, sort of, more blunt. 

32
00:02:01,280 --> 00:02:03,540
The argument we're making is exceedingly 
simple. 

33
00:02:03,540 --> 00:02:07,000
Here it is, Obama is ahead in Ohio, 
right? 

34
00:02:07,000 --> 00:02:10,253
It's not a magic trick. 
And so then after the election on 

35
00:02:10,253 --> 00:02:13,242
November 10th, and he was shown to be 
right and got this, sort of, flawless 

36
00:02:13,242 --> 00:02:18,050
prediction. 
The blog post that this last quote is 

37
00:02:18,050 --> 00:02:19,790
taken from. 
Whoops, excuse me. 

38
00:02:19,790 --> 00:02:23,625
That the last quote was taken from was 
describing why he started the 538 blog in 

39
00:02:23,625 --> 00:02:26,872
the first place. 
And he says, look, you know, the bar set 

40
00:02:26,872 --> 00:02:31,078
by the competition was invitingly low. 
Someone could look like a genius simply 

41
00:02:31,078 --> 00:02:34,662
by doing some fairly basic research into 
what really has predicted power in a 

42
00:02:34,662 --> 00:02:38,777
political campaign. 
And so what really had predictive power 

43
00:02:38,777 --> 00:02:42,830
in this case was the state polls 
themselves aggregated right. 

44
00:02:42,830 --> 00:02:46,910
So, historically the state polls 
aggregated did a pretty good job of 

45
00:02:46,910 --> 00:02:52,710
predicting the outcome of the general 
election and so that's was he did. 

46
00:02:52,710 --> 00:02:56,870
Now there wasn't In quantifying the 
uncertainty and certainly in presenting 

47
00:02:56,870 --> 00:02:59,988
these results. 
There's a lot of beautiful interactive 

48
00:02:59,988 --> 00:03:02,886
visualizations that he created in order 
to sort of convey these ideas to the 

49
00:03:02,886 --> 00:03:05,523
public. 
And that's one of the points I want to 

50
00:03:05,523 --> 00:03:09,350
make about this is that getting the 
answer in some cases is the easy part. 

51
00:03:09,350 --> 00:03:14,570
It's in interpreting the results and in 
convincing others of the result. 

52
00:03:14,570 --> 00:03:18,800
By presenting them usually through 
visualization, that can be the hard part. 

53
00:03:18,800 --> 00:03:24,790
This is one of the themes that we'll come 
back to throughout this course, okay. 

54
00:03:24,790 --> 00:03:28,540
so, so just to summarize it though, I'm 
not sure I said this. 

55
00:03:28,540 --> 00:03:33,956
The simple methods plus enough good data. 
You know wins, that trumps more 

56
00:03:33,956 --> 00:03:38,650
sophisticated methods in many cases. 
That's another theme we'll come back to. 

57
00:03:38,650 --> 00:03:41,450
Alright. 
So something else related to the campaign 

58
00:03:41,450 --> 00:03:46,138
before we move on to other topics. 
Was the system that the Obama campaign 

59
00:03:46,138 --> 00:03:49,870
used for their data driven ground game to 
so speak. 

60
00:03:49,870 --> 00:03:54,556
The Build you sort of target, direct, you 
know, targets to specific categories of 

61
00:03:54,556 --> 00:03:57,718
users. 
And so what they did was they built and 

62
00:03:57,718 --> 00:04:01,806
maintained a really significantly sized 
massive voter database. 

63
00:04:01,806 --> 00:04:05,334
And used it to design these highly 
tailored messages to very, very specific 

64
00:04:05,334 --> 00:04:08,612
groups. 
So, you know the mother of two in a small 

65
00:04:08,612 --> 00:04:12,440
town in Ohio who tweeted about the 
environment and mentioned organic 

66
00:04:12,440 --> 00:04:18,142
vegetables on her Facebook page. 
And you know, who had voted in 2008, and 

67
00:04:18,142 --> 00:04:22,520
had registered on Obama's website but had 
never donated. 

68
00:04:22,520 --> 00:04:26,720
Okay, you know, she would get a message 
from Michelle Obama that highlighted 

69
00:04:26,720 --> 00:04:31,770
Barack Obama's environmental policies. 
Okay. 

70
00:04:31,770 --> 00:04:34,637
And so in order to design these message, 
what you had to do was, kind of do add 

71
00:04:34,637 --> 00:04:38,410
hawk hypothesis testing about what might 
work and what didn't. 

72
00:04:38,410 --> 00:04:42,387
You kind of had to slice and dice this 
da, data at, kind of interactive speeds. 

73
00:04:42,387 --> 00:04:45,681
And this is another theme that we'll 
return to is the need for these kind of 

74
00:04:45,681 --> 00:04:49,290
adhoc interactive analysis. 
And the systems they use for this are 

75
00:04:49,290 --> 00:04:51,791
pretty interesting too, you know, this 
was a SQL database, a very fast one 

76
00:04:51,791 --> 00:04:54,835
called Vertica. 
And we'll talk a little bit about what 

77
00:04:54,835 --> 00:04:58,099
makes Vertica special, I hope, toward the 
end of the course. 

78
00:04:58,099 --> 00:05:00,979
but it is a SQL database and so SQL 
sometimes gets a bad name in data science 

79
00:05:00,979 --> 00:05:04,144
context. 
And sort of the old guard, that can't be 

80
00:05:04,144 --> 00:05:07,246
possibly used for analytics and doesn't 
really make sense in today's era, but, 

81
00:05:07,246 --> 00:05:09,880
you know, don't believe it. 
Right? 

82
00:05:09,880 --> 00:05:14,688
It has a role to play in many cases. 
And so here, you know, they did use 

83
00:05:14,688 --> 00:05:18,978
Hadoop, right, to do the aggregate 
generations of anything not real-time, he 

84
00:05:18,978 --> 00:05:22,900
says here. 
But for the speed-of-thought queries 

85
00:05:22,900 --> 00:05:26,025
about the data they used, this Vertica 
database. 

86
00:05:26,025 --> 00:05:33,674
Okay, and so we'll come back to systems, 
in, in a, several segments from now. 

87
00:05:33,674 --> 00:05:36,510
Okay. 
So moving on this was around the same 

88
00:05:36,510 --> 00:05:41,066
time, you know when Hurricane Sandy made 
landfall, one of the things that struck 

89
00:05:41,066 --> 00:05:48,148
me was. 
The fact that visualizations of available 

90
00:05:48,148 --> 00:05:53,460
data were starting to emerge in real time 
in response to the storm. 

91
00:05:53,460 --> 00:05:57,552
And there's some very nice examples of 
people who have used Twitter data to 

92
00:05:57,552 --> 00:06:02,580
analyze or to produce a map of where the 
power was going out. 

93
00:06:02,580 --> 00:06:07,260
And in this even simpler case Josef 
Fruehwald got public data from local 

94
00:06:07,260 --> 00:06:11,950
weather stations. 
And just from two different weather 

95
00:06:11,950 --> 00:06:18,174
stations and just simply plotted them. 
And so this is the barometric pressure 

96
00:06:18,174 --> 00:06:23,142
over the course of this, you know, few 
day period. 

97
00:06:23,142 --> 00:06:29,370
Two days I guess in two locations, 
Atlantic City and Philadelphia. 

98
00:06:29,370 --> 00:06:32,740
And so you can see this enormous dip is 
the storm passing through. 

99
00:06:32,740 --> 00:06:37,370
You can also see the time lag between 
Atlantic City and Philadelphia. 

100
00:06:37,370 --> 00:06:41,087
And you can also see the intensity is 
probably a little bit higher given that 

101
00:06:41,087 --> 00:06:45,352
the barometric drop is more significant 
in Atlantic City. 

102
00:06:45,352 --> 00:06:48,590
Okay. 
So couple things here. 

103
00:06:48,590 --> 00:06:53,067
One is pulling data down from the web and 
re purposing it sort of in real time. 

104
00:06:53,067 --> 00:06:56,763
Or at least in, in short time, not real 
time, to produce visualizations and then 

105
00:06:56,763 --> 00:07:01,356
publishing those back onto the web. 
I think this is very much the characters 

106
00:07:01,356 --> 00:07:04,584
of data science activities. 
In this particular example, there's not 

107
00:07:04,584 --> 00:07:09,070
necessarily a large data set involved. 
But re-purposing data that was collected 

108
00:07:09,070 --> 00:07:11,922
for a different purpose is, is a theme 
that we'll come back to, and again, we 

109
00:07:11,922 --> 00:07:17,976
see the ad hoc nature of this as well. 
Fine, so another plot here is wind speeds 

110
00:07:17,976 --> 00:07:22,276
and they sort of peak out at 40 at 
Atlantic City which is green. 

111
00:07:22,276 --> 00:07:25,872
And you can see that Atlantic City is 
indeed more intense here, and the gray 

112
00:07:25,872 --> 00:07:31,800
here is arrow bars. 
so again, another variance on the same 

113
00:07:31,800 --> 00:07:35,276
data. 
Okay, so changing gears a little bit, 

114
00:07:35,276 --> 00:07:40,556
this was a study, the title here is 
called The Expression of Emotions in 20th 

115
00:07:40,556 --> 00:07:45,615
Century Books. 
And so what they were interested in is 

116
00:07:45,615 --> 00:07:49,775
had the words that we choose to use in 
our collective literature changed over 

117
00:07:49,775 --> 00:07:53,294
time. 
And eventually does that tell us 

118
00:07:53,294 --> 00:07:56,940
something about, sort of, culture or 
civilization? 

119
00:07:56,940 --> 00:08:00,478
I find the scientific inquiry sort of 
compelling, but what I think is most 

120
00:08:00,478 --> 00:08:04,429
striking about this. 
And why I wanted to include this example, 

121
00:08:04,429 --> 00:08:09,020
is that the methodology that they used 
here is pretty straight forward. 

122
00:08:09,020 --> 00:08:13,150
You could do this yourself with not a 
significant background in either 

123
00:08:13,150 --> 00:08:18,832
technology or in statistics. 
Or even if in you know, linguistics, or 

124
00:08:18,832 --> 00:08:22,330
anything. 
And so this is what they did, right? 

125
00:08:22,330 --> 00:08:26,537
So the first step is kind of a doozy. 
This is take all the books written in the 

126
00:08:26,537 --> 00:08:31,362
20th century and digitize them. 
Well that would be a non starter, none of 

127
00:08:31,362 --> 00:08:34,964
us could do that but that's okay. 
Google's already done it for us and 

128
00:08:34,964 --> 00:08:39,500
they've made the data available at this 
URL and so you can go check that out. 

129
00:08:39,500 --> 00:08:43,376
What they've done is digitize the books, 
done the character recognition on it and 

130
00:08:43,376 --> 00:08:49,091
produced these Ingram data sets. 
So these are tables of data where each 

131
00:08:49,091 --> 00:08:54,283
row has an ingram and followed by the 
year and the counts of the number of 

132
00:08:54,283 --> 00:09:01,007
times that the ingram occurred. 
This has already been broken down and 

133
00:09:01,007 --> 00:09:04,340
processed into a form that's digestible, 
okay. 

134
00:09:04,340 --> 00:09:06,600
And so what's an ingram. 
Well, it's pretty simple. 

135
00:09:06,600 --> 00:09:10,440
A one gram is just a single word like 
yesterday. 

136
00:09:10,440 --> 00:09:15,372
A five gram, an example here is, the 
phrase, analysis is often described as. 

137
00:09:15,372 --> 00:09:20,350
Okay, and so in this study, they just 
ignored everything but the one grams. 

138
00:09:20,350 --> 00:09:24,838
and then they took some subset of, some 
subset of those one grams, and assigned 

139
00:09:24,838 --> 00:09:28,640
them a mood score. 
So, how did they do this? 

140
00:09:28,640 --> 00:09:33,260
Well you can imagine that certain words 
are, are charged with a particular mood 

141
00:09:33,260 --> 00:09:37,956
or associated with joy or sadness or fear 
and so on. 

142
00:09:37,956 --> 00:09:42,786
And you can also imagine that synonyms of 
those words might also be associated with 

143
00:09:42,786 --> 00:09:46,375
that mood. 
And so this analysis sounds non-trivial, 

144
00:09:46,375 --> 00:09:49,770
and it is, but once again, that's already 
been done for you. 

145
00:09:49,770 --> 00:09:54,192
There's a, a resource on the web called 
WordNet, where they've done this kind of 

146
00:09:54,192 --> 00:09:58,200
effect analysis. 
And so the authors of this paper were 

147
00:09:58,200 --> 00:10:01,601
able to take the digitized books from 
Google. 

148
00:10:01,601 --> 00:10:06,445
or already broken down into Ingram, and 
the [INAUDIBLE] scores from WordNet. 

149
00:10:07,600 --> 00:10:10,360
And then do this calculation, which, you 
know, may sort of looking intimidating if 

150
00:10:10,360 --> 00:10:13,466
you're not used to staring at these 
mathematical expressions. 

151
00:10:13,466 --> 00:10:18,165
But it's actually pretty simple. 
This is the count of a particular, word 

152
00:10:18,165 --> 00:10:27,440
in the set, the set being the set of word 
net word, which is not as all the words. 

153
00:10:27,440 --> 00:10:30,540
Only some words are able to be scored as 
mood. 

154
00:10:30,540 --> 00:10:35,190
And then you normalize by the count of 
the number of occurrences of the word 

155
00:10:35,190 --> 00:10:38,818
the. 
So why do they do that? 

156
00:10:38,818 --> 00:10:43,918
Well you need to normalize over something 
in order to account for the fact that 

157
00:10:43,918 --> 00:10:49,810
perhaps we just write more books in 2005 
than we did in 1937. 

158
00:10:49,810 --> 00:10:51,790
Or we've been able to digitize more 
books. 

159
00:10:51,790 --> 00:10:55,139
So we need to normalize by that total. 
Well, why don't I just normalize by the 

160
00:10:55,139 --> 00:10:58,678
total number of words? 
Well, the reason is that the, the word 

161
00:10:58,678 --> 00:11:03,288
the is a better indicator of prose than 
the total number of, of words. 

162
00:11:03,288 --> 00:11:07,000
And this is because we've also, 
apparently we've also started to produce 

163
00:11:07,000 --> 00:11:12,360
more sort of captions in figures, and 
more sort of technical language. 

164
00:11:12,360 --> 00:11:17,535
And more sort of formula and more 
expressions, more non-prose, utterances 

165
00:11:17,535 --> 00:11:21,910
in these, in these, books. 
And therefore we can sort of skew the 

166
00:11:21,910 --> 00:11:23,752
results. 
We really want to capture, in our 

167
00:11:23,752 --> 00:11:26,834
language, when we actually write full 
complete sentences, however these words 

168
00:11:26,834 --> 00:11:28,860
being used. 
Okay. 

169
00:11:28,860 --> 00:11:35,555
And then you add those up and you divide 
by the total number of words in the set. 

170
00:11:35,555 --> 00:11:40,859
And then there's one more transformation 
here that should look familiar to you if 

171
00:11:40,859 --> 00:11:45,715
you sort of recall your high school 
statistics. 

172
00:11:45,715 --> 00:11:51,175
And so you subtract the mean and divide 
the standard deviation. 

173
00:11:51,175 --> 00:11:55,291
Okay. 
So this is normalizing with respect to a 

174
00:11:55,291 --> 00:11:58,610
normal distribution. 
Okay. 

175
00:11:58,610 --> 00:12:00,975
But that's about it. 
You know, there's a count, and there's a 

176
00:12:00,975 --> 00:12:04,140
division. 
and then there's two data sets that you 

177
00:12:04,140 --> 00:12:06,598
can pull from the web. 
And they're big but they're not 

178
00:12:06,598 --> 00:12:08,900
exceedingly big. 
They fit in memory on most of your 

179
00:12:08,900 --> 00:12:11,208
laptops nowadays. 
So it's a, it's a significant 

180
00:12:11,208 --> 00:12:13,910
computational task, but nothing that 
requires Hadoop. 

181
00:12:13,910 --> 00:12:19,124
You can do this in a weekend if you, had 
thought about it. 

182
00:12:19,124 --> 00:12:21,962
Okay, so I find that pretty compelling. 
Fine. 

183
00:12:21,962 --> 00:12:28,662
So, these are the results, this is joy 
words minus sadness words so this is the 

184
00:12:28,662 --> 00:12:36,850
z-score for joy and sadness. 
And you can see that there's sort of a 

185
00:12:36,850 --> 00:12:43,116
big dip after World War two. 
And that's one of the points they make in 

186
00:12:43,116 --> 00:12:48,896
the, in the paper. 
And then you can see this sort of thing 

187
00:12:48,896 --> 00:12:56,194
start to increase in the late 90s, okay. 
I won't try to analyze this for the 

188
00:12:56,194 --> 00:13:00,590
scientific value, I'll just present the 
results. 

189
00:13:00,590 --> 00:13:04,930
What I think is maybe more interesting is 
this one. 

190
00:13:04,930 --> 00:13:11,161
So this is now emotion words total, minus 
random words total and there's a sort of 

191
00:13:11,161 --> 00:13:17,070
prominent downward slope over time. 
So what is this, what is this, what's 

192
00:13:17,070 --> 00:13:20,624
going on here. 
Well, apparently you can make the 

193
00:13:20,624 --> 00:13:26,805
argument that we're using fewer emotion 
words over time. 

194
00:13:26,805 --> 00:13:30,776
Okay. 
That said, there's a bit of an uptick in 

195
00:13:30,776 --> 00:13:32,910
this red line. 
And so what does that represent? 

196
00:13:32,910 --> 00:13:37,377
Well, that's fear words. 
And, y'know, you can imagine some of the 

197
00:13:37,377 --> 00:13:43,841
reasons why there might be an increase in 
fear words, since the, since the 1980s. 

198
00:13:43,841 --> 00:13:46,780
Okay. 
So this is pretty fun though this is a 

199
00:13:46,780 --> 00:13:49,111
significant analysis that can be done 
just by taking these, these data sets 

200
00:13:49,111 --> 00:13:51,582
that they didn't have to prepare 
themselves. 

201
00:13:51,582 --> 00:13:54,280
All right. 
And then the other point I want to make 

202
00:13:54,280 --> 00:13:56,532
about this. 
This is just a copy and paste of a 

203
00:13:56,532 --> 00:13:59,640
segment of the papers that this paper 
cites. 

204
00:13:59,640 --> 00:14:01,900
And I just was struck by the titles here, 
you know. 

205
00:14:01,900 --> 00:14:05,035
Quantitative Analysis of Culture Using 
Millions of Digitized Books. 

206
00:14:05,035 --> 00:14:08,710
Quantifying the Evolutionary Dynamics of 
Language. 

207
00:14:08,710 --> 00:14:11,974
Frequency of word use predicts rates of 
lexical evolution through Indo European 

208
00:14:11,974 --> 00:14:17,970
history. 
song lyrics, I mean linguistic markers. 

209
00:14:17,970 --> 00:14:22,731
I mean what strikes me about this is that 
you know linguistics anthropology 

210
00:14:22,731 --> 00:14:27,428
history, culture. 
These studies are becoming hard sciences 

211
00:14:27,428 --> 00:14:30,940
by the virtue of data driven methods, 
right. 

212
00:14:30,940 --> 00:14:35,380
So all science is becoming data science. 
Right? 

213
00:14:35,380 --> 00:14:40,116
And therefore data scientists have a lot 
of power in this regime, it's a great 

214
00:14:40,116 --> 00:14:46,557
time to be you know, a data geek okay. 
You know, there's being a journalism as 

215
00:14:46,557 --> 00:14:48,494
well. 
I mean one point, I promised you for the 

216
00:14:48,494 --> 00:14:51,810
slide in here about this. 
But you know when the Wiki leaks material 

217
00:14:51,810 --> 00:14:55,775
came out, you know, what, you're not 
going to pour yourself a pot of coffee 

218
00:14:55,775 --> 00:15:00,781
and pore over these materials. 
You know, print them all out and sort of 

219
00:15:00,781 --> 00:15:03,495
go through them, one by one, you're 
going to write algorithms that do this 

220
00:15:03,495 --> 00:15:06,642
kind of an analysis. 
You know, word-use analysis, look for 

221
00:15:06,642 --> 00:15:09,266
email chains and dialogues, these sort of 
computational methods, in order to 

222
00:15:09,266 --> 00:15:13,067
analyze that material. 
That's on journalism itself is a 

223
00:15:13,067 --> 00:15:18,332
computational enterprise, is a data 
science problem, not a or at least 

224
00:15:18,332 --> 00:15:24,127
amenable to data science technique. 
So, you know, as a data scientist the 

225
00:15:24,127 --> 00:15:27,102
world is your oyster. 
Alright. 

226
00:15:27,102 --> 00:15:30,987
So, let me pause there. 
And we'll pick up with a couple more 

227
00:15:30,987 --> 00:15:34,120
examples before moving on in the next 
segment. 

