1
00:00:05,798 --> 00:00:06,547
[MUSIC]. 
Okay. 

2
00:00:06,547 --> 00:00:13,085
So, what's a second reason for this 
decline effect for the truth wearing off? 

3
00:00:13,085 --> 00:00:18,267
Another one is just people make mistakes 
and in some cases there's is This is a 

4
00:00:18,267 --> 00:00:20,420
fraud. 
And so there's some evidence here that 

5
00:00:20,420 --> 00:00:22,300
these kinds of things are going up as 
well. 

6
00:00:22,300 --> 00:00:28,610
as measure, one way to measure this is by 
the number of paper retractions, and the 

7
00:00:28,610 --> 00:00:32,905
number of retractions has gone up fairly 
significantly as Norton pointed out in 

8
00:00:32,905 --> 00:00:37,550
2011 in an article in Nature. 
So, from the period of 2001 to 2011 there 

9
00:00:37,550 --> 00:00:44,260
has been a tenfold increase in number of 
paper retractions, but only in 1.44 fold 

10
00:00:44,260 --> 00:00:49,580
increase in papers. 
Themselves. 

11
00:00:49,580 --> 00:00:53,580
So this plot is from two different 
repositories, PubMed and Web Of Science, 

12
00:00:53,580 --> 00:00:55,960
neither one of which can include Computer 
Science papers by the way. 

13
00:00:55,960 --> 00:01:03,202
But you can see that it's gone up pretty 
significantly. 

14
00:01:03,202 --> 00:01:05,740
Fine. 
So, not going to dive into too much 

15
00:01:05,740 --> 00:01:09,350
detail into this Phenomenon. 
But I do want to give you one statistical 

16
00:01:09,350 --> 00:01:14,799
tool that is sometimes used to detect 
fraud. 

17
00:01:14,799 --> 00:01:21,946
And that's Benford's Law. 
So, if you're not familiar with Benford's 

18
00:01:21,946 --> 00:01:26,180
Law, it's a fun one to be familiar with. 
So this. 

19
00:01:26,180 --> 00:01:32,630
Law predicts the distribution of the 
first digit of data. 

20
00:01:32,630 --> 00:01:37,568
Okay. 
And if you think about it without think- 

21
00:01:37,568 --> 00:01:45,870
-- if you think about it without thinking 
too deeply, you might intuitively think 

22
00:01:45,870 --> 00:01:50,060
that this would be fairly uniform. 
If you have random data, you might see an 

23
00:01:50,060 --> 00:01:55,910
equal number of 8s and an equal number of 
3s and an equal number of 2s and so on. 

24
00:01:55,910 --> 00:02:01,200
But it turns out that the distribution is 
not random un- -- in, in some 

25
00:02:01,200 --> 00:02:05,360
circumstances, in many circumstances. 
Well, it's not even, I should say. 

26
00:02:05,360 --> 00:02:06,920
It is random. 
It's not even. 

27
00:02:06,920 --> 00:02:10,980
The distribution looks more like this. 
You will get more 1's in the first 

28
00:02:10,980 --> 00:02:14,070
position than you will 2's and more 2's 
than 3's and so on. 

29
00:02:14,070 --> 00:02:17,672
Until up to 30 percent of the numbers 
you're measuring are 1's. 

30
00:02:17,672 --> 00:02:23,430
So this should, you know, blow your mind 
a little bit, okay. 

31
00:02:23,430 --> 00:02:26,210
So some examples. 
Of this before we explain what's going 

32
00:02:26,210 --> 00:02:28,875
on. 
taken from there's a nice website 

33
00:02:28,875 --> 00:02:32,990
testingbenfordslaw.com where they pulled 
real data in, and, and showed the plots. 

34
00:02:32,990 --> 00:02:37,460
This is the number of Twitter users by 
their follow, by the number of followers 

35
00:02:37,460 --> 00:02:39,080
they have. 
Sorry not the number of Twitter users, 

36
00:02:39,080 --> 00:02:40,917
Twitter users by the number of followers 
they have. 

37
00:02:40,917 --> 00:02:47,070
Have, okay? 
Alright, so the, the, the list of numbers 

38
00:02:47,070 --> 00:02:52,140
is just the number of followers. 
The number of Twitter users that have, 

39
00:02:52,140 --> 00:02:58,860
whose number of followers begins with the 
digit 1 is 32.62%, if you can see that. 

40
00:02:58,860 --> 00:03:03,370
And the red dot is the prediction made by 
Benford's Law. 

41
00:03:03,370 --> 00:03:08,430
so not too bad. 
The number of instances that have the two 

42
00:03:08,430 --> 00:03:14,576
as the leading digit is 16.66% and 
there's the red dot predicted by 

43
00:03:14,576 --> 00:03:17,460
Benford's Law. 
So not too bad. 

44
00:03:18,750 --> 00:03:24,470
The distance of stars from earth in light 
years Follows a similar pattern. 

45
00:03:24,470 --> 00:03:27,440
This is remarkably close to Benford's Law 
prediction. 

46
00:03:27,440 --> 00:03:35,155
Government spending in the UK between the 
period of time May through September of 

47
00:03:35,155 --> 00:03:39,480
2010. 
Now, you might imagine there's some 

48
00:03:39,480 --> 00:03:43,350
selection bias on this particular 
website, and I can't guarantee that 

49
00:03:43,350 --> 00:03:48,520
there's not. 
But, you know, with some reading plus a 

50
00:03:48,520 --> 00:03:51,360
little bit of trust that I've hopefully 
built up. 

51
00:03:51,360 --> 00:03:58,250
I hope to convince you that this is not 
fully explainable by this website, 

52
00:03:58,250 --> 00:04:00,815
choosing particular data sets for which 
this is true. 

53
00:04:00,815 --> 00:04:03,796
Okay. 
Google Books, the number unique 1-grams. 

54
00:04:03,796 --> 00:04:08,785
And we talked about one grams several, a 
couple weeks ago. 

55
00:04:08,785 --> 00:04:14,790
Terms, essentially, are what grams are. 
Okay, again, pretty good prediction. 

56
00:04:14,790 --> 00:04:19,880
So, before we say, before we give the 
intuition for this. 

57
00:04:19,880 --> 00:04:22,720
If there, intuition's a little bit 
tricky, but if we attempt to give the 

58
00:04:22,720 --> 00:04:26,300
intuition to this. 
You can use it to detect fraud. 

59
00:04:26,300 --> 00:04:30,560
And so this was attempted, or this was an 
experiment was done by Diekmann in 2007 

60
00:04:30,560 --> 00:04:35,990
to see if this could be used for 
detecting scientific fraud. 

61
00:04:35,990 --> 00:04:39,900
And so what he found was that first and 
second digits of published statistical 

62
00:04:39,900 --> 00:04:44,830
estimates were approximately Benford 
distributed in real studies, okay, or at 

63
00:04:44,830 --> 00:04:48,710
the very least they had kind of a 
monotonically decreasing distribution. 

64
00:04:48,710 --> 00:04:51,770
So ones, more, more ones than twos. 
And more twos and threes and so on. 

65
00:04:51,770 --> 00:04:55,390
If not exactly Benford's Law. 
And then what he did was asked subjects 

66
00:04:55,390 --> 00:05:00,290
to manufacture regression coefficients. 
You know? 

67
00:05:00,290 --> 00:05:03,450
Basically fitting a line manually. 
And found that the first digits were 

68
00:05:03,450 --> 00:05:06,820
actually hard to detect as anomalous. 
But that the second and third digits 

69
00:05:06,820 --> 00:05:10,590
deviated from expected distributions. 
And so, this distribution that I gave you 

70
00:05:10,590 --> 00:05:15,080
in, in the last few slides. 
And, and that I gave you in the, Actually 

71
00:05:15,080 --> 00:05:17,921
I guess I didn't give you the actual 
formula here. 

72
00:05:17,921 --> 00:05:22,510
Okay so the distribution that we'll be 
discussing is only for the first digit 

73
00:05:22,510 --> 00:05:26,660
and the skew in that distribution 
actually gets suppressed as you go the 

74
00:05:26,660 --> 00:05:29,174
second and third digits. 
But Benford's Law can also be used to 

75
00:05:29,174 --> 00:05:36,170
express different, yet still measurable 
distributing of second and third digit. 

76
00:05:36,170 --> 00:05:41,120
And so, there, the second and third 
digits deviated significantly. 

77
00:05:41,120 --> 00:05:45,240
And so the conclusion was, this is, it is 
potential, potentially useful as a fraud 

78
00:05:45,240 --> 00:05:48,415
detection tool in scientific data. 
And actually, there are instances where 

79
00:05:48,415 --> 00:05:52,610
Benford's law has been admissible as 
evidence in court in cases of fraud, and 

80
00:05:52,610 --> 00:05:59,730
it's been used by reporters and so forth 
to argue for evidence of fraud in cases 

81
00:05:59,730 --> 00:06:05,030
of, voting, election, and other kinds of 
Accounting data on the sort of global 

82
00:06:05,030 --> 00:06:07,760
scene, okay. 
So what's going on here? 

83
00:06:07,760 --> 00:06:12,635
Well, one way to think about the 
intuition here is, imagine a sequence of 

84
00:06:12,635 --> 00:06:17,060
cards labeled with a particular number. 
1, 2, 3, 4, 5 all the way up to, you 

85
00:06:17,060 --> 00:06:22,998
know, 999, 999 fine, okay. 
And so put them one by one in a hat, in 

86
00:06:22,998 --> 00:06:27,282
order. 
And at each, every time you throw a card 

87
00:06:27,282 --> 00:06:35,136
in, measure the probability that a random 
selection from the hat would produce a 

88
00:06:35,136 --> 00:06:40,000
card where the first digit is one. 
Sorry, this isn't very well said. 

89
00:06:40,000 --> 00:06:50,036
I, it's not. 
It's not drawing the number 1, drawing, a 

90
00:06:50,036 --> 00:07:08,910
card, where. 
The first digit is 1. 

91
00:07:08,910 --> 00:07:10,790
Okay. 
So what does that probability look like? 

92
00:07:10,790 --> 00:07:14,740
Well, this figures a little bit 
misleading because the X-axis is on the 

93
00:07:14,740 --> 00:07:18,960
long scale, but if you just look at the 
heights. 

94
00:07:18,960 --> 00:07:23,345
What's going on here is on the y axis is 
the probability of drawing a, a 

95
00:07:23,345 --> 00:07:26,860
particular digit. 
The blue line is associated with the 

96
00:07:26,860 --> 00:07:30,410
digit one, the green line is a digit two 
and so on. 

97
00:07:30,410 --> 00:07:35,084
And this was generated by a simulation, 
where you, you know. 

98
00:07:35,084 --> 00:07:39,049
I really did select random numbers from a 
distribution. 

99
00:07:39,049 --> 00:07:44,130
And I really did order them in the manner 
described in the previous slide and put 1 

100
00:07:44,130 --> 00:07:50,270
in and measured the probability drawing 
it value of number 1 or with the vertices 

101
00:07:50,270 --> 00:07:55,196
of just 1, okay. 
And so if you think about this, the 

102
00:07:55,196 --> 00:07:59,370
numbers You know, what's, what's happened 
here, this is where 1 and this is 10 and 

103
00:07:59,370 --> 00:08:04,380
this is 100 and so on. 
Well, as soon as I put a 1 in, the chance 

104
00:08:04,380 --> 00:08:07,300
of drawing a card with a 1 on the front 
is 100%. 

105
00:08:07,300 --> 00:08:11,348
When I put a 2 in, it's now 50%. 
When I put a 3 in it's 33% and so on. 

106
00:08:11,348 --> 00:08:17,600
But then as soon as I get to 10 It goes 
back up to a higher percentage again and 

107
00:08:17,600 --> 00:08:20,620
it stays there, 10, 11, 12, 13, 14, and 
so on. 

108
00:08:20,620 --> 00:08:26,240
And then as soon as I get to 20, it drops 
down a little bit. 

109
00:08:26,240 --> 00:08:29,740
And so, that's why it's climbing here, 
through the tens, and then it starts 

110
00:08:29,740 --> 00:08:33,400
dropping again. 
Again sharply, okay. 

111
00:08:33,400 --> 00:08:37,170
And so the point here is that under this 
model, values where the first digit is 

112
00:08:37,170 --> 00:08:40,870
one always are in the hat already by the 
time you get to the twos. 

113
00:08:42,610 --> 00:08:44,420
And the twos are always there before you 
get to the threes. 

114
00:08:44,420 --> 00:08:46,325
And the threes are always there The fours 
and so on. 

115
00:08:46,325 --> 00:08:51,390
And so you end up with these height, 
these peaks are lead by the ones. 

116
00:08:51,390 --> 00:08:53,880
And the area under this curve. 
Although, remember this is a log plot so 

117
00:08:53,880 --> 00:08:58,450
the area's not quite right. 
represents the probability of, of drawing 

118
00:08:58,450 --> 00:08:59,990
this. 
And so the probability does actually get 

119
00:08:59,990 --> 00:09:02,760
higher. 
Okay. 

120
00:09:02,760 --> 00:09:05,980
And there's a few different other models 
you can, you can find if you read up on 

121
00:09:05,980 --> 00:09:08,220
this. 
And I encourage you to there's other ways 

122
00:09:08,220 --> 00:09:10,260
to sort of think about the probability 
here. 

123
00:09:10,260 --> 00:09:13,840
And there is actually a closed form 
expression for Benford's Law as well. 

124
00:09:13,840 --> 00:09:19,920
Okay. 
One of the limitations here, it's not 

125
00:09:19,920 --> 00:09:25,640
always true, and one of the key, or the 
key situation in which it's Applicable 

126
00:09:25,640 --> 00:09:29,745
is, the data set has to span many, many 
orders, well many, many orders of 

127
00:09:29,745 --> 00:09:30,950
magnitude. 
Right? 

128
00:09:30,950 --> 00:09:32,950
It can't be values between, I mean think 
about it. 

129
00:09:32,950 --> 00:09:37,730
You have values between 50 and 90 and you 
select those randomly, well you're not 

130
00:09:37,730 --> 00:09:42,070
going to get any numbers with the first 
digit as 1. 

131
00:09:42,070 --> 00:09:43,400
Okay? 
And similarly if you do it from 1 to 100, 

132
00:09:43,400 --> 00:09:48,430
well you, you have this effect a little 
bit, but not universally. 

133
00:09:48,430 --> 00:09:50,770
So you want to span a lot of orders of 
magnitude. 

134
00:09:50,770 --> 00:09:54,760
And so this is sort of illustrated by 
this plot here. 

135
00:09:54,760 --> 00:10:00,120
In that the red areas are. 
Represent the probability of selecting 

136
00:10:00,120 --> 00:10:03,620
something with the first digit as one. 
And the blue areas are the areas where 

137
00:10:03,620 --> 00:10:06,903
the first digit is eight. 
Okay. 

138
00:10:06,903 --> 00:10:10,930
And so as long as you span enough, orders 
of magnitude, you get much more area 

139
00:10:10,930 --> 00:10:16,040
where the first digit is Okay. 
But, if you take a narrower case, well, 

140
00:10:16,040 --> 00:10:18,980
then the probability is more defined by 
the distribution itself and there's not 

141
00:10:18,980 --> 00:10:26,140
enough, chance for this area into the 
curve under the digits one to, to, to get 

142
00:10:26,140 --> 00:10:28,550
big. 
Okay. 

143
00:10:28,550 --> 00:10:31,810
So fine, I just want to introduce you to 
that law, mention that it can be used to 

144
00:10:31,810 --> 00:10:37,570
detect fraud and then connect the fraud 
a, plus chance of mistakes back tot his 

145
00:10:37,570 --> 00:10:42,440
original context we're in of trying to 
understand. 

146
00:10:42,440 --> 00:10:47,610
The weakness of, of, of statistical 
results. 

147
00:10:47,610 --> 00:10:52,450
Perceived increasing weakness of a, 
statistical results. 

148
00:10:52,450 --> 00:10:52,450
[BLANK_AUDIO] 

