[MUSIC]. Okay. So, what's a second reason for this decline effect for the truth wearing off? Another one is just people make mistakes and in some cases there's is This is a fraud. And so there's some evidence here that these kinds of things are going up as well. as measure, one way to measure this is by the number of paper retractions, and the number of retractions has gone up fairly significantly as Norton pointed out in 2011 in an article in Nature. So, from the period of 2001 to 2011 there has been a tenfold increase in number of paper retractions, but only in 1.44 fold increase in papers. Themselves. So this plot is from two different repositories, PubMed and Web Of Science, neither one of which can include Computer Science papers by the way. But you can see that it's gone up pretty significantly. Fine. So, not going to dive into too much detail into this Phenomenon. But I do want to give you one statistical tool that is sometimes used to detect fraud. And that's Benford's Law. So, if you're not familiar with Benford's Law, it's a fun one to be familiar with. So this. Law predicts the distribution of the first digit of data. Okay. And if you think about it without think- -- if you think about it without thinking too deeply, you might intuitively think that this would be fairly uniform. If you have random data, you might see an equal number of 8s and an equal number of 3s and an equal number of 2s and so on. But it turns out that the distribution is not random un- -- in, in some circumstances, in many circumstances. Well, it's not even, I should say. It is random. It's not even. The distribution looks more like this. You will get more 1's in the first position than you will 2's and more 2's than 3's and so on. Until up to 30 percent of the numbers you're measuring are 1's. So this should, you know, blow your mind a little bit, okay. So some examples. Of this before we explain what's going on. taken from there's a nice website testingbenfordslaw.com where they pulled real data in, and, and showed the plots. This is the number of Twitter users by their follow, by the number of followers they have. Sorry not the number of Twitter users, Twitter users by the number of followers they have. Have, okay? Alright, so the, the, the list of numbers is just the number of followers. The number of Twitter users that have, whose number of followers begins with the digit 1 is 32.62%, if you can see that. And the red dot is the prediction made by Benford's Law. so not too bad. The number of instances that have the two as the leading digit is 16.66% and there's the red dot predicted by Benford's Law. So not too bad. The distance of stars from earth in light years Follows a similar pattern. This is remarkably close to Benford's Law prediction. Government spending in the UK between the period of time May through September of 2010. Now, you might imagine there's some selection bias on this particular website, and I can't guarantee that there's not. But, you know, with some reading plus a little bit of trust that I've hopefully built up. I hope to convince you that this is not fully explainable by this website, choosing particular data sets for which this is true. Okay. Google Books, the number unique 1-grams. And we talked about one grams several, a couple weeks ago. Terms, essentially, are what grams are. Okay, again, pretty good prediction. So, before we say, before we give the intuition for this. If there, intuition's a little bit tricky, but if we attempt to give the intuition to this. You can use it to detect fraud. And so this was attempted, or this was an experiment was done by Diekmann in 2007 to see if this could be used for detecting scientific fraud. And so what he found was that first and second digits of published statistical estimates were approximately Benford distributed in real studies, okay, or at the very least they had kind of a monotonically decreasing distribution. So ones, more, more ones than twos. And more twos and threes and so on. If not exactly Benford's Law. And then what he did was asked subjects to manufacture regression coefficients. You know? Basically fitting a line manually. And found that the first digits were actually hard to detect as anomalous. But that the second and third digits deviated from expected distributions. And so, this distribution that I gave you in, in the last few slides. And, and that I gave you in the, Actually I guess I didn't give you the actual formula here. Okay so the distribution that we'll be discussing is only for the first digit and the skew in that distribution actually gets suppressed as you go the second and third digits. But Benford's Law can also be used to express different, yet still measurable distributing of second and third digit. And so, there, the second and third digits deviated significantly. And so the conclusion was, this is, it is potential, potentially useful as a fraud detection tool in scientific data. And actually, there are instances where Benford's law has been admissible as evidence in court in cases of fraud, and it's been used by reporters and so forth to argue for evidence of fraud in cases of, voting, election, and other kinds of Accounting data on the sort of global scene, okay. So what's going on here? Well, one way to think about the intuition here is, imagine a sequence of cards labeled with a particular number. 1, 2, 3, 4, 5 all the way up to, you know, 999, 999 fine, okay. And so put them one by one in a hat, in order. And at each, every time you throw a card in, measure the probability that a random selection from the hat would produce a card where the first digit is one. Sorry, this isn't very well said. I, it's not. It's not drawing the number 1, drawing, a card, where. The first digit is 1. Okay. So what does that probability look like? Well, this figures a little bit misleading because the X-axis is on the long scale, but if you just look at the heights. What's going on here is on the y axis is the probability of drawing a, a particular digit. The blue line is associated with the digit one, the green line is a digit two and so on. And this was generated by a simulation, where you, you know. I really did select random numbers from a distribution. And I really did order them in the manner described in the previous slide and put 1 in and measured the probability drawing it value of number 1 or with the vertices of just 1, okay. And so if you think about this, the numbers You know, what's, what's happened here, this is where 1 and this is 10 and this is 100 and so on. Well, as soon as I put a 1 in, the chance of drawing a card with a 1 on the front is 100%. When I put a 2 in, it's now 50%. When I put a 3 in it's 33% and so on. But then as soon as I get to 10 It goes back up to a higher percentage again and it stays there, 10, 11, 12, 13, 14, and so on. And then as soon as I get to 20, it drops down a little bit. And so, that's why it's climbing here, through the tens, and then it starts dropping again. Again sharply, okay. And so the point here is that under this model, values where the first digit is one always are in the hat already by the time you get to the twos. And the twos are always there before you get to the threes. And the threes are always there The fours and so on. And so you end up with these height, these peaks are lead by the ones. And the area under this curve. Although, remember this is a log plot so the area's not quite right. represents the probability of, of drawing this. And so the probability does actually get higher. Okay. And there's a few different other models you can, you can find if you read up on this. And I encourage you to there's other ways to sort of think about the probability here. And there is actually a closed form expression for Benford's Law as well. Okay. One of the limitations here, it's not always true, and one of the key, or the key situation in which it's Applicable is, the data set has to span many, many orders, well many, many orders of magnitude. Right? It can't be values between, I mean think about it. You have values between 50 and 90 and you select those randomly, well you're not going to get any numbers with the first digit as 1. Okay? And similarly if you do it from 1 to 100, well you, you have this effect a little bit, but not universally. So you want to span a lot of orders of magnitude. And so this is sort of illustrated by this plot here. In that the red areas are. Represent the probability of selecting something with the first digit as one. And the blue areas are the areas where the first digit is eight. Okay. And so as long as you span enough, orders of magnitude, you get much more area where the first digit is Okay. But, if you take a narrower case, well, then the probability is more defined by the distribution itself and there's not enough, chance for this area into the curve under the digits one to, to, to get big. Okay. So fine, I just want to introduce you to that law, mention that it can be used to detect fraud and then connect the fraud a, plus chance of mistakes back tot his original context we're in of trying to understand. The weakness of, of, of statistical results. Perceived increasing weakness of a, statistical results. [BLANK_AUDIO]