I'm glad you've chosen to learn a little bit more about power laws. It's important that if you're going to claim that your network has a power law degree distribution, that it actually does. Because if it has a power law degree distribution with an exponent, you know, between two and three, This means that things are going to spread very effectively, they're going to be different strategies and immunizing the network and the resilience of the network to node failure as we'll learn in a little bit, is going to be affected as well, so we don't want to clean that, our network is power law if it's really not. So let's see what power laws are about. Power laws are different from the normal distribution, the normal distribution looks something like this, you have the average height for a human male, and some are taller and some are shorter. And if you look at the max to min ratio even looking at the Guinness Book of World Records, the tallest man is only several times taller than the shortest man On the other hand if you look at a city like New York City, it's 150 times, a 150,000 times, larger than a small town of Deerfield, Virgina. We have seen this before. This is the power law distribution on a log, log scale. Oh sorry, this is it on a linear scale. Where we see, like, this very sharp L. And here it is on a log, log scale. And we see that straight line. That's going to be our signature. You see these distributions in many different. Domains. And here, I should just say, these are cumulative distributions. So rather than looking at the probability that, for example, in the first graph, that a word occurs, exactly, 200 times. We're going to be looking at the proportion of words that occur 200 times or fewer. And if your distribution is really power law. Then this cumulative distribution, which is going to be its integral, is also power law with its exponent being alpha-1. So. If we see a power law in the cumulative distribution we're actually seeing a, a power law distribution. And so you can see if you look at word frequency, that is how many times in Moby Dick did the word whale occur versus the world the, the word the. Citations, how many citations did each scientific paper receive, web hits, this is the data set that I analyzed back when I was a grad student, books sold, among best sellers. Telephone calls received, now you might imagine that this is some sort of call center that received you know, over a hundred thousand calls in a single day earthquakes, many, many small earthquakes, a few really large earthquakes. Yet more power loss greater diameter Peak intensity of solar flares. Casualties in wars lots of small skirmishes and then, a few really large very costly in terms of human life wars. Income distribution. A few very rich people and lots of not-so-rich people. Name frequency, lots and lots of Smiths and not that many, I don't know. [laugh] You know, lots of people have less, less common names. Populations of cities. We just talked about that. A few mega-cities and many, many small towns. The para-law distribution written out. We've seen this before, is that the probability of X, say a word occurring 100 times is equal to some constant times X to the minus alpha where alpha is typically between two and three. And if we take the log of both sides, as if we were plotting on a log-log scale, we indeed see a linear relationship between the log of the probability of observing X and the log of the value X itself. We often hear power law networks refer to as being scale-free. What this means is that the power law looks the same no matter what scale we look at, so whether it's in the range of two to 50 or 200 to 5000. And this is only true of a power law distribution because if you change x by some multiplicative constant that probability is just going to have another constant in front. You also hear power laws sometimes referred to as Zipf distributions. George Kingsley Zipf was a Harvard linguistics professor who first looked at the distribution of frequencies of different words in text. And some words are much more common than others and there are all sorts of fun models trying to figure out why this is. And one such model is that you're throwing down letters at random but you also have, you include the space and so the words are generated. As just random letters, and, And terminated by stasis and that actually gets pretty close but turns out to not be the correct model. But when Zipf was looking at say, for example distribution of, of words and texts or the sizes of cities, another subject he studied. He turned it in terms of the size of the Rth largest city and so he got power loss with an exponent data. And you can actually match the two exponents because saying the Rth largest city has uninhabitance means that R cities have n or more inhabitants, right? All the others are smaller, so this is in fact our cumulative distribution. And so you can flip things around and find that beta is actually equal to one over alpha minus one. So. If you have a power law distribution with an exponent alpha of two. This is true of the number of unique visitors from AOL to websites. This is a very old data set, I think 1997 or so that I studied as a, as a grad student, When you plot the ranked plot, you're going to get an exponent beta of one. Exactly as you would expect. There is also a, a kind of a 80-20, rule, so, Using the exponent alpha you can actually figure out whether it's 80-20 or 90-10 or whatever else. And this is just the amount of wealth that is in the hands of the richest P proportion of the population. And for example, for wealth distribution in the US, which is getting more and more skewed, the alpha exponent's 2.1, meaning that the richest twenty% of the population holds about 86% of the wealth. Now let's get to the meat of the matter which is how do you fit power law distributions? How do you know that the degree distribution that you measured on your network is actually power law? And. What you start out with is this histogram. You know? I have this many nodes with degree one, this many nodes with degree two. Blah, blah, blah, blah, blah, right? And you may think, well, it's supposed to be a straight line on a log, log plot. So, what I'm going to do is I'm going to plot it on a log, log plot. And I'm just gonna try and fit a line to it. And this is problematic, Well, we'll see how its problematic. Let's first see what it looks like on an artificially generated data set. So I just generated, lots and lots of random variables but I generated them in such a way that they're distributed according to an alpha exponent of 2.5, so we're going to know that a fitting method that gets close to 2.5 is actually, actually a, a better one. So we're going to start with our little linear scale plot and you see it already seems to zero out by the times it gets to twenty. And over the whole range you see this L which isn't very informative and this is it on a log, log plot and here actually looks pretty good. It looks like we can plot this. There is a basic problem which is that here we have 10's of 1000's of observations right. This is something like a 100,000 of the data points that are. Generated were, you know, the number. I guess this might be three. Right? But then where we get into numbers that are over a 1000 right there many (missing bins) right because we just didn't get any number of, you know, 2300 and 73 or something like that. Maybe we didn't see that in this data set and so we have a zero. Which, you know, you can't take the log off so it's just missing on this plot. And each of these points becomes an idividual observation versus we have a 100.000 data points going into. Here, meaning that you know, we should be weighting this more. But we're not if we do a simple linear regression. And in fact, what you see happening is if you do a best regression fit, it's going to be totally misled by these points out here and you know, the missing data that isn't represented. And you're going to get an alpha fitted exponent that's way too low. Alright. And so here it's actually, saying oh the alpha may be something like 1.6. Which, you know, it's not right? We know that the data has, came from a distribution where Alpha equals 2.5. And this is just reiterating. You, few bins here, many bins here but you know we actually have more, more data here than out here. So the first solution is to bend logarithmically. That is, you want to actually get more data here so you're going to have wider and wider bends. And they're going to start you know, one, two, four, eight, sixteen, 32, or you can, you can decide what exponent you want. And then you get these nice evenly spaced data points which you can then fit. And if we do this we get an alpha of 2.41 which is actually pretty decent. But what it's doing as. Well, this is kind of smoothing out the data and losing some information, right? When you aggregate into these bins, you don't... There may be interesting features that you just kind of smoothen out and that you won't recognize. So the second solution is to do cumulative bending, like all those pots I've showed you with solar flares and family names and so on. Those were cumulatively bent. There's no loss of information because Each value captures, you know, how many, exactly do you observe, You know how many observations do you have of this value or lower? And so you don't have kind of zeros that are falling off or something like that. And so, you can do just the regression and get this, cumulative exponent, which is going to be alpha minus one, so you'll have alpha, and that's also a fine way to do it, and here, we get alpha minus one is 1.43, and that's much closer to the actual 2.5. So that's, that's, perfectly fine. There's also the question of where do you want to start fitting. Now, I drew from a pure power-law distribution. Because I was generating thick data. [laugh]. But in reality, you know, data isn't going to be so pure so there may be some deviation on the lower end. For example, if you look at The number of links that different university websites get, Each university gets at least a minimum number of links. So you're not going to get, like, the, the, I guess from your point of view, you know, lots of universities with only one link to their website. I mean, they wouldn't be a university if that was the case. So you may see something that's somewhat like this. So really you just want to fit out here, you want to fit the power law there, but not be mislead by the low value. So you can establish an xmin, a minimum value. The load which you're not going to be fitting. So if we look at, for example citations, maybe the power lies only evident for papers that were cited at least 100 times. And so you can set, X-Men two to be 100 and then fit the tale. There is the very, very best [laugh], method. Which is this max likelihood method. That calculates using every one of your Data points. What is, you know the observation that you have, the data points that you have, what is the most likely distribution? What exponent would it have that would have produced the data that you see. And if you want to use this kind of method. You can download, and I'm going to provide the link separately, from Aaron Clauset's site. And it's basically Python and R codes that, you know, you can use either one. [laugh]. You don't have to use both. Even that lab code, I think. That, given a data set, we'll do. Well, we'll first evaluate, you know. Does this is this a good candidate for a power law fit versus another kind of distribution such as the log normal or the stretch exponential etcetera. And then if it is a good candidate for a power law fit, what is the power law. So if you're getting really serious about characterizing your network I would recommend doing it properly. And this is the way it's done these days. So here is some exponents for real word data. You can see that most of them tend to be between two and three, but drop a little bit below two. Now this is weird because these distributions would have, you know, the kinda infinite variance, so most likely they have a drop off you know, the power law can't extent indefinitely. So those are different assets that word networks, many real world networks are power loss. So if you look at film actors this network is who has acted with whom in a movie and this is, there's a funny story around this Which is that the Sun Belt Conference which is the annual social networking conference had a data analysis challenge which was to take the Internet movie database and plot out the network, all the actors who has co starred with whom And calculate things such as who're the most central actors. And little did they realize that there's a certain kind of movie, namely pornography, that produced very high centrality actors because A single actor can, kind of, star in many, many movies, right? They're made at a very rapid pace, and so they... Until the, the, the, people participating in the challenge, you know, met up at the sundown conference and started comparing notes, you know, no one really expected that the porn industry would be so very central in the movie collaboration graph. Anyway there's the telephone call back which we already talked about right and who calls whom. Email networks, sexual contact networks which we talked about as well, the world wide web we talked about, the physical internet also paralyze some very large hubs Peer-to-peer networks, for example for file exchanges, some, some computers tend to be very central in these networks. And then different kinds of. Biological networks where you have either metabolites or proteins interacting with each other. But it's important to note that not everything is power law. So bird species abundance is not power law. Size of wildfires is not power law. Even though some wildfires are very big and many are small. Even networks that appear power law at first, especially the social ones, for example, email. If you just take who has emailed whom ever, excuse me. You may get a power law. But once they start restricting that like, the other person actually replied or maybe there was a little bit of back and forth like some actual interaction. It ends up being not power law because people. You know there really limits to how many relationships you can maintain and so you lose the power law property. The electrical power grid, not power law. I guess the thesaurus, there isn't really a word that could be a synonym for every other word. The networks of company directors, so this is if you look at the boards of directors of different companies, so, how, on how many, you know, how many other people did you jointly sit on boards with? And even very, you know, hardworking, very prominent individuals only sit on so many boards, and so they only have, you know, a limited degree in such a network. So here's just an example on the AOL visitor's data set and if we try to fit it directly we get an alpha that's way to low if we do exponentially wider bins, we get a seemingly reasonable slope of 2.1 But you know there's, it's real data so if we do this kind of exponential bending we may lose information about you know, there's actually a little bit, a deviation from a parallel where we don't have that, as many sights as you would expect with, With very low numbers of visitors. And then again in the tail it doesn't quite match up to the power law. And so you know, what you may end up doing is fitting a power law with an exponential cut-off, which just says up to this value kappa. It's a nice power law and then after that it starts to decay exponentially. And there could just be different reasons for this. For example, constraints such as how many people you could ever interact with. Or you know, you can only there only so many individuals. Available to connect to, et cetera, right there could be different reasons for these cut offs. So, just to wrap up. Power laws are cool and intriguing, but you should make sure that your data is actually power law, before you boast to others.