Hi. Welcome to Week three of Social Network Analysis. This week we'll be tackling Centrality. Centrality tells us which nodes are important in the network. This goes beyond just looking at social networks and which nodes might be popular. And to looking at all sorts of different networks, trade networks. Who are the most important importers, exporters? nodes that bridge different regions and facilitate trade. In biological networks. Which genes, for example, play a very crucial role in regulating entire systems and processes. And infrastructure networks, which nodes are critical, such that if you were to remove them you would significantly impede the functioning of whatever the network is supposed to do. So identifying such important nodes is, is pretty critical and it may be one of the first things you want to do when you look at a network is look at what nodes are more central. If you go to the website, moviegalaxies.com, What you'll see are a bunch of networks that have been constructed out of the scripts. So they parts, I'm guessing automatically, through the scripts, and see this character address this other character at this point in the movie script. So let's draw an edge. So if you look at, for example, the 1993 movie, The Fugitive, you can see immediately that there's this central node here, Kimball. And this is Dr. Richard Kimball, who is the fugitive. Falsely convicted of, Murdering his wife. And you see here a secondary important individual who is Deputy Gerard from the US Marshals Service who is pursuing Kimball throughout the movie. And this you can immediately infer without ever having seen them. Well you can't infer all those details but you can infer that both Kimball and Gerard are central actors. Similarly if you look at Heap you see a similar pattern. Where Detective Hannah is, pursuing Neil who's like a, a long time crook, right? And you, you start to see, these kinds of patterns. Now, of course if the, if the movie just has one central actor. For example, in Memento, there's just Leonard. And he's the most central. So even though he may not actually remember having all these etches, all these interactions because what happens is he keeps, he keeps losing his memory, We can see, in fact, that over the course of the movie, he's the most central actor. The question is though, is counting edges enough? So what I'm showing you here, is a fragment of the network. My first social network that I ever gathered data for. And this was, these were Stanford homepages. So take yourself back to 1999 and this was before Friendster, before Facebook, before MySpace really. People were still constructing social networks. Online, and the best way that they knew how, which was to link their home pages together. And the way that I caught on to this was that well, this was actually back in, in, my undergrad days. A friend of mine had linked from her home page to mine and said hey Aida, how come you're not linking back to me? And so I started... Actually my home page had a friend's. Page that linked to my friends, and once I was at Stanford and in grad school and starting to study networks, I grew curious. You know, how many people actually did this. And so I gathered the data, and this is the local neighbor, my local network neighborhood in the Stanford home pages. I linked to Orkut Buyukokkten. Who here looks very sensual. He, he is also linked to or, or links to many other friends. And in fact Well, I mean he was sort of the center of my social universe in, in graduate school. Now the complete picture though is something like this, right. So Sure, you know, if you look at degree, Orkut may be one of the highest degree individuals in the network, but still you would probably say that he's peripheral to this whole network. However, that didn't last long, because Orkut actually created Club Nexis, which was it had exactly the same functionalities as Facebook, was created around the same time as Facebook, but was for Standford students. And there he definitely was the center of the network Then when he eventually started working for Google, excuse me, he, He created there for social networking service that was called orku.com and was incredibly you know, successful there so much so that, Well where that service flourished was in especially in Brazil and in India, and so when he went to Brazil he actually needed a security detail. He'd become so, so very central thanks to starting this hugely popular social network. But back to this picture it tells us that back in 1999 when Orku was still a new grad student. Maybe you know, his close friends have come to appreciate him. But he was still not the center of the Stanford social life. So what kinds of centrality measures could we use in order to make that distinction? So there are many different notions of centrality. We talked about in-degree and out-degree before and we'll review them but there are also notions of between-ness closeness and we'll also talk about eye conductor centrality. So in-degree is just the number of incoming edges that a node has and here X has higher in-degree than Y, Y happens to have in-degree zero. I have here plotted the network of oil experts using the NBR United Nations trade data. In fact I think they, they based the data in imports rather than exports but you know you can because they consider them more reliable but you can kind of symmetrize the, the two and I use the category that, that corresponds to petroleum and petroleum products. And you can see this, every node is a country and there's a directed edge from the exporter to the importer. So just to check your understanding which countries have high in degree, that is, they import petroleum and petroleum products from many others? Now let's talk about out degree, out degree is just the number of edges that eminate form a particular node. So here, I have just modified the same network of oil trade. And I have sized, the nodes by their exports. And I've also colored according to the ratio of exports to imports. So if that ratio is high, the node is more blue. If that ratio is low, the node is more red. So, for example, here, the US imports much more than it exports. So just to see if we know what's going on. Which country has low out degree but exports a significant quantity of petroleum products. In this First visualization. You may have seen Singapore actually have a non-significant out degree. And that's a little bit puzzling because Singapore is a city-state. So where is it getting this oil to export? So what I've done is I have excluded you know, everything except. For crude petroleum, that is the actual oil as opposed to derivative products. And here you can see, that Singapore becomes primarily an importer. It has indegree, but it doesn't really have significant outdegree. So. This illustrates the point that when you're talking about centrality, it's very important to know first, what does your network represent. What is it's scope? Here scope is not that much of an issue because we have all countries. However, you know, what exactly do the edges mean? Is it a crude oil or is it refined oil, even a little tweak can make a significant difference. It's you just have to know what your network is, and when you're talking to others you have to honestly represent that as well. So let's put some numbers to it, let's quantify centrality. If, and let's switch to an undirected mode and networks for the time being. So, in undirected degree all we're doing is we're counting the number of edges you have so if it's a friendship graph it just counts the number of friends. And the basic assumption here is that, it doesn't matter, how many, friends your friends have, or, you know, anything else. You're just counting the number. That's, that's all that counts. You can normalize this measure, by dividing, by the maximum, number possible, for the degree, which is n minus one, as we saw, you know, just, just last week. So. Here. This node has a normalized centrality of.5 because it's connected to three of the possible six. Other nodes that are in the network and although I don't personally normalize centrality degree that often you will just for my own purposes I, I like having the raw count. And especially with large networks it's you know, have a network of a 1,000,000 nodes no one's is going to have you know all this normalized degree centralities are going to be.0000 right because you are not going to be connected to all the other nodes that's ridiculous. But a lot. Social analysis software does like to normalize. So I would just like you to get used to these normalized scores. And of course you can just multiply by n minus one to get the, the, the degree back, the number of neighbors for each node. Centrality refers to an individual node. But there is a network wide measure called centralization, that characterizes the entire network and what it does is it captures the inequality and the distribution of centrality. So it could be the inequality of degree between the nodes. And for this you know, it, it doesn't. Matter that much what particular measure of centralization used? You could just use the standard deviation in degree. And I have used this for example, when studying trading networks for commodity future's contracts. Right? We just looked at standard deviation in degree. But we also threw in something called the genie coefficient, which varies from zero if everyone has the same degree to one if one node has you know, all the degree and all the other nodes have. Degree zero, which you might see in a star, formation. There's also Freeman's General Formula for Centralization and what you look at there is you sum over all the nodes, you take the maximum value that you observe in the network, you subtract The centrality of that particular node and you do this for all the nodes, and then you normalize by this number of, of pairs, right. And it's just one lower because you get 0s for the node that has the, the highest, centrality. So. Let's look at a few examples. This is with Freeman, centralization. Here, as promised in the star formation, you have perfect centralization. You got 1.0. Here in the line graph, The nodes look pretty similar as far as degree is concerned. We'll see later that using other centrality measures, right, the one that's in the middle might end up being more central. But here, all we're doing is counting degree. So, all these nodes look pretty similar, giving us a relatively low centralization value. And this similar thing happens for this kind of butterfly network where, again, the nodes have either Degree two or three, and it's all very similar. I mentioned these trading networks that I had analyzed. So these are brokers and they're trading in futurist contracts having to do with the S and P 500. And in this case what I've done is I've sized the nodes by their n degree, so the larger the, the node appears, the higher its in degree, and I've colored them by out degree. So the larger the out degree, the, The darker the note. And in any case you can see that here the network has very high in centralization because one node really dominates it has a very in degree and most of the others have relatively even in degree and lower in degree and here on the right you see that actually all the nodes have relatively even indegree and no one is really hogging all the N-links. So, there is an example where we could actually use centralization of the home network and correlate it with other variables. For example, the returns and volume and other kind of financial properties of the system at that time. But let's return to this question of things that to, that simply counting the number of edges doesn't capture. If we look at this graph degree may very well capture what's going on but if we look at. This network, Right, this node has Degree two which makes it kind of less importance and in the degree sense then these two nodes and you know about of equal importance so these two nodes. But you know, in two degree, we think that this node must be important because everyone here who may want to get in touch with any one here is going to have to go through this node. So how, how would one capture that. Similarly here, degrees has all of these are equal, but this one if it wanted to visit any of other nodes it would just make two hops max. And all the other nodes could potentially have to make three or four hops. If they want to reach some of the individuals in the network. So how do we capture that? Well let's return to the Standford Social Web. Right. Here again, we have that Orkut is very central locally. And, certainly for my social, neighborhood.. If any of us want to get in touch with under grads and so on, may be, you know, Orkut would be our best, choice and he would kinda be between us and the rest of the network. But, overall, his betweenness might not be as high. As for example, some of the other nodes kind of closer in, and his closeness, in a sense that he is not, you know, very proximate to the rest of the networks should be low as well. So let's first talk about brokerage. This is where a node lies between other people, here X plays more of a brokerage role than y, because. These two nodes would have to go through X to talk to these two nodes. And there is a bunch of research into this. In particular, Ron Burt at the University of Chicago Business School, has studied. What advantages individuals who occupy such roles might have. And he's found things such as, they're more likely to generate creative ideas. If he studies them within an organizational context, their ideas are more likely to be accepted. They're more likely to receive favorable reviews, and they're more likely to be credited. Now of course, there's a little bit of an ambiguity. Is it that the person enjoys these advantages because they occupy that position in the network or did they come to occupy that position, because they're a special kind of person who does generate creatives ideas, etcetera? Right? And this is a little bit of an, ongoing, debate. The consensus seems to be that, you know, if you, if you put someone incompetent in that position, they're not going to suddenly, start to blossom and do really, really well. But if you have a good person in that position they're an even, in an even better position, or can, can accomplish even more when they're in the brokerage position. And this goes back to the notion of constraint. And that is that, if you are not a broker You experience constraint in the sense that your contacts are talking to each other. Which means that you can't say one thing to one and another thing to another because they're communicating with each other. And so you are constrained in what you might be able to pull off. On the other hand, if you're in this brokerage position, you can turn to one part of your network and say hey I would do business only with you. And they, you know, they have no reason not to believe you. But you could turn around and, to the other side of the network that you're kind of hm, bridging and you can say oh you guys are my favorites and no one would be the wiser. And so this is what's called like when you're brokerage position you actually experience low constraint. So let's look at this measure of betweenness which is going to capture brokerage. The basic idea is that you're asking how many pairs of individuals in the network. Would be connected through you, on their shortest path to each other. So are you on the shortest path between those, other two nodes? And, it's possible for there to be multiple short paths. In that case, what you're going to get, so here, you're I, and you are on the. One of the shortest paths between J and K, or one or more of the shortest paths between J and K. But the credit you get is the number of shortest paths that you're on divided the number of shortest paths that J and K have. And then, you're between-ness is the sum over all pairs of nodes, JK. And you can normalize this, again I, you know for between us it might make more sense because these values can indeed become very astronomical but you're just normalizing by the number of pair of vertices excluding I itself so for what fraction of the other nodes are you on the short path? So let's look at a few examples on Toy Networks here. The central node has between this ten. Because there are, five four / two, because we don't care about direction. Different pairs that go through the central node in order to reach one another. But each one of those nodes individually has between a zero.'Cause no one has to go through them to got to anyone else. So let's return to our line graph. And here, naturally A has zero betweenness because it's not on the path between any other nodes. B has betweenness three because A relies on it to reach every one of the three other nodes. C has betweenness four, because it is on the path between A and B, and D and E. Note that in this case C gets full credit because there are no alternate paths, for example between B and D, the only way they can reach other is through C. If we go back to the butterfly network, remember this central node has degree two Now it has the highest between, as it said between this nine because each of these nodes relies on the center node in order to reach the other three nodes in the network. On the other hand, these peripheral nodes here have between a zero because, you know you wouldn't go out of your way in order to reach any of the other nodes. They, there are, they are on those shortest paths. Let's look at this particular toy network. Here C and D each have between this one. How do you calculate that? Well, c is on a path from a and b to e, and so is d. So imagine first that d is not there. Well, if that were the case then c would have between this too. Because it's on the path room a to e and then the path room b to e. However d is on the exact same path. And so they share credit and they each get between this one Quiz question, what is the betweeness of node E? Here's, an, older version of my Facebook network. I've sized the nodes by degree. The, larger the size, the higher the degree. And I've colored them by betweenness. The redder the node, the higher the betweenness. And you can see some nodes here that have, relatively low degree, but they have high betweenness because they bridge, for example, my, undergrad network, with my University of Michigan and research, network. Another quiz question, find a node that has high betweeness, but low degree. And now, find a node that has low betweeness but high degree. What if it's not so important to have so many direct trends? So doesn't matter whether you want to do this power play where you're brokering between different parts of the network and you want to make sure that they have to go through you. What if instead all you want is that you have. Easy access to a large part of the network. You want information to reach you in a short number of hops, you want to be able to disseminate the information you have. Therefore it's not about brokerage, it's about how far away the rest of the network is from you. For example, in this network, this node has degree one but because it's friends with the highest degree node this means that the number of hops from it to anyone else in the network is only one extra hop then it is for the most central node. And in fact has easy access through just that single connection, so. We want to capture this no, notion of a small number of hops. So by definition, the closest centrality is the sum of the distances between the note in question I, and all other notes. [inaudible] it's J. And you're going to sum these distances and invert them. Now, there's a problem in that, if your node has multiple connected components. That is, there are nodes in other components than yours, your distance to them is infinite. And when you sum infinity with other things, it means your, your closeness is actually zero. Right? Because here, you're inverting. So in order to avoid that, an alternative measure of closeness is to take this inversion inside of the sum. And so you take the inverse of. The distance to every other node, and you sum those value. So. All those nodes and other components, they just contribute zero to your closeness and, but you do get credit for having short paths to nodes that are in your same component. And you can normalize this by n minus one, the number of other nodes in the network. Here is a toy example where if we look at the closeness of A, it has distance one to B, distance two to C, three to D, four to E. We sum those, we divide by n-1 and we invert getting a closeness of.4. Finally we can also look at our other toy network and here again the middle node has high closeness. The ones right next to it also have relatively high closeness. Here the middle node has high closeness but the other nodes being directly connected to this hub, that is the spokes they actually have relatively high closeness as, as well. They're only two hops away from any other node. So looking at this network, can you tell me which node has relatively high degree but low closeness?