As promised, I'm going to be telling you about some applications of network centrality, that I actually did myself The first is A hospital patient transfer network. There are thousands of, transfers between hospitals in the United States. And my colleague Jack [inaudible], was doing this study based on Medicare records. So people who are on Medicare who constitute roughly half to 60% of these transfers. When they transfer between hospitals, there's a record of it. And he, he was analyzing this data set. It makes for a very pretty, network. But it also highlights a potential problem and that is that there are these highly resistant bacteria that sometimes take hold in hospitals and their because of their drug resistance it is very, very difficult to treat patients who acquire these infections and it is actually rather, rather costly right because your first line of defense the antibiotic will not work. And there is some documented evidence that patient transfers have lead to also transfers of these bacteria who were hitching a ride with the patients when going to a new hospital. Now. Pretty much, it is each hospital trying to prevent these infections on their own. So measures that they can take, for example, are, keeping an incoming patient in a separate area. And when someone enters this area, they have to be wearing a mask. They have to wear gloves. They have to, they have to follow protocol. Of course, this actually costs quite a bit of money. So the question is, could you, formulate more effective. Infection prevention strategies if the hospitals could actually coordinate within this network, and pull their resources, and put prevention resources at hospitals where the total number of infections across all hospitals would be minimized. If you did the allocation correctly. Now what we assumed was that the more resource a hospital has, the less likely it is to acquire an infection. But once it is infected it cannot really root out such an infection and it will actually be sending along infected patients to a neighboring hospitals so. You know, one, one solution, of course, would be to, forbid all patient transfers. However, there are many benefits to such transfers. Starting from some hospitals being overbooked, to their intensive care unit can't take any more, patients. And so they transfer to a second hospital. Or, a hospital does not have the needed expertise, and a neighboring hospital does. And, in general, patient outcomes are better when patients who need to be transferred or transferred. So. We really do just want to allocate some prevention resources as supposed to, you know shutting down whole hospitals or something like that or forbidding transfers. So there are different Strategies that you could use, you could just allocate money at random and here what this shows is what the, the number of infections would look like after, you know, a year And, so, you can see that, you know, there are really some hot spots, some hospitals that would have been very likely to be infected, if an infection start in a random hospital. So, you see, well, okay, we know that some hospitals have more transfers than others. And, so we tried a, a strategy that would target those hospitals with high degree. And here, you know we have the choice of in degree and out degree, some hospitals receive more patients, some hospitals spent, send more out, and what we found was, most effective was to take The geometric mean, that is the in degree times the out degree, and the square root of that quantity. And you can see that this produces a significant reduction in infection rates. However, we also know about betweenness. Right? And we know if a hospital is sort of bridging, say the East and West coast then Being able to, you know, target those hospitals are kind of between a lot of other hospitals. That might be even more effective, and indeed it is. When you target based on betweenness as opposed to degree, you reduce infections even further. And finally, what, what is most effective is, actually simulating the infection spread and, running a greedy algorithm to see which. Hospitals should be targeted, and these typically are a high degree and high between these hospitals, but you can do even better by simulating in this greedy manner. And a collaborator of mine at CMU [inaudible] actually showed that an eigenvector approach was most effective. It was a really good approximation to this greedy algorithm. And so, we kind of ran the full gambit from degree to between this to eigenvector in solving this problem. Of course I am saying solving. [laugh] It is kind of solved in, in theory, of course, from a practical policy standpoint. It is not clear how you would get hospitals at a national level to, to pull resources. But still it is an interesting thing to consider if there are, for example, federal funds that can be allocated between different hospitals. The second example, I have already mentioned a little bit when we were talking about preferential attachment and scale-free networks, and this was identifying expertise in online question and answer forms. And the reason why we may want to identify expertise, is that there is this response time gap between when an expert asks a question in Sun's Java forum, they may be waiting for hours or days, but when a newbie asks a question there their question is answered right away, and the question is could. You speed this up, by matching experts with questions that were post by experts, in order to, to get those taken care of sooner. And so what we did was we, we gathered a whole lot of data from Sun's Java form. And you know we had what is this, 200,000 users and 800,000 edges, so plenty to play with. But you know, with real world data the thing is always you have to decide how you're going to construct the network, so you could construct it in an unweighted way. What we wanted to do was have, for example, if you have A who posted two questions, and then you have B and C. Who answered those questions. You have some choice. You could, for example, do an unweighted network where A has directed edges to B and C. And in this case, you know, we are saying B is more expert than A because they could answer the question. We could weigh the edges by the number, or weight to the edges [laugh], sorry. By the number of threads that, that the two notes intracted in. We could have them share credits, so both b and c answered a question. They would each get half of, well, weight one-half added to whatever other weight they had, or we could even introduce a little bit of back flow. And this just means, if you asked a question that was answered, it means that at least it was sort of decent question. So, we may have a little bit of the weight of the edge going back. Now, as I mentioned before, there is very uneven participation, especially when we look at nodes that post lots and lots of replies. And, in addition to looking at these kind of individual note centralities, we can look at the whole structure of the network. And something that I have not really discussed before is the bow-tie model of the web. It is very simple. You have the strongly connected component. Remember, on the web this means that you could actually, if you are very patient, and kept clicking, get from one page to any other within that component. And then you have the in component. From this component, you can start on a web page, and end up in the strongly connected component, but you cannot get back out. The out component means that. You can get out of the stronger connection component into this region, but you cannot get back in. And then you can have other sorts of structures. And what is true of the web is that this, this At least a decade or more ago was that you had this very symmetric bow tie. And the, the symmetry is not there, it is very skewed in Sun's Java form because well you do have a strongly connected component but it is actually relatively small. Only about twelve percent are interacting in such a way that they are all kind of helping each other in a generalized reciprocity kind of way. So maybe A did not A replied to B, and B did not reply directly to A, but B replied to C who, replied to, to D, etcetera You have a large in-component. These are people who are asking questions, but not necessarily replying or if they do reply they didn't, reply to someone, in here. And so they all composed this in-component, and there is a small out-component. These are people who only answer, but do not but do not ask. Okay, but back to the, the question of centrality. As I mentioned before, it was very clear. You know, that some nodes were very active and very central. But what we wanted to know was whether indeed degree was the best way to, to capture centrality. And so, in order to evaluate, we had human readers, read the posts and judge the expertise level of the repliers. And what we found was actually a very high correlation over a number of centrality metrics between the human rated expertise. And just this, you know, whatever centrality measures, we were deriving. And so, in a way, that was nice. Because you basically could not mess up. You could use any measure and do really well. But what was disappointing to us was that page rank was not, the clear winner. And the reason why we had hoped that page rank would be, You know, a good, a good centrality measure is if you think about it, if someone is. Expert enough to answer the question of another expert who themselves has been helping others. Then, you know, that kind of propagation, should tell us that this expert is so much better than this other expert etcetera. So, the recursion thing apparently was not working in this form. So, we wanted to know why, and we did look at the break down of the human reading versus the Versus the kind of centrality measures we were getting and that was not really telling us why page rank was not performing all that awesomely. And then Jun Zhang built a simulator. So from last week you may be familiar with, hey Sometimes just simulating how a network forms can give you insights into whether you know, you understand the dynamics that are leading to the network, right? And the dynamics may be really the things that you are actually after. And so, here we. The parameters were, who is going to ask more often? Is it going to be the newbies or the experts? And the one that seemed to resemble more of the real network, just because of the skew, was that newbies ask lots and lots of questions. And then, who is going to answer whom? We wanted to have a particular matching algorithm. And so, in one case, we had the best preferred. That is, no matter, you know, what the expertise of the person asking, the. Expert who has the, the most difference in expertise, that is, their, their, as. As expert as possible over what is being asked, is going to answer and we had another Model which was the Just Better model so that people who are answering would be choosing to answer questions that would be challenging to them. So if a newbie asks a question, then just the people who are a little bit more expert than newbies would be answering the newbie questions. And the top level experts who cannot save their badnwidth and only answer the questions of those who are like almost as expert as they were. And this actually lead to two different kinds of topologies. In the best preferred network you can see, so this is. A, a circle layout where I have just ordered the nodes by the expertise that we assign them in the simulation and you can see a lot of the questions are answered by the experts and also, so here they are showing larger nodes they are more central in the network. Here with the just better, you can see that the 2s are answering the 1s, the 3s are answering the 2s etcetera And, in fact you have that the experts are kind of being pushed to the periphery of the network and so it makes sense so I will And so it make sense that when we look at algorithm performance that If the underlying dynamic is this, you know, the best guys are always answering, which is in fact what we see, right, they are not actually discriminating. Than things such as in degree, or kind of z scores of in degree, how much higher is the, is the degree than what you would expect, at to, you know if you had a balanced in out degree for the different nodes That should perform better. And so we kind of figured out what the underlying dynamics were by seeing which algorithms is performed well. On the other hand, in the just better simulation, if people really were, mindful or their bandwidth, and matching their expertise to the level of the question. Then, in fact, we would see here, page rank is in, yellow, that the page rank algorithm would be the one that is most, Most predictive. So I have not discussed the hits ranking algorithm, it is also kind of a web ranking algorithm but you know, it is just to illustrate that there are lots and lots of different centrality measures on hits. The basic definition is that you have hubs and authorities. Hubs are going to link to or point to authorities and authorities are pointed to, by good hubs. Right, so you have this kind of bimodel, Definition, I'm going to iterate, again, some, matrix, computations, and in this case, we can figure out that if you had the just better model a hits algorithm would not perform very well at all, because most of the questions are being asked by the newbies. They are the hubs, but their questions would be answered by the, you know, by the barely experts, right? So you would get the barely experts being ranked highest by the hits algorithm, and so we can. And sort of predict the hits algorithm would not work in this context versus the pay train algorithm whip. And so this is just to say that you have. Many, many different, choices of centrality measures. And sometimes it is nice to throw in the kitchen sink and see what works. If you were a social scientist you may try several different centrality measures. And you stick them into a big regression, where you are trying to figure out, you know, what specific outcomes do you care about. For example, if it is a person's well-being, because they are these are degree centrality matter more than their betweeness centrality. If you are looking at an online community who is going to stay active in the community? Does it matter how many contacts they have? Or does it matter that they are actually bridging different communities? For example, if they have high betweeness, what you might expect is that if one community, or one subset of people, starts being less active in that community, they would still have another that could potentially still be active. But on the other hand, if it is really high betweeness. They may just be bridging different, Different groups without really finding a home in, in that online community. So they might be more likely to quit. So the thing is, you don't know ahead of time. Just as with the Java forum, we, we expected paid rank to perform well and it, it did not actually perform any better. Is actually slightly worse than just looking at the Z scores of in-degree versus out-degree. And that is a little bit of the magic [laugh] of doing, especially exploratory session network analysis that You know, you can go in with the standard set of centrality measures. And then you really think about the problem. You really think, you know, how can I capture. For example, with expertise, maybe it is not so much how much you answer. But do you answer a lot more than you ask, right? It make-, it sort of makes sense, right? You are, you are sort of, [inaudible]. You know, maybe that is a better measure. And so. I mean, it does contribute to the problem of there being lots and lots of difference in childing measures. But I also think, practically speaking, it is good to, it is good to customize. And, you know, you can always use the standard ones, but sometimes, you know, if you are thinking deeply about your problem you, you may come up with something new and, and different and interesting. So in the next video I will be talking about how to fit power loss properly. I consider this video optional because, In a you know, it is going to be a little bit technical but, if you are ever going to claim that you have a scale three network you must watch this video and learn out how to fit power loss properly, because I do not want you running around making false statements after taking my class. Okay so I will see you maybe in the next video then.