So this brings us now to the notion of spam farming, right? Where the idea is, what I kind of alluded to where this t-shirt seller creates a fake set of web pages that they all link to his own web page, and all these web pages in the anchor text says that the target page is about movies. And this is what is known as spam farming. So, Google versus spammers, round two. So right, once the Google became a dominant search engine, really spammers become, become trying to figure out ways how to trick the Google search results. And what they did is they created what is called spam farms, where basically the idea is that you want to concentrate and collect the page link, and kind of funnel it towards a single target page. And there are, there is many kinds of web spam and many kinds of link, link spam. And in particular, you, many, many times I am sure you have visited a webpage where, where you see, where you come and see things like this, where, that basically just have a set of hyperlinks to some other web page. And the idea is exactly as I mentioned, that these pages basically funnel their page rank score importance to the, to the high value target web pages. So the way we can think about this now, is that we want to manipulate the structure of the, of the web graph in order to create new links in such a way that given web pages will get high importance. Conceptually, we can take the web and split the web into three types of web pages. We can call we can call them these classes based on the spammer's viewpoint. So for example, inaccessible web pages are basically pages that the spammer cannot touch. So these are pages on the rest of the web that spammer cannot touch. Then we have a notion of accessible pages. These are basically pages that the spammer can touch. So for example, spammer can add fake blog comments, spammer can add fake posts to, to various types of pages. And all these posts would kind of point to the target page. And of course, the spammer has also its own set of pages. We call these the pages the web spammer owns. And these are completely controlled by the spammer and, you know, may spam multiple domain names. There may be millions of these pages, and so on. So now the question is, what can the spammer do? And the spammers' goal, right, is really to maximize the page rank score of a given page t. Right, so there is this this t-shirt selling webpage, let’s call it page t, that the web spammer wants to improve its page rank score. So the technique the web spammer will use is that it will get many links from accessible pages pointing to the target page t. And this way, they will create what is called a link farm, such that the, all the page rank importances for these pages kind of funnel their importance back to the, to our target page t. One possible strategy for a link farm is created here. So basically this is a topology of how a link farm may be organized, right? So the blue, the blue cloud shows the inaccessible part of the web. Then, this inaccessible part of the web has in and out links to and from the accessible part of the web. What the web spammer can do, make, make this accessible part of the web, as I said before, these are kind of blog, blog posts and things like that. They can create comments that, that link into the target webpage t. And then, what the web spammer can also do, they can take these webpages that they own, and they can make all these web pages both point to the target page t, and the target page t can point back to these pages. And the idea is that there is, the number of these pages is huge. We will call the number of these pages to be m. And think of them as millions of farm web pages, because they are very cheap to, to create. So now, this is actually one of the most common and most effective link farm topologies. Now let's start to compute, and let's try to con, convince ourselves, what is the PageRank score of Node D? So basically the page we want to boost its PageRank score. So, to do this we will do the following. Let's, let's use the, the symbol x to denote the PageRank score contributed by all the accessible pages, here denoted as blue node, and how much PageRank they contribute to t. And let's use the Y to be the, to be the PageRank score of node t. So now, first thing we want to compute is what is the PageRank score of every of the farm pages that the web spammer owns? That is very easy to compute. We know what is the score of node t. And it is only the node t that links to the, to the farm pages. So the node t takes its PageRank score y, divides it evenly among all the M pages that are owned by the spammer. And gives them data fraction of, of its PageRank score to each one of them. And then, of course, each of these owned spam pages also receives a fraction of the score due to the random jumps. The random jumps happen with probability 1 minus beta, and there is N pages in total on the web, so that is 1 minus beta over N. So now, given that we now know what is the score of every web page that we own, this is denoted as the red nodes. Let's also compute what is the value of y. So the value of y, y is the page rank score of node t, is simply x, which is the amount of page rank contributed by the accessible pages, plus beta times M. And now the the pages the contribution of page rank scores from the pages that we own. So this is beta times y divided by M plus 1 minus beta plus N plus 1 minus beta divided by N. Where M now is the number of pages that the spammer owns and N is the number of pages that are total on the web. Okay? So if you think about this and multiply with beta M and solve the system. What we, what we get is that y equals x plus beta squared y plus some constant terms. All right? And what we will do, is we will take this last term 1 minus beta over N, this is very small because the web is huge, and is large, so we will ignore this one. So let's keep looking at what we get. What we basically get, is we get something that is like y equals x over 1 minus beta squared plus M divided by N, plus some constant, where this constant is beta over beta plus 1. Okay? N is the number of pages that the spammer owns, and N is the number of pages that are on the web. So what this means is that the page rank score of our target page t, equals basically the amount of page rank score that comes from the accessible part of the web, plus the ratio of M to N multiplied by some constant. So what does this mean? It's basically that the more web pages the web spammer owns, the bigger the M, the higher the score of the target page y will be. So, in some sense, spammer can create arbitrary large number of pages. So M can be arbitrarily large, which means that the page rank score of the target page t can also get arbitrarily large. And of course in reality, N is huge, right, the size of the web graph is huge. And M doesn't need to be that large because all we need to do is we need to boost the, the score of the target page t not to be the most important page on the web, but to be more important than, than, some other pages on the web. So even not, not too big link farms can already have a big effect, right? So, now this is basically the problem, and the question is, how do we now go and detect such links, link farms on the web?