Okay, so I guess this slide is this I guess you can say this slide but I guess this is true. That no all this models on the left side right there are a multiple of those right. So many of those are in fact special cases of Markov logic in the sense that you can represent them, right. Markov networks, Markov random fields, when I say Bayesian networks what I mean is that essentially it's as good as converting it into a Bayesian network with a Markov network and then representing as Markov logic network. Of course, you lose some independences. >> You lose some? >> Yeah, yeah. So it is, in the sense that, any distribution which is in the network, you could present here. But you lose, yes. [CROSSTALK]. >> You will produce things which. >> Yes, yes. >> Not necessarily [INAUDIBLE]. >> Dependent, right, yes. So you will lose some independencies, that is true, right. So, but [UNKNOWN], right? So which essentially a special case in that restricted sense, at least in the case of Bayesian network. But Markov network it is obviously true that, you can just have propositional sort of right, [UNKNOWN] Markov network, yeah, right. So, right, and in principle what happens that if you make [UNKNOWN] which means you don't have variables then this propositional model becomes essentially enigmatic, right. But what is more interesting is that you have this explicit notion of interdependence and [UNKNOWN] structure, which, which many of these models [UNKNOWN] right. So, that is, that is what you have, What is even more interesting is contradiction to first-order logic. When it's a first-order logic, at least here, when, an infinite weight first-order logic. Right? So, in the limit of all the weights sending to infinity. In the same limit, right? So, meaning if your Markov works, if all of them together turn to infinity, then you can show that. This exactly becomes first-order logic. Right, so all the weights which are satisfied, will get the probability of non-zero probability. And all the weights that are not satisfied will get the zero probability. >> When you say [INAUDIBLE] does it mean that you can replace logical and pyramid with that? >> Logical entailment with infinite case. And when you take the limit of the [UNKNOWN] probability, yes. >> A implies B which implies C. >> Yes. >> That follows from [UNKNOWN]. >> That follows from nine friends, if you take the limit of [UNKNOWN] yes. In the finite case. >> [INAUDIBLE] >> Meaning your network should be finite, finite number of constants. Yes, yes, yes, yes, but that should not be too difficult to see. >> That's, that's very straight forward. >> Yes, that's fairly straight forward. Yeah, yeah. So, all of them, all the decided weights will get the same probability, so their key assignments, each one will get one byte key. And others you'll get zero, it is for W will give you, right, because even if there is one difference, one formula which is not satisfied, then the denominator will really overpower the numerator and get to zero. Right, but what is more interesting is that if you are satisfied, if you have a satisfiable knowledge base, right. So your knowledge base is something which can be satisfied and if positive weights, then those assignments, which satisfy all the formulas are the modes of distribution. Right. So remember I talked about this sort of [UNKNOWN] probability. What you're saying is that, if your knowledge base in fact has an assignment which can satisfy all the formulas, those are the assignments which will get the maximum probability during distribution. So which is very, very good, right which is a very nice probability. In particular Markov logic allows contradictions between formulas, which is not the case for pure first-order logic. That's what we're looking for, okay. >> [INAUDIBLE]. >> Mm-hm. >> [INAUDIBLE] >> Mm-hm. >> [INAUDIBLE] >> Mm-hm. >> [INAUDIBLE]. >> no. >> Because like, smoking implies cancer. >> Uh-huh. >> But they're not talking anything about cancer implies smoking. So this [CROSSTALK]. >> So what you're saying not cancer implies not smoking? >> Right, so actually, so you're saying that [UNKNOWN] right, so first-order logic doesn't have any emotional directionality, right? I mean, I mean it is, it is [INAUDIBLE] or cancer, right, yeah. >> [CROSSTALK]. >> Yeah, yeah, yeah, so you saw, in this case it is little bit misleading the example that I gave that was more for orientation, but they are purely logical formulas. We are not talking about causality yet, right. So there is no causal semantics right. Although that intuition may help sometimes in terms of the weights that we get, but the mathematically its not causing, okay. Okay. So, how are we doing on time? How much I mean? >> 15 minutes. >> 15 minutes, okay. >> May be skip it [UNKNOWN]. >> Okay. >> And go to an example. >> Example, sure, yeah. So let me just, what we, what I will do is that I will define the Inference and Learning tasks. I will not go in detail of that, there is some interest in that then we can talk off later. And then I will give you some example that will help I think. Right. So just couple of slides. [COUGH] So the Inference problem is, I think we talked about that right? Given some nodes in the network right? So green nodes are given, and lets say you know the weights. So then you find the margin of probability of the nodes on the network, or what is the most likely set up of the network, right. So in this case the green nodes are given and you're saying, what is the probability Ana smokes, or Bob smokes, or Bob has cancer. Right, so that is the inference problem, and, right. So, this is called marginal interference results, or something called MPE: Most Probable Explanation, right. So, that, deals with the question of, what is the most likely state of the, the nodes together? And, that can be different from the marginals, maximum marginal state, okay. So, that, is a different interference problem. And, right, so I think we talked about that, that. Okay this is fine. So let me just play one or two more slides. Marginal inference, what you do is just put the, put the formula the, state into the equation of the, of the problem distribution. One little thing is that, typically you're given some evidence, right? So, given that evidence, you fix those nodes to evidence, and then marginal inference is find the probability of Y which are your query right? So then z depends on x. Because you fixed those evidence. And rest of the things remain the same, and this is the distribution then you can find the probability of y in principle. And doing this exactly is exponential in number of states of y, right, number of variables y. So typically you have to resort to other approximate sometimes. And I'll skip that page, so I'll skip that, that slides are propagation and there are some interesting work here where we can find clusters of nodes which we have similarly, but, but let me skip that. So this is timing. Okay. [BLANK_AUDIO] Right, so now, learning parameters. So learning parameters is, that given this formulas, right. So the formulas typically some domain expert can give you, right. So given the formulas, what are the rates that you would like to learn. Right. So let's say you, we assume that you have the data. Right. Let's assume full observability. So example, there are three constraints. So the data is in the form of an additional database, right? So you'll say that in my domain Ana smokes, Bob smokes, Ana has cancer, Bob has cancer, and the other friends relationship. So, given this, and you assume that the, it's a closed world assumption, that anything not in the database is false, then from that, essentially you use the principle of maximum likelihood, right. So, you know what is probability of this particular assignment of things with the variables in your network. So find the weights which maximizes getting this weight. Right? And that is the maximum like-node principle, and you can follow the standard application [UNKNOWN] and get, get the rates, right. So that is [UNKNOWN]. >> [INAUDIBLE]. >> Is there any [UNKNOWN] here? Right? So, what you can do is that you can start with some prior, and that will give you some regular addition. Right? So you can start with some prior [UNKNOWN], yeah. And, in fact, there is some work there, if you're, do, learn regular [INAUDIBLE] that was the [INAUDIBLE] actually go towards here. Yeah. Right. So the, the, so, Daniel Lowd and Pedro Domingos had a paper at ECML, and I think that compares many, methods for learning. Second order methods also, so you can get that, if you're interested. So the learning structure corresponds to learning the formulas, right. So, let's say, if the formulas are not given to you, then how do you learn those? And standard ILP based techniques can best be used, right. Now it's problematic so we have to relate your likelihood to see how good those are, and there are many different techniques, many more advances techniques which have been proposed. I will not go in details of that. And [UNKNOWN] has done lot of work on that. Right so he has a couple of three or four papers at ICML, a succession of papers which really improves those, those learning structure techniques. But, but let me just give you one, one thing. One idea that, you can either learn the network from the outside or [UNKNOWN] says that I have this formulas. And then are there more formulas that can be learned? Right, for example, given this formula as you can see the other formula that can be learned, and which is friends x, y and friends y, x they are essentially, this is a similar calculation. Right, so from the data you can potentially learn this formula. So what do you typically do for your real problem is that maybe you can start with a formula that a domain expert gives, and then you refine those formulas, based on the data, right so that's a typical setting. If you are very confident that you know the formulas, then you can directly [UNKNOWN] right. So that is the system. Okay, right, so I should mention that one of the reasons that it has become so popular and so many people have really adopted this for their own problems, is that there is software called Alchemy. And in fact there are at least four or five different implementations which have come about. All of them, I think, are freely available online. [INAUDIBLE] implement Markov logic, and they give various inference algorithms, various learning algorithms, and many different features. So you can just, after this talk, if you're interested, you can just go back and download that, and it should work, you know, as, just fine. And, of course, when you really want to make it work for your own problem, there may be issues, right. Because it's just like any other, any other subsystem. And one key issue that I should mention is that, it is very easy to blow up your network. Right, because you can write two or three formulas and if it is very high then you can sort of [INAUDIBLE] which is very, very big. So that is something that you do very careful when you write the formulas, not to blow up your network. And there are techniques that deal with that explicitly, in terms of inference engine, but it is good to be, you know, pragmatic and doubt such right formulas, which clearly do not lead to not very big network, right, very, yeah, so that maybe a, one challenge that, or something that is sort of an odd, as you use this system.