Okay, this is the, start of the second of our sub-topics in the lectures on Uncertainty and Inference, with a Big Data Analytics Virtual Summer School. Now we're going to talk a little about inference and hopefully things will now be migrating in a direction that seems a bit more practical or at least getting to more practical. So, here we're going to talk really about a, a very small subset of what you would really think about as the topic of inference and statistics. We're going to talk about some basic concepts and then the principle of maximum likelihood. And we'll use the principle of maximum likelihood, to illustrate the concept of the uncertainty of the estimate and talk about desirable properties of estimates, which as it turns out, maximum likelihood estimates tend to have, which is a handy thing. So, basic concepts of inference. Remember we talked about a process distribution which I might represent with a random variable x, having a, pdf or pmf. Let's talk about pdf's from now on, and you'll just understand that what I mean, is either pdf or pmf as appropriate. So the process distribution, is f of x. And let's say that what we are interested in, is in knowing something about what the expected value of that distribution is, what is the most typical outcome I could expect, if I took a random draw from that distribution. Let and we'll call that expected value mu sub x. So what would I do about this? I'm, I'm sure you have seen this before in statistics classes,you take a sample, and let's say, we're going to draw a sample of size capital N from the process distribution. And there's my cartoon picture of sampling there, which is supposed to look like a filter, and so, I'm going to let some of those realizations pass through. And I will obtain a sample size in which I have written now as Y1 to YN, over there on the right. So, note that the distribution I'm interested in, is a scalar, it's a distribution of a scalar random variable. But the sample that I obtain, is an n dimensional object, living in n dimensional space. The form of the probability distribution of the sample random vector is determined by the form of f, not surprisingly, that's where we got the observations from, or where we got the values of y from. But also by the sampling procedure that was used, by the nature of that screening process. I will just make a comment that that's actually very important for those of us that work in areas like remote sensing. Because going back to what I said earlier in the first module about observational studies and experiments we kind of have to live with the data that we get. And in remote sensing, the samples that we get are, what you might call the systematic samples, in the sense that, satellite orbits tend to repeat. So we tend to see things always at the same time of day, for example, if our satellite is in polar orbit. So in that case, you would not be able to say that, that sampling, that filter, that screen there was sort of a, kind of an unbiased representation of all possible outcomes that we could have gotten. So we do need to keep that in mind, but for the moment, let's ignore that. So our sample is identified as the random vector Y1 to Yn there. And it has a PDF of it's own. And if each draw of y from f sub x, was independent and from the same distribution which we call independent and identically distributed, IID for short. Then the joint pdf of all n of those random variables Y1 through YN, is simply the product of their individual probabilities. And that is what I have written down at the bottom of the page there. It might look a little funny but the joint probability of the vector little Y, which I'm, I'm now, being very careful now to have the argument to my function there be a lower case letter, because it's a realized value. Is whatever the pdf of capital Y1 is evaluated at the observed value of little y1 and so on, through that multiplication and I can write that sort of more compactly with the product operator at the lower [INAUDIBLE] on the left and all I've done is to absorb that each of the y's actually has the same distribution as x because that's where I drew the y's from [SOUND]. Okay, so let's look at a really, really cartoonishly simple example here. Let's pretend, I take a sample of size three, from my process distribution. So I've Y1, Y2, and Y3. And that's a random vector, it's a single point, in a three dimensional space. And it is shown on the right, there's a cube that's a three dimensional space and I've identified the sample I actually got with the green dot. So and the question is, what can learn about mu of x? Front the sample that we have, Y equals little y. So the PDF of the sample, actually is related to whatever mu sub x is. It's going to that, there, where the center of that distribution is going to have an effect on where in that three dimensional space, that green dot is. So it would seem, that a logical way of trying to get at what mu of x is, is to ask ourselves, what the probability is, of obtaining the y we actually obtained, for different values of mu of x. And that is the principle of maximum likelihood. Here what I'm trying to show, is I have three copies of the three-dimensional space, and I have the same sample, which is the location of the colored dots, in each of the three. because I got, I got the answer I got. But what I've done is I've color-coded it, according to the probability that I would get that answer, that I would get that sample under three different choices of mu of x. And then I've made the whole picture a lot more compact, by simply drawing those probabilities as a function of mu of x and the little tiny postage stamp sized graphic on the right whose labels you can't read, probably, the X axis in that plot is labeled mu x. The Y axis I call likelihood and the title is likelihood. so, if we think of the PDF of the sample viewed as a function of the parameter we care about mu of x that's called the likelihood function of mu of x. And instead of writing F sub Y of Y and mu of x which we could do. We're going to write it as L of mu of x of y to sort of highlight the importance of the choice of mu here. We're now, basically going to think of y as fixed at the value we got and we're going to twitter around with mu sub x, in order to try to maximise that probability. The maximum likelihood estimate maximizes L for the real life sample All right. So, you might have seen this example too. Another freakishly simple thing but it gets the point across. If the original distribution of X, that we care about, is Gaussian, with an unknown expected value, mu sub x and a variance equal to 1. Let's pretend we know that. Then we know what the form of f sub x, is. And we know what the form of f sub y is. Because, those Ys were drawn independently. So the joint distribution of the three Ys, is just the product of three copies, of the Gaussian distribution, with a mean at mu x and a variance of 1. I'm going to write that, as I did at the bottom to include mu X as a variable in, in the notation there, just to emphasize the dependence on mu. In cases where we know the form of f sub y, it's possible to solve for the maximum likelihood estimate analytically. By finding the derivatives of the likelihood function and setting them equal to zero. Okay? It's often easier to solve for the log of the likelihood, which is a fair thing to do, because the log of a likelihood will have a maximum same place as the raw likelihood does. And this is how you would do it. I probably don't have to go through this. This is why we end up with the sample mean, y bar, as the maximum likelihood estimate of the unknown parameter mu x. So, it's not just because, it seems like the right thing to do, it's because it maximizes the likelihood. Okay. So now let's talk about uncertainty. That was all, I think, fairly straightforward and fairly intuitively reasonable. The part that may be a less little intuitively reasonable is that the one sample we got, the green dot in the cube over there, was only one of many different things we could have got. And I've tried to indicate that, with all the other light green dots that I've put in that box too. Each of those, represents, a different sample that we could have gotten. And there is a mu hat of X, or a maximum likelihood estimate, that goes with each and every one of them. And it is the distribution over all those possible samples and their values, their corresponding values of mu hat sub X, that defines the uncertainty of our estimate. And that's all we're trying to get at in inference and with a quantification of uncertainty. So, let's redraw the picture, as I've done here. It's the same three-dimensional box, with the dark green dot and the light green dots on the left. And I've simply listed there in a color that I hope everyone can see. The different mu hat x's that we would get from from each of those light green samples. And if I made a histogram of those or better yet, I was able to draw the probability density function of those mu hat of x's, I might get something that we see on the right. And there would be desirable properties of the way we would want that PDF to look. In other words, the way we want our estimate theta hats of x to behave. We would want it to be unbiased, meaning that we would want the peak of that distribution, its expected value, to be, at the true value of mu sub x. And that's the unbasedness condition there, its saying the same thing. Which is to say that the expected value of the difference, between the estimate and the target, mu sub x would be 0. And we would also like that distribution to be very narrow around that true value. Or around whatever the expected value is, which is to simply say, that we want the variance of the statistic mu hut sub x to have, to be smaller than, the variance for any other estimate. And any other estimate which I've called Mu tilde here. And estimate that has both those properties is called minimum variance unbiased. M-V-U-E. Minimum variance unbiased estimate. And it's something that we normally strive for if not require when we invent ways of getting estimates. Now, it's true that we could choose any function of the sample, as our estimator of theta hash. I could choose something that has nothing to do with the sample at all, if I wanted to. I could choose any single theta hat, or any single value of y, of g of y if I wanted to. But we have to be careful because if estimators don't have those desirable properties, then we don't have any guarantees about the conclusions that we draw from them. So, typically, we choose g so that it's the sample version of the, thing we want but as I illustrated earlier, that's not just because it seems like the right thing to do, it's because of the legacy of the maximum likelihood maximum likelihood derivation. When the thing we care about is a moment of the original distribution. There's a fair chance that that's actually the right thing to do, but it may not be optimal in the sense of being MVUE. Here's a key point I think you can think of that statistic that you choose to compute, from the sample, as a form of data reduction. We've been using in our example, we've been using a sample of size three, which is a point in three dimensional space. Most samples are much larger than three. It would be, rather silly, to take a sample of size three in most cases. And so, suddenly we're in very high dimensional spaces and it's not really possible to draw pictures that way or to even operate, with those things. So we're looking for good forms of dimension reduction to apply to our sample in order to get good estimators of the things we care about. And there is a concept out there called statistical sufficiency which says that if the theta hatch statistic we choose the g function we choose to apply. Contains all the information about theta that is contained in the original Y then theta have is called the sufficient statistic for theta. And there is a formal probabilistic definition in terms of the PDFs of the various densities involved, but I'm not going to put that here, is easy to look that up. Now it may be hard, to solve maximum likelihood equations. Particularly if you don't know what F is and even if you do, well of course you can't solve the maximum likelihood equation if you don't know what F is. Even if you do know what F is, if F is non-Gaussian or some other very nicely behaved distribution. It can be very hard to solve maximum likelihood. If the the y's are not independent and identically distributed, it can be hard to solve for maximum likelihood because then you have a hard time saying what the join distribution is. So I already said if N is greater than three it can be hard. And it can be hard if we're talking about arbitrary process distribution parameters of interest. Not simple things like means or other moments. So, that's the end of this module and in the next module, we're going to look at some things that you might be able to do in the absence of knowing what f is. Here are a couple of references that I like, a couple of textbooks. The top on is actually an upper division undergraduate-level text. The bottom one is a first-year graduate text. They're both quite good. And I think we said that so, let's move on.