Hi. Welcome back to Part 2 of Basics of Inference. Lecture module for the Caltech-JPL Virtual Summer School of Big Data Analytics. So in this section what we're going to talk about is the Central Limit Theorem Confidence Intervals. We'll introduce a Bayesian formalism for inference and then we'll take a step back and summarize a few things. The material on the Central Limit Theorem is based largely on Tom Ferguson's 1996 book called A Course in Large Sample Theory which is an enormously. Useful reference. You do have to know a little bit of math though, to appreciate it. So when last we spoke, at the end of the last module we talked about the difficulty of computing maximum likelihood estimates and their sampling distributions if you don't know what the underlying probability distribution is. That would seem to be. a, a show stopper for, for doing maximum likelihood. [SOUND] But there's a miracle out there. And the miracle is called The Central Limit Theorem, and it really is quite something if you stop and think about it. [SOUND] The Central Limit Theorem says that if y1 through yn are a sequence of iid random variables, like. That which constitutes a sample. Each with a same expected value and the same variance, both finite. Then the distribution of a random variable, s of n, which is essentially a sum. A rescaled sum of the y's. Tens to the standard normal, that's a Gaussian distribution, as the sample size goes to infinity. Or gets large. In other words. What I've shown down at the bottom there. If you take the cumulative distribution function of that random variable sn and you look at how it goes as n gets large, it converges to something that looks like that expression there in the middle of the bottom line. That's the Gaussian. The integral of the Gaussian PDF. And that really is quite a stunning fact if you think about it. There's. One wonders why that should be true. And in most statistics courses at least at the undergraduate level. No one ever actually does tell you why this is true or proves to you why it's true. And the reason is because it does require some more sophisticated mathematics based on something called moment generating functions. And you can find that proof if you're interested in that you can certainly find that in many places probably. Probably even in Wikipedia by now and it's really something to look at, it's quite, quite extraordinary. But the result can be written in any of the following three equivalent ways it seems the most natural way to write it is in terms of the sum random variables, the sum of the y's, where the y's are the random variables. And the notation there has a little arrow which says that the distribution of that object on the left converges, in distribution to a Gaussian distribution with the mean n times the individual random variable expected values and it's variance is n times the variance of the individual variables. [SOUND] that. Convergence in distribution thing could be a very long explanation. That is essentially convergence of functions, so if you've had a math class that talked about what it means for functions to converge, we're talking about distributions function, distributions functions converging. In the sense. There are other senses of the, of convergence and at a certain point I'm going to stop by bothering to distinguish between them because it's not important at this level. But do be aware that, that expression represents a particular kind of mathematical convergence. [SOUND] An equivalent formulation is the expression in the middle. [SOUND] This is probably the one you might be most familiar with. It says that the sample mean converges to a Gaussian distribution [SOUND] with the sa, with the mean, or whose expected value is the expected value that you seek. And whose variance goes down with the sample size. That may be the one that, that probably is most familiar. And the last one down there at the bottom. It's a little mysterious looking. But all it is, is the same version as on the line above, but centered in the sense of the thing in the parentheses is simply the deviation of y bar from the thing that it's targeting. And then we're blowing it up by a factor a square root of n. And the reason we want to do that is because that random variable will have a mean of zero and it will have a variance that looks like the variance of the original observations. And we're going to come back to using that kind of a random variable that transformation of a random variable a little bit later. And so it might be good to get used to looking at it right now. Let's see, that limiting distribution on the right is sometimes called the asymptotic distribution of the statistic. There is also a central limit theorem for independent but not identically distributed random variables. It's call the Lindeberg-Feller central limit theorem. And it has some extra conditions that have to be met. I'm not even going to talk about what that is. You can look that up in a book. Just know that it exists and you can go find it if you're actually dealing with a situation where you have a sample that is not iid, but it's just id, independent. There's a Central Limit Theorem for random vectors, not surprisingly, and it looks like the expression here. I'm showing only that centered and scaled version. And there is central limit theorem for functions of random variables and vectors which I'm showing down at the bottom and this is a consequence of the fact that. When I transform by the function G I have to account for that in the covariance, because my random variable is already centered, it doesn't impact the main. This theorem is ha, is called, at least in Tom Ferguson's book, it's called Cramer's Theorem. The Central Limit Theorem and Cramer's Theorem premieres are extremely useful because many estimators end up being functions of sums or averages of iid random variables or vectors. The sample variance for instance, is the function it is a function of two different. Of a transformation of a sum and using combining the Central Limit Theorem with Cramer's Theorem, we can obtain the fact that the centered and scaled sample variance has a Gaussian distribution with mean zero and variance. As I've shown it there. That mu sub y four is the fourth central moment of y. Okay the sample correlation coefficient for y variant random vectors actually obeys the central limit theorem too. There wasn't enough room on the page to show you what the asymptotic variance was. Of this, of this expression. Like everything else it can be looked up but it's too ugly to write down and the point is that the sample correlation coefficient does converge to something we know as the correlation gets large. I'm sorry as the sample size gets large. So other kinds of statistics for which the central limit there are holds are sample quantiles, rank statistics which are things like if I take a sample and I want to look at say the maximum or the minimum or the third largest element, those are statistics. They're called order statistics. Chi-squared statistics, remember that? Observe minus expected squared over expected. Extrema, maxima or minima and many others. So it holds in a surprisingly high number of cases. It also holds for dependence sequences of random variables. Which is maybe a little surprising. I was, I think I'll just take a minute here to tell you. What a stationary m-dependence sequence is. It's kind of a daunting name, but m-dependence simply refers to the fact that blocks of variables are sort of travel together. They move together. Independent sets separated by. It length, separated by m in indices are independent of each other, but they're dependent within. And a stationary distribution of a sequence means that, that distribution, that the mean and the variance of that distribution do not change as a function of your position with the, within the time series, within the sequence of random variables. So, let's not worry too much about that. Here's the CLT for dependent sequences of independent, of stationary independent sequences. And again, I, let's not dwell too much on math here. But everything on the page should look okay. In fact, if the y's were independent, instead of dependent. The second to last line in the equation would only have the very first term on the right side of the equal sign, because the variances would simply add if everything was independent. But if they're not independent, they're partially dependent out to some degree, then you have to cope with all the covariances that arise from that dependence. And let's just call that whole thing on the right hand side of the equal sign times squared. But we certainly can say, that the sample mean obtained from a stationary independent sequence, the centered and scaled version of that statistic, does converge to a Gaussian distribution, with the right mean and with a variance that we can know. That is assuming that. We can know all the terms in the middle parts of that point. Now, maximum likelihood estimates do have a really nice property which is that they're centered and scaled versions. They make if theta had is a maximum likelihood estimate, subtract off the target theta so now we are looking at the deviation. And then re-scale it by multiplying it by the square root of n. That thing converges in distribution to a Gaussian distribution with a mean of zero, which means unbiased. And a variance that we can figure out. the, the key is I'm going to, I'm going to, I'm going to rain on that parade in a second. But we can figure it out if we know what f is. That think on the right, i of theta is called the Fisher Information. And it is a measure of the information content, of the rand in the, in the random variable y, about the vector theta. Oh, I'm sorry, about the quantity theta, the parameter theta. So if you, we, let's define something that that expression at the top line of the equal sign in the middle, I'm sorry, the top line of the eq, first equation in the middle. I think that's the Greek letter psi, if I'm not mistaken, psi of y comma theta is called the score function. And it is simply the derivative of the log likelihood. And remember it was that derivative of the log likelihood that we set equal to zero and solved for to get the maximum likelihood estimate. So we've actually seen that thing before and if we let y be random in that expression, then the, what's called the Fisher Information is just the variance of the score. And this, this may be starting to go off into a bunch of words that have been packed together that don't mean anything to you anymore. But as we've seen a couple of times already the, the punch line of the story is that we can know what this is in principle. And in fact the inverse of the Fisher information is something called the Cramér–Rao lower bound which is the minimum variance that any estimator can have. So what this is telling us is that maximum likelihood estimates are what's called asymptotically optical which means if the sample size is large enough and you have the luxury of knowing what f is. So that you can compute these things, then you know, you're doing the best you can possibly do. But then there's the catch, right? Which is that in most practical situations, we may not know what f is, but we'll come back to that. Let's talk about confidence, let's do a quick detour into a confidence interval. Remember we had the sampling distribution of our statistic and the sampling distribution of our statistic basically quantifies everything we know about the behavior of the statistic if it were to have been calculated repeatedly over multiple random samples. So. Something like a maximum likelihood estimate I've now, I've, I've drawn the multiple samples in the cube with the green dots, I've written down all the different thetas I might have computed from each of those samples, and I've written I've just wrote down a particularly nasty looking probability density function. For theta hat here. Just to make the point that not everything in the world is Gaussian or even symmetric, and what we know, if we knew what the distribution was of at least the centered version of our statistic, which would be theta hat, which is computed from y, I'm now calling it theta hat of y instead of g of y. And I look at how it deviates from the true value that where trying to estimate theta. Then I know that with probability 0.95, by definition that random variable falls between the number l of 0.025 and the number u of 0.975 and if I know that then I can reverse engineer and expression that looks like. The first thing in the equation under the first bullet, which is designed to tell us something about how likely it is that our interval, computed from our sample y, actually contains the true value that we're after. And people are sometimes tempted to regard that as a statement about the probability of obtaining a particular value of theta, which is the true parameter. And if we're going to be good classical statisticians we can't, we don't want to say it that way. Turns out if we're Bayesians we're going to be allowed to say it that way and so we'll get to that in a second. But the statement, the only thing that's random in that probability statement is y, and therefore the upper and lower limits of the confidence interval. So, that's just a quick introduction there you probably remember that if our statistic follows a Gaussian distribution, with mean zero and variance one. [SOUND] Then that lower limits ends up being minus 1.96, which is what you'd look up in a table, or get r to tell you. And the upper limit ends up being plus 1.96. But this is a much more general concept [SOUND] than just the Gaussian situation. [SOUND] Okay. So, now let's talk about Bayes, Bayesian's again, Bayes Theorem again. You remember Baye's theorem from one of the probability modules. And I'd like to reiterate that Baye's theorem is just math and like any other math. Whether it's a good thing to do or a bad thing to do to invoke it depends on how confident you feel about the assumptions that are necessary to invoke it. So up to this point, we've been good frequentists. We've been good classical statisticians who think about probability as long run relative frequency. And we've treated everything having to do with, we've treated theta as a fixed but unknown quantity, a kind of an ordinary variable in the algebraic sense. And our inference was based on playing this sort of sensitivity game with theta. Saying how does the probability of observing the sample I got, which I call the likelihood. How does that change, as I sort of turn the knob on potential values of theta? So, there's nothing random about theta. There was only, sort of the sensitivity of the likelihood to different choices of theta. But if you're a Bayesian, you can go a little bit further. And. I'll, I'll try to tell a, a small story straight, sometimes I don't get the story straight and it ends up being embarrassing, but I'll take the risk. I would like to play a game and what I'm going to do is I'm going to flip a coin and before I flip the coin, I'm going to ask you what you think the probability is that the head, that heads will come up on the coin. And you'll probably tell me that it's a half and then I flip the coin and let's say it does come up a head. Well, the coin, the random variable has been realized and at that point the probability of it being a head is one, because it was a head. Now we'll play the game again, and what I'm going to do is I'm going to flip the coin, but I'm going to put my hand over it so you can't see it. And now I'm going to ask you what you think the probability is that the coin came up a head. And you'll probably tell me it's a half, again. Because. The fact is, that even though the coin is either a head or its not at that point, you still don't know. And you're willing to use probability to express the fact that you don't know. So, if you're okay with that, then you might be a Bayesian. In fact, you probably are a Bayesian. And if you feel that way then you probably also feel that it's okay to treat theta, that unknown parameter that we're interested in as a random variable instead of some fixed but unknown quantity. And rather than play the sensitivity game, We will use probability to describe the con, we only, we will use the conditional probability distribution of theta, given y to describe what we think the the behavior of random variable theta is. And in order to do that, we have to assert a marginal distribution. For theta, because if you remember back to the probability module, when we first wrote down Bayes theorem. You can see it over here on the right in grey letters. At the bottom of the page I rewrote it there. It's the probability of b given a times the probability of a over the probability of b. That's all Bayes theorem says. And if random variable let's say what I'm interested in, is random variable theta, the conditional distribution of theta given y. I don't know what that is, but maybe I have a much better idea, of what the conditional distribution of y, given theta is. So, if I'm drawing from a Gaussian distribution. And I don't know what the expected value is. I don't know what mu is. I don't know what mu what, I'm sorry, what's what y given, what theta given, I'm sorry, if I'm drawing from a Gaussian distribution. I don't, and I don't know what the expected value is, rather than simply twiddling with potential values of theta that maximize the likelihood, I'm going to actually use Bayes theorem to compute the, what's called the posterior distribution of mu given the sample that I got. And that was not particularly well done, but I'm just going to. Don't buy it. Let it be. So here's my cartoon of being a Bayesian and I'm doing Bayesian confidence intervals. So the I think the cartoon starts with a little graphic all the way over on the left. Which is what I'm going to postulate or what I'm going to declare to be what's called the prior distribution or the marginal distribution of theta. Now what I'm doing, remember what I'm doing is I'm basically assuming theta and y are jointly distributed. And therefore I can, I could get the I have this notion of a marginal distribution of theta. And I have a distribution of y given theta, because let's say I take a random draw from that distribution f sub theta, I get a particular value of theta and then I generate a bunch of potential samples from a population that has that value of theta. Well, if I'd had picked a different value to theta to start with, from the distribution, I would get a different set of samples, like I do in the se, say the second cube there. And for each potential realization of the random variable theta, I might get a different set of sample of green dots there. And would get a different little sampling distribution. Of theta. I'm sorry, of sampling distribution of theta, that's correct. So essentially what we're doing as Bayesian's is we're saying, we're just using conditional probability to say I'm going to average over all possible choices of theta, with the weights for that averaging being give by fs of theta. And I'm going to use that, to get this posterior distribution of theta, after I've seen the data y or the sample y. Just one little note is that, here's a place where we might definitely want to use a sufficient statistic, instead of working with the original sample, even though I've drawn the picture as if. As if I could work with the original sample here. If I'm a Bayesian, I can write down a Bayesian confidence interval, which will take the form of being a probability statement about theta. And by the way, it's also a probability statement about l and u, because those two things are functions of the observed data y. Okay, all right, so I'll make a few editorial comments now in the way of a summary, which is that it all boils down to how you want to model your unknown parameter. Is it a random variable, or is it a fixed but unknown value? And the, I guess the anecdote I like to give is we've probably all done the calculus problem. Where we're told that a bath tub is filling up at a certain rate because water is pouring in from the spigot, and it's draining out at a certain rate because the plug, the stopper isn't in properly, and we would like to know how long it'll take for the bath tub to overflow. And if that's really your bath tub. Then you darn well care whether the assumptions about how fast the water's coming in and how fast it's going out are right. And it's the same thing here. You know, you have to think about the consequences of making a wrong choice, or the consequences of getting a bad estimate in this case. As to whether you want to be a Bayesian or a Frequentist about how you're doing it. Do you think you have reliable information about theta, that you can bring to bear on the problem by asserting a prior using Bayes theorem? Or don't you, do you want to fall back to sort of the sensitivity version of the problem? So in my opinion, the Bayesian formalism is more complete, more flexible, and lends itself to conditional modeling, and therefore is actually. In a lot of cases at least in scientific applications, where we think we know something about how things work, a really good way to model things. But whether you are a Frequentist or Bayesian, you still have to know or assume things about the distribution in order to use any of these techniques. And that's where, you know, that's where I think that, that People become disappointed and disillusioned with formal statistics because that's, a lot of the time that's not true. And so in the next two modules what I'd like to do is talk to you about a couple new procedure that have come on the scene in the last, it's actually been 20 or 30 years, so they're not actually that new. That might help us out in the situation where we can't use the Central Limit Theorem to help us out with what the form of the density is, or what the form of the sampling distribution is. So that's where we'll go next Let's see. I did, I had these references here, I believe I've shown both of them before, or all three of them before, so. I'll leave you to explore those references on your own.