Hi and welcome back to the JPL-Caltech Virtual Summer School and Big Data Analytics. I'm Amy Braverman from the Jet Propulsion Laboratory, and in this module I'd like to talk to you about the Bootstrap, which is a statistical, technique for [SOUND] estimating or approximating sampling distributions of complex statistics. [SOUND] As you might recall from the previous module, whether you're a Frequentist or a Bayesian, you need to know something about the sampling distribution of the statistic that you're interested in. And how it depends on, the true or realized value of the population parameter that you're trying to estimate, or learn something about. and, we've looked at a bunch of [SOUND] analytical methods in the last module. Which we can use when we know the underlying probability density function of the process. We call that the process distribution. Or when we can invoke the Central Limit Theorem and we have a very large sample. [SOUND] however, those things are, in the real world, sometimes hard to have at our disposal. We sometimes don't know, whether we can use the Central Limit Theorem. We sometimes don't know whether what the form of our underlying probability process distribution is, or the sampling distribution of the statistic we're interested in either for that matter. So the question is can we get information. About the distribution of our statistics and how it might depend on the true perimeter we care about without making those assumptions and I want to apologized for the abuse of notation that, is embodied in my little expression for f sub Theta hat. Bar capital Theta there. I said way back at the beginning how important notation is, and that we would represent random variables with capital letters and their realized values with lower case letters and I would seemed to have violated that by using Theta hat for both purposes. In that expression. But, Theta hat is so traditionally the, notation used for a sample's statistics that I can't tear myself away from it. So, in this module we're going to look at a technique called Bootstrap. In fact, two particular versions of the Bootstrap, I would like to say that. I made, great use of some notes from Professor Aniron, Anirban Gupta, DasGupta at Purdue University in the Department of Statistics. And Professor Charles Geyer in the School of Statistics at the University of Minnesota. And a, book that I'll give you a reference to, called Politis, Romano, and Wolf. 1999's book's subsampling. So the Bootstrap, the ordinary Bootstrap is quite old by now. I think it was first, popularized in about 1979 by Brad Efron and Rob Tibshirani from Stanford. And they were, looking for a way to, solve this problem of how can we learn about the sampling distribution of a statistic when we don't want to make assumptions? This is what's called a non-parametric method. Parametric referring to having to make assumptions about parameters, or about the form of a distribution function, or a density function, or a mass function. [SOUND] so. [SOUND] Let's say we have an iid sample, Y1 through YN, as we had before, from a distribution F. [SOUND] Earlier we called that distribution F sub X, but I'm just going to call it F now. [SOUND] And we have a statistic, which as before, is a function of the sample. And we want to know its sampling distribution. First thing to notice is that the sample itself actually defines a distribution. We call that the empirical distribution function of the sample. That puts mass 1 over N on each of the realized values that belong to the sample. So let's call this, empirical distribution function F hat sub N. And I'm careful to put the N on there because that distribution changes as I have more, and more elements of the sample. So, the graphic on the left is just a plot of an F hat, sub N for a simple example where I drew a sample of size five. I think I got these, I might have gotten these from Gaussian I'm not sure. So there they are, there's five numbers there and you'll recall that for discrete distributions the cumulative distribution function is a step function as it is here, because essentially what I'm going to do is I'm going to sample from the sample. And therefore, the only values I can get are these five values. So, [SOUND] let's define a resample, as a sample of size N, which was the original sample size, drawn from F hat N with replacement. And let's call those realizations Y star 1 through Y star N are those random variables. So notice I am sampling, with replacement which means every time I draw something I put it back so I can get the same thing twice and I'm going to draw a sample of the same size as the original sample, I call that a resample. On the resample I'm going to compute the statistic Theta hat. And I'm going to call that Theta hat star because I used the value Y1 star through Y1 star to obtain it. I'm going to do that, B times where B is a large number, and you may ask me how big B is and there's actually a longer answer to that than you, than you might, think but we'll get to that. And we will estimate the cumulative distribution function of our statistic, Theta hat. Which is computed from the original sample. By a cumulative distribution function I'm calling H Boot. So now I've switched from F to H. I've kept the N in there to show that it's going to be different depending on how large the original sample is. And I subscripted it with Boot, to indicate it's a Bootstrap distribution. And all it is, is the, probability that a random draw from the sample itself, that's Theta star. Realizes a value that's less than or equal to t. And I'm going to approximate that with a Monte Carlo estimate where after doing this capital B times I simply count the number of my Theta star B's that were less than two. I add them up and I divide by that big num, that big B. And there you have it, it's a Monte Carlo estimate of that probability for that particular value of t. You may or may not have seen the notation one, with an argument there, that's called an indicator function, and it simply takes on the value one if, it's true and zero if it's not. Okay, so here's a simple example, here's my original sample that's copied from the previous page, and let's say that I'm interested in the median. I'm interested in the median and by the way, that's a particularly nasty choice and in the examples that I show you, we're going to see just how well this doesn't work as well as how well it does work for a choice of, of, of something that's a little complicated like the median. If I showed you this for the mean. It would all come out beautifully, and we'd all go home and say that's nice, but wonder what we might have learned. So, the true median of this sample is 2.82. And then I, draw B additional resamples where you can see that now, in some cases, I have gotten the same value more than once. And on each resample I compute the median. And the, PDF or PMF of g is going to be approximated by the histogram of those Theta hat stars, as the CDF is going to be approximated by that Monte Carlo average. [SOUND] Okay. So, now we're at the really important part. We really want the true CDF of Theta hat. But we only have one value of Theta hat, namely the one computed from the original sample. We don't know what the true CDF is. Let's call it H sub N of t. We can, by this, by this procedure that we just described, compute what H Boot sub N is, in the way that we described. And our hope is that H Boot is close to H, because if we're going to make inferences based on H Boot and claim they apply to Theta hat. Then that would have to be true. So, the question is, is that true? And to, get after that point, we might think about the following things. There are really sort of two main sources of error in making that, approximation. Probably the really important one is treating those resample draws. As if they were draws from the original population rather than draws from the empirical distribution of the one sample we actually obtained. The second source of error is the Monte Carlo approximation of the probability. But that can be driven down by making B sufficiently large. It's the other one that's really important. Okay, so Now we do know that, F hat the empirical distribution function of a sample of size N, converges to, the true distribution function F as the sample size gets large. In other worse, as N goes to infinity. And I specifically didn't put a little d over top of that, although I probably could have. But I believe that we want to be a little more careful there so I'm just going to leave that kind of vague and loose. And refer you to books to determine just what sense that convergence is in for the purposes of the rest of this lecture your common sense is as good as any. So if that's true then the wise star ends are getting more and more like the original Y ends so that's good. Does that alone guarantee, that the approximation of HN by H Boot is good enough? Not really, no. Unfortunately not. It depends on a bunch of things. It depends on the nature of that transformation g, in other words, on the form of Theta hat. And on some of the underlying properties of the true distribution function F. The main point I want to make here, we'll go into that last point in a minute, but the main point is, when we say the Bootstrap works, this is what we mean. We mean that H Boot there, is what's called a statistically consistent estimate. Of H sub N. In other words the Bootstrap distribution, converges for N going to infinity to the true distribution so, hopefully, that comes as a shock to you. Because to say something works doesn't mean it always does work. In this sets, that sounds, sounds like a, contradiction. But, wh, when we say the Bootstrap or any of these methods work, what me mean, in fact, for all that matter we mean, the central limit theorem, they work because they converge the right thing as N gets large. We haven't said how large N needs to be. So we need to keep that in mind that if we do an experiment with a finite N we have to worry a little bit about how big that N is and whether it's big enough. Okay so now I'm going to, going to finish off this module with a bunch of examples here that hopefully I can lead you through in a clear way. So the example I want to talk about. Is estimating the median from a sample of size 3,000. For the the median value of a Chi-squared distribution. So what's shown on the left here is a Chi-squared distribution. I used R to simply plot this. And the true median of a Chi-squared distribution with three degrees of freedom, which is what I used here so that it would be, pleasingly skewed is 2.366. If I take one sample of 3000, from a Chi-squared distribution where everything is independent, I've made a histogram of what's, of what's in that sample. And I've shown you where the median is and so for that one sample that I took, obviously the median from the sample, is not the same as the median of the true distribution. So I think we all already knew that, but I wanted to make that point, clear. So simply assuming that 2.415 is. The true median would not be a good idea. So now, lets play another game that we can play because we know what the true distribution is. What is the true distribution of the sample median? Well if I take 10,000 samples of size 3,000 from that Chi-square distribution with three degrees of freedom. And I compute the median for each sample and I make a histogram. I'll plot it and I get what's on the left there. And I put a Kernel density smoother on there to make it look a little smoother. And I'm going to call that the true distribution of the sample median for the purposes of what we're talking about here because, 10,000 is a pretty big number and that's probably fairly close. So what I would like is, that Theta hat is centered at the right value 2.365 which it appears to be because I believe that was the value from the earlier page, and that it has small variance. And it certainly does appear to have small variance here. That's nice. Now it turns out, that, it's actually easier to look at this centered and scaled version of the sample median, rather than the raw version. And principally, that's because what I'd really like to see straight away, is that, that distribution is centered at the right value. And that it has a small variance and that variance number, the way it's, shown there on the left, is so small that I might not be able to, see it well enough, of course I could have used more decimal spaces. But when I start comparing this to other ways of estimating the sampling distribution of the sample median it's going to be more convenient to look at this, centered and scaled version of the statistic. Which I've plotted on the right and you can see the distribution pretty much looks exactly the same except now it appears to be centered at nearly zero, not quite exactly zero, but nearly zero. And it has a variance of 7.118 which doesn't really mean a whole lot right at the moment, but it will in a minute. So, this is the distribution that I assert we, this is the target. If we could reproduce this distribution then we'd be happy. Okay so now let's look at the ordinary Bootstrap. If I take 3,000 sample of size 3,000 and compute a sample median for each one. From the, I'm sorry. From the one sample that I have, we take the one sample that we have and we take 3,000 resamples and compute the median for each one. Make a histogram of that distribution. What we get is the thing on the left and it's centered and scaled version's the thing on the right. Now, one thing you might notice is that in the centered and scaled version on the right. What I really would have liked to have had there, where I have Theta hat is q50, q.50. But I don't know what q.50 is, so I am actually approximating the centered and scaled distribu, Bootstrap distribution of the sample median. With something that's not quite exactly it because I substituted Theta hat. The value of the sampling median for the sample that I actually got, in in place of q50, but hopefully this is a fairly good approximation, at least in terms of the variability and the bias of the sampling distribution of the sample median under the Bootstrap. okay, so if we look over on the right. We can see that this distribution is centered at 0.257 which is rather far off of zero, or so it would seem, and it has a variance of 9.063. So we could say that the performance of the Bootstrap, in this case, I don't know whether you think of that as being not so good or. Good it depends what, sort of what you need for your application. But let's just remember these numbers because this does appear to be biased and the variance is too high, by a factor of two out of seven. Okay, so. Not great but Better than nothing, I think we can say that. The Bootstrap does, work in a, in the sense of being statistically consistent where things are, nice and you've probably heard nice in your math classes before. Nice in this case means we're working with a nicely behaved. Transformation g, or form of Theta hat like the mean. And if the true process distribution has a finite variance, or has some other, nice features about it and I've listed some of them here. now, you know, we may want to do this for things like sample skewness and of the sample correlation coefficient. Without knowing what F is, how we would ever know that is have four, or six, or eight finite moments and a positive variance I'm not sure. For positive variance it's probably okay. We can probably look at our sample, and get a feel for that. But these other things sort of take us back to. The problem we started with, which is if we don't need these, if we don't know these things, then we might be playing with fire here if we draw important conclusions from it. I will point out the last one, which is that sample quantiles of which the median is one, this does tend to work for the sample, quantiles but n has to be really big, possible much, much bigger than 3,000. So if you weren't, depressed enough by that already, then here's some more things to be depressed about the Bootstrap. Which is here's when it doesn't work. These badly behaved. Transformations or sample statistic calculations where, things aren't continuous or differentiable we said that it does tend to work for the quantiles but N has to be very large because the, the median is not a continuous function and N has to be very, very big before we can overcome that. Things like sampling on the uniform distribution, where the, the upper or lower limit of that uniform distribution is the parameter we care about, that's tough. And rather unfortunately, the, things go wrong and the same sorts of situations where the central endotherm goes wrong. Now what can we do about it? There, there has been a new development relatively recently, meaning in the last 20 years, where we have something new called the M out of N Bootstrap which can fix some of these problems for some situations. And I think, we're going to look at that now. The M out of N Bootstrap is, just like the ordinary Bootstrap. Except that we are not going to resample N things with replacement, where N was the original sample size. We're going to sample fewer than N things. We're going to sample M things. And the condition is, that M needs to be big but it can't be too big. It needs to, the ration M over N needs to go to zero as both M and N go to infinity. And other than that, there's not a lot of guidance as to, what M should be. So that's the main question here, is how do you choose it? Rules of thumb, that obey this rule are things like N to the half and N to the two-thirds. I've seen one suggestion that 2 times N to the half power, 2 times the square root of N is a good number. I, played around in the examples that I'm going to show you a little bit and I was certainly able to find a bunch of M's that didn't work. I found one, that sort of did work if you recall, I didn't emphasize it but the, sampling distribution, in other words the form of this histogram and the density curve over plotted on it. The ordinary Bootstrap was pretty ugly looking it seemed to have a couple of modes and it really didn't look very pleasing. Here are the results for the M out of N Bootstrap I think I'll direct your attention immediately to the right side of the page because that's the one we can easily compare. And what we'll see is that while that second mode hasn't been eliminated it has been somewhat reduced. So the distribution is looking a little more symmetric. Which is a good thing, but maybe not quite symmetric enough. The bias has been reduced to 0.111 here and the variance has been reduced considerably in fact, the variance is now getting really. A bit closer to the 7.11, I think it was, or, I believe it was about 7.11 was the variance of the true sampling distribution of the median. So that's something of an improvement, at least in terms of the variance, and the bias. Now I will say that that's just a cartoon example, there are a zillion Bootstrap. Methods out there, designed for specific situations and with specific kinds of improvements. And here's a list of just a few of them. probably, the one I'll direct your attention to are these two, two top items on the right side. The block Bootstraps, which are Bootstraps for dependent data. Where essentially instead of drawing, individual items with replacement from the original sample to create a resample, you draw things in blocks. So you're attempting to preserve, some amount of dependence within block by drawing whole blocks at a time. But of course, the catch is there you have to know what block size to use, and that is not an easy problem either. So you can check out in any number of places. You can, search on these terms, and you'll find a lot of information about these. Here are, sort of, three books on the subject. An Introduction to the Bootstrap by Brad Efron and Ron Tibshirani. [SOUND] That was almost the original book. That's a good place to learn things. I believe the methods that are discussed in there are the basis of an R package. It's either Boot or Bootstrap. There's a more sophisticated book, called Resampling Methods for Dependent Data which discusses the problem of using the Bootstrap on depended data. It's quite mathematical. And the subsampling book which we're going to talk about in the next module, we're going to talk about that technique. It's called subsampling. That's quite highly mathematical but truth be told, I found the discussion in the first chapter. To be really interesting it described the Bootstrap and then described this other method called subsampling and contrasted them and I thought was pretty useful. So now we'll look at this method called subsampling and see how we can do, subsampling is related to the amount of in Bootstraps. We'll how we can do.