Welcome to the very last module in the Inference and Uncertainty section of the JPL Caltech Virtual Summer School on Big Data Analytics. This section is about a technique called Subsampling. And I'll remind you with this little outline slide. Where we are in the discussion. And this is the same introductory or outline slide that we had at the beginning of the last section, in which we discussed the bootstrap. And both the bootstrap and subsampling methods are nonparametric techniques for approximating or estimating or understanding the sampling distribution of a statistic that we want to use to estimate a process distribution parameter of interest. When we don't know the analytical form or the parametric form of the underlying probability distributions that gave rise to the sample. And the, that is often the case in a lot of applied works so these methods are not surprisingly very popular and perhaps potentially very useful for some of the things that we're doing here. I will not repeat the asterisk at the bottom, which is an apology for using some notation. I did that in the last module so we'll just live with that. So this section is reasonably short. We're just going to introduce subsampling and then I'm going to make a few comments and give you a few cautions about using any of these methods. The material in this section is based substantially on notes provided by Professor DasGupta in the Department of Statistics at Purdue University and Professor Charlie Geyer at the School of Statistics at the University of Minnesota and on the book by Politis, Romano and Wolf called Subsampling, which I'll give you the full reference to at the end. So remember in the bootstrap we talked about drawing resamples from the one sample we actually have. Now we're going to do something slightly different called drawing a subsample from the one sample we actually have. And a subsample is a series of draws from the one sample we actually have without replacement. A resample was a draw with replacement, and it was either a draw of the same size as the original sample, namely capital N, in the case of the ordinary bootstrap or a draw of some smaller number of observations which we called m. But in either case in the bootstrap it was with replacement and now we're drawing without replacement and we're only drawing m. So let's call the subsample v sub one to v sub m and the crucial difference here may not seem obvious is that when we drew the resample with replacement, we were not drawing from the original distribution f, we were drawing from the empirical distribution of the sample f hats of n. Here we are drawing from the true distribution f. And if you think about that, you think you're drawing a subsample of a sample. If you had just gone out this morning to draw the subsample directly, or a sample of the size of the subsample directly from the population, you could've gotten this very collection. And you may be thinking to yourself, but wait a minute. The only values I can get are the ones that were in the original sample so how can that possibly be true? Well, the answer to that is that everything we're doing here is treating the sample itself as random, so all the probabilistic statements that we're making are implicitly averaging over all possible samples we could have gotten. So now you're taking a sample from a sample where the sample from which the secondary sample is taken could be any one of a number of things that you could have gotten from the true distribution f. So that tiny change of sampling without replacement rather than sampling with replacement is huge here. So we now have v one to v m a sample from the true distribution f. And we compute theta hat star as we did with the bootstrap for each resample. We do that capital b times to obtain a collection of theta hat stars. Which are estimates of the statistic of interest based on the resample sorry on the subsamples. And we're going to estimate the CDF of theta hat by this thing that I'm now calling l sub m l sub, sub N comma sub, sub comma n. You know what I mean. And I'm going to talk right at the moment, now I'm going to talk only about the centered and rescaled statistic. And I'm centering again here around the value of the statistic, the value of theta hat computed from the original sample. And I am looking at the distribution of the theta hat stars and I'm going to do the same kind of Monte Carlo approximation that we did with the bootstrap. So, again our question is do, does that distribution function converge to the true distribution function of theta hat? And the reason subsampling is kind of another sort of minor miracle here is that this actually only requires two relatively weak conditions. It requires that we know that the centered and scaled version of theta hat converges, and here I'm going to say the word converges in probability, which is one particular sense of convergence to some distribution function that isn't pathological. In other words, it isn't point mass at some one potential value. And that the scaling exponent there alpha which tells us you might call this the rate, the rate at which n has to get large in order to make this work. Is something just that it exists. We don't have to know which either of these things are we just have to know that they exist. So both things seem like relatively weak and general sorts of assumptions that we might actually be justified in making. As with the m out of n bootstrap we have to choose the subsample size and it has to be alike with the M out of N bootstrap. M has to be big but not too big relative to N. And if this is true you, you can prove that the distribution function, the subsampling distribution function converges in probability to the true distribution function as N goes to infinity. Now there are a million, I think there's a whole page of technical conditions that go with this mathematically, and if you want to see them you can go to page 43 of the Politis, Romano and Wolf book. It's actually very interesting but I think that the essence of it is if you're willing to assume that those assumptions are not too restrictive, and in my opinion they aren't, then you can just go ahead with this in this way. So what this means is that the subsampling technique will quote unquote work in the assumption sense in situations where the bootstrap won't. So it will work for stationary and dependent sequences and extremes in a lot of places where the bootstrap either doesn't work well or is really hard to use. So the catch once again you have to choose M and you don't know exactly how to do that. You might have to play some games, so here's a cartoon illustration of subsampling. For a really simple example of, sample of sample of, original sample of size 5 like we had with the bootstrap and a subsample size of 3. So here are the top line here is the original sample 2.82 is the original median of the original sample. And then I've show v one to be vb as the subsamples and their corresponding medians and I've replaced the empirical distribution function on the left with a true distribution function of a pie square three in order to make the point. And alluded to this already, yes it's true that it all depends on the sample we actually have. But we are in the sense of of asserting that this methodology quote unquote works the there's an averaging process going on over all possible samples. And maybe this is a good place to say, that no matter what you are doing, if you just got unlucky and got a lousy sample that isn't really representative of the true process, then you've got trouble. And, you know, there's just, there's no way around that in any statistical methodology. If all you know is the sample you have, then you pretty much have to rely on that sample as having been more or less typical. Now you can do something about the uncertainty that you quote with your estimate. And that is in part what these methods are doing. So, here are the same kind of results we showed for the bootstrap in the M out of N bootstrap, only for the subsampling version of it. I'll direct your attention to the graph on the right. The PDF there, the estimated PDF of the centered and scaled statistic, looks very similar to what we got from the M out of N bootstrap and that's probably not surprising because I set M and N and B all to be the same. And the technique is really rather similar, it's just the only difference is pulling with replacement versus pulling without replacement. And what we can see is that the bias here has been reduced way down, it's now only .015. So it looks like that one difference, resampling with and without replacement made a difference to the bias. It also made a difference to the variance, the variance has now become too small. So that's a danger, we did better on the bias but worse on the variance. so, let me summarize the results in case you can't remember all those numbers from the previous pages which I can't either. So the top line of this table is what the right answer is. At least according to my Monte Carlo approximation with 10,000 trials. We would like it if our centered and scaled distribution had an expected value of minus 0.49 and a variance of 7.118, and we see that the ordinary bootstrap does not do very well. It does not do as well as the M out of N bootstrap. The M out of N bootstrap actually outperforms the subsampling with respect to the variance, but not with respect to the main. So the question is, which one of these things would be the right thing to use and the answer really depends on what your problem is and on what you care about. If you're going to choose between having a less biased estimate or a less variable estimate, that's really a question, a, a decision has to be made in the context of a particular problem. So I'm going to, come to the last slide here of this whole set of lectures on inference and uncertainty and I'm going to quote and paraphrase Dr. Charlie Geyer and I really like this because I think this is actually pretty easy to remember. It says, the right thing is samples of the right size, which is capital N from the right distribution. The Romano and Politis thing which is subsampling is samples of the wrong size from the right distribution. The bootstrap thing is samples of the right size from the wrong distribution. Both wrong things are wrong, we would like to do the right thing but we can't because we really only have one sample at our disposal. So which wrong thing do we want to do? And that goes to the comment of it really depends on what your problem is. Whether you care more about bias or more about variance. And I want to close the entire set of little lectures by saying the following thing, that I think these methods, subsampling and bootstrapping are really the way to go for modern data analysis. They are computationally intensive. And certainly if you're talking about sophisticated machine learning algorithms, which may take a long time to run in the first place, might be a little difficult to do them, but I think these things are on the right track. The caution is that what it means to work is an isotonic guarantee and it's a mathematical guarantee that requires assumptions that we don't necessarily know to be true. So, what I like to do is whenever I have a real problem, I try to cook up a synthetic problem that contains the elements of the real problem that I care about. And then I do a bunch of experiments to try to see how much difference making different assumptions makes. And how sensitive the method itself is to being wrong about certain assumptions. And that can help to decide which wrong thing we want to do and perhaps what cautions we might want to issue along with any conclusions that we draw particularly when it comes to people who might be making policy decisions based on the conclusions that, that that we have drawn from the data that we have. So with that I'd like to say thanks and thanks for paying attention. And we'll see you in the exercises session I believe.