And so I'm gonna talk about bootstrapping now. And show how the bootstrapping algorithm works. To do things like computing a standard error for, say, the estimate of value of risk. All right. So talking about the boot strap. The boot strap is a very modern, innovation in statistics, and it's part of the computer revelation in statistical analysis. So twenty years ago, when you, wuh, you know, back when I was studying statistics in college, I was actually an undergraduate Stat major at Berkeley. I was a Stat-Econ double major. And, during that time actually the bootstrap was just being developed so this was the middle 1980's. And most of the theorems that were proved that justified the use of the bootstrap was done in the say early 1990's. So it's a very modern statistical technique. And something that is only feasible because we now have fast, cheap computers, right. So, you know, when I was a. I hate to say this when I started college in 1982, you know, we were still. I mean in terms of computer technology, I was the first. And let me see. When I was taking my Computer Programming Class that was the first year that we moved away from punch cards. They used to do computer programming by not, you know, typing things into a terminal but literally you, you punched holes in cards. And that, those it was the computer instructions. And then you fed these cards into a machine and then that compiled your program [laugh]. And if you made a mistake, you had to go and repunch a card. Okay, so, I mean, I missed that by one year. [laugh]. So, I mean, and, and in the lec-, you know? I mean it's amazing to see what, you know, the advances in computer technology. If you go from essentially the beginning of the 1980's until now. I mean, it's, it's quite incredible. You know, what we're able to do essentially on your iPhone in terms of computing, you know, is much, much more powerful than you know, most people could have done on a, you know, a lot of computing platforms. And so the bootstrap relies on being able to essentially simulate many, many, many values you know, in a computer simulation. And because computers are fast and cheap now, and we have programs like R, doing the bootstrap is, is completely trivial. And to a large extent, the bootstrap eliminates the need to do all this, you know, very technical complicated mathematics. Used to be when you took statistics courses you just learned all these formulas and have to grind out all this stuff. Well you don't have to do that anymore. If you want things like a standard error, you don't have to memorize a formula. All you need to do in, in practice is run the bootstrap. So bootstrapping, one of the two great computer simulation innovations in statistics. The other being Markov Chain Monte Carlo Methods, which is a simulation technique used in what's called basing statistics. The motivation behind bootstrapping is like I said. Before modern computers doing statistical analysis involved mathematics, and probability theory, to the derive formulas for standard errors and confidence intervals. Very often these formulas are approximations, their nasty looking formula and often these approximations rely on the central limit theorem and require having very large samples. Okay. And the innovation is, with modern computers and statistical software like R, The bootstrap, which is a, a subset of what's called a resampling method, can be used to produce standard errors in confidence intervals without the use of formulas. And they're often more reliable than the formulas because they don't rely on a bunch of the assumptions that the formulas are based on. So, the advantages of bootstrapping is you know, many fewer assumptions about the underlying statistical behavior of the data. For example, you don't, need to rely on everything being normally distributed. The bootstrap works in a general context. The bootstrap often gives you greater accuracy. You don't have to have a large sample in order to use the bootstrap. You can have a sample size of five. And the bootstrap works just the same way as if you have a sample size of 20,000. So, you don't need to rely on the central limit theorem, for example, to justify what you're doing. And the boot strap has great generality. If you want to construct a bootstrap standard error for the mean, and you want to construct a bootstrap standard error for valiant risk, you do the same technique for both of the quantities. There's not one bootstrap for the mean and another bootstrap for the valiant risk and so on. The same bootstrapping algorithm works for anything that you want to do. So it, it has a great generality. Okay. So let's illustrate bootstrapping in the constant expected return model. Now, so we have continuously compounded returns. Have a mean and an error, which is our random news. The random news is IID normal with mean zero and a volatility sigma squared. So we have two parameters to estimate, the mean and the variance, alright? And we wanna construct a standard air for the mean, and a standard air for the variance. And a 95 percent confidence interval for the mean and a 95 percent confidence interval for the variants, okay? So, we have an observed sample. So this is the data we download from Yahoo. So, let's say we have 100 options, 100 monthly returns on Microsoft or Starbucks, and our goal is to compute the standard air of the mean, the standard air of the volatility. And we wanna compute 95 percent confidence intervals for, mu and sigma. Now. The previous lecture, we derived analytic formulas for these standard errors. Okay? We went through the math. And it turns out that the standard error for the mean is the estimate of the volatility / the square root of the sample size. This was an exact result. And then we had an approximation to the standard error based on the central limit theorem, that says, the standard error for the volatility is the volatility / by two the sample size. Okay? And if we want to construct a 95 percent confidence interval, we can take estimate + and - two the standard error. Okay? Now, we just happen to know these formulas, so we can use them. Now what we wanna do now is, use the bootstrap to construct a bootstrap standard error for the mean. A bootstrap standard error for the volatility. And a bootstrap confidence interval. So how does bootstrapping work? Turns out bootstrapping is very, very close to the Monte Carlo simulation that we just did, okay. We ran a Monte Carlo simulation of the constant expected return model. We specified values for the mean and the volatility, and then we generated pseudo-data by, you know, using the computer to generate normal data with the specified mean and the variance. Bootstrapping works in a similar way except we're not going to simulate data from a normal distribution. We're gonna resample with replacement from the original data. So, we're gonna treat our original data as if it were. Well we're gonna treat our original data as the sort of like balls in a box. And yet you know in your stat class you think about, you know, randomly drawing balls out of a box, the bootstrap works exactly the same way. The balls are the observed data, and the bootstrap sample is randomly picking observations out of our observed sample and creating a new sample. Okay? So, the bootstrapping algorithm, and this is what's known as non-parametric bootstrapping. Alright, so what does the word nonparametric mean? Anybody? Or the, the obvious answer to something that's non-parametric, it's something that's not parametric, right? [laugh] So what does parametric mean? >> If I tell you I have a parametric model what does that mean? >> I told you mounting something. [laugh] You know. >> No. >> So a parametric model is any model that has parameters, right. So the Constant Expected Return model has two parameters. It has a mean parameter and it has a variance parameter. So if you can write down a mathematical formula for a model that has parameters, you have a parametric model. And if something is non-parametric, well, you, you don't have a model. Okay? Something that's non-parametric is just based on the observed data. So in nonparametric bootstrapping, we're not gonna generate our data from this model. We're going to resample from the observed data. So, what we're gonna do is we're gonna create capital B, bootstrap samples we're, where we're gonna sample with replacement from the observed data. And each bootstrap sample, is going to have capital T observations. That's the same number of observations as the original sample. >> And, I'm gonna denote a bootstrap sample, like this so in my first bootstrap sample is gonna be say R11 to R1t so I have T observations and this, this one here represents the, the first sam, bootstrap sample. And the star means that its, its a randomly re-sampled observation from the original data. Then remember our, our original sample is R1 upto R capital T. So when we look at this first observation here. So what we want to do is we want to think of these observations like balls in an urn. And we just. Stick our hand in and we pick out an observation. And so we might pick out observation 22. And that becomes this observation here. Then we do sampling with replacement. So we take that observation, we put it back in the sample, right? And then, you know, we reach in and we pull out a new observation. Say it's observation 55. That becomes the second observation here. And then we put it back in here. And then we reach in again, we pull out another one. Until we get capital T observations. Okay? So that's the idea of random sampling from the observed, sample. Yeah. >> What are, what one? >> Yeah. So here, the first one represents bootstrap. Sample. One, right? And this is, so we're gonna create, say, a hundred samples of size P. And this is my first sample. And so notice that there's a one subject here that represents the, the first sample. The capital B sub-script here, that my last bootstrap sample. Okay? Now the second sub-script represents observation numbers. So this is, sample observation number one, sample ob, observation number two. Yeah, of course. [inaudible], 'cause your sampling with replace. So the question is, in a bootstrap sample, can I have duplicate observations? And yes, you can, because you're sampling with replacement. So you're not gonna rely on one sample. You're gonna have many, many, many bootstrap samples. It could be that one of these you could have, you know, the same observation could end up here 50 times, right? That's just like flipping a coin and you get 50 heads, right? It can happen, but it's very unlikely that it will. But, yeah. >> How many permutations of the observations can you make? >> T factorial. >> T factorial, right. So if T is 100 you have 100 factorial possible samples. That's a lot, right. But if you have five observations you only have five factorial, right. So the, the, the disadvantage of the bootstrap with a small sample size is you only have a. Fine, you have really a finite number of permutations. Right? So even though the bootstrap works, it, it doesn't work as well if you get a sampled size of three or four for example. So alright, so I so we just random sampling with replacement from a population and we create our one pseudo sample and, and the last sample. Now. The rationale and the intuition about re-sampling from the observed data, is well me again, think about what you're doing with statistics. You don't know the model, that the underlying true model. Right? But you observe the data. The data comes from whatever that model is. So if your, wanna create a representative sample, well re-sampling from the observed data is a very intuitive way to create a representative sample from your unknown model. So, So the distribution of each boot-strap sample, is exactly the same as the probability distribution of your observations. And so the idea of the bootstrap, when you do the sampling to create your sample, you want the sampling mechanism that you're going to use to pull observations out of the data to be the same as your underlying assumptions about the probabilistic behaviour of your data. So if you believe your data to be a random sample from a population. Then when you do the bootstrap, you want a random sample from your observed, observations. If you think your observations are a covariant, stationary time series but that. The each observation is uncorrelated with the rest. Then you can still random sample from the population cause that would preserve the uncorrellatedness in the data. If your sample observations are correlated with one another, then when you sample from the bootstrap. From your observed data you want to preserve that correlation. So in that case you're going to sample in blocks. So if you think R1 is correlating with R2, then your gonna wanna sample two adjacent observations at a time for example. And so that the bootstrap would preserve Whatever as- assumptions that you have in, in, in the data. You have an autocorrelated series. The whole idea to do the bootstrap correctly is you want the bootstrap to be able to capture the autocorrelation in the data. And the procedure known as block bootstrapping, that is boostrapping adjacent observations, and so you randomly take out blocks of the data, will preserve the correlation structure in, in the sample. So we're just gonna use the simple, Random bootstrapping for the stuff in this course, but, you know, as we'll see in R, they have functions for doing the block bootstrapping and other things like that. Cuz when you're bootstrapping, so, here the sample is size T or it can be a sample of size M. Your bootstrap sample is always the same size as your observed sample. Right, cuz again you want to capture, I mean the idea is, you are using the bootstrap to compute the, say a standard error for the mean. Your mean estimate was computed on a sample of size T. So, when you use the bootstrap. You're gonna use the bootstrap to compute say a standard error. You want the bootstrap sample to reflect exactly the same size sample as, as your data. Its because you are using the bootstrap to evaluate a statistic that is based on a sample of size T, so you want each of your bootstrap samples to have exactly the same size. If it has a different size, then it's not going to correspond to, you know, the sample that you observe. The bootstrap sample that you observe is, is randomly drawn from this. The distribution of the bootstrap sample is exactly the same as the distribution of the observed data. If the observed data is normal, your bootstrap sample is normal. If your. Observe sample follows a gamma distribution. Your bootstrap sample will follow a gamma distribution. So the whole point about the bootstrap is, whatever the distribution of your data is, your bootstrap sample will follow the same distribution as that data. So it doesn't rely on assuming things are normally distributed and so on. I don't know. I once heard a statistician saying that the bootstrap is the closest thing to magic in statistics that you'll find. Because, you know, you sort of do this and you think, you know, wow this is so easy. But it turns out to be so powerful. And, you know, and it's, it is. It's one of these where it was a, it was a truly revolutionary concept in the area of statistics. To make, you know? Our lives so much easier [laugh]. And we don't have to learn all of these horrible formulas anymore. Alright so, so boo random sampling for a population. Okay, so hopefully you understand what that means. So once we have these bootstrap samples. What do we do with it? Each bootstrap sample we, we calculate whatever statistic of interest that we want. So if you want to compute a standard air for the mean, than on each bootstrap sample we calculate the mean on the bootstrap sample. And so then we'll have capital B. Estimates of our mean. Okay. And so now we have. So any statistical wants this could be we can compute the mean. We can compute the standard deviation. We can compute the value at risk. Anything that we can compute on the sample is, could be theta, right? So, so now we have capital b value of our statistics. Okay? Now, what do you do with these capital B values of the statistic? Well, we use the bootstrap distribution. We could, if, you know, for example, suppose we wanna know, what is the probability distribution of our estimate? Well, that's f of theta hat. We can just look at the histogram of the bootstrap. That will give us an estimate of the probability curve of our estimator. Right? Does it follow a normal distribution? Well, if this looks like a normal distribution, then it does. If it doesn't look like a normal distribution, then it doesn't. If we wanna es-, evaluate bias of a statistic, we look at the mean of the bootstrap relative to the sample mean that we calculate from the actual data. So we can use the bootstrap to estimate bias. If we wanna know the standard devia-, the standard error of an estimator. Screen saver. If you what to know a standard error of an estimator we just calculate the standard deviation of theta hat on bootstrap samples. 'Kay? So, the two things that people are most interested in, is an estimate of bias. So, a bootstrap estmate of bias. So, I'll use a subscript boot to represent something calculated from the bootstrap sample. So, what is, what is bias? It's the expected value of the estimator minus the truth. The bootstrap estimator of the bias. You take the bootstrap mean and you subtract off the sample estimate and that's the bootstrap's estimate of the bias of your estimator. Okay? So it's bootstrap mean minus your sample mean. If I want a bootstrap estimate of the standard error, well, what is the standard error? The standard error of an estimate is the standard deviation of the bootstrap values of your statistic. So the sample standard deviation is one over the number of bootstrap samples minus one. And we look at the Bootstrap estimate of your statistic, minus the. Bootstrap mean, and then you square it, and you sum over all the bootstrap. This is just the sample standard deviation of your bootstrap values, 'kay? Notice that, 'kay, the bootstrap standard error is just the sample standard deviation across the bootstrap. This is trivial to compute. It doesn't matter how complicated your theta hat is. For example, theta hat could be our value at risk, and our bootstrap standard error of value at risk, is just a sample standard deviation of value at risk computed on each of the bootstrap samples. Alright. Now, what if we want to construct a 95 percent confidence interval using the bootstrap? So, here there are sort of two approaches to take. Remember, Last we, last time, if an s, we were talking about constructing a proximate confidence interval as being estimate plus or minus two times the standard error. That's justified if the underlying probability curve of the estimator isn't normal. Well, in the bootstrap, you can look at the histogram of your bootstrap values of theta. If the bootstrap histogram looks like a normal curve, then you can use our nice, easy formula estimate, plus and minus 2x the bootstrap standard error, to compute a 95 percent confidence interval. Okay? If the bootstrap doesn't look like a normal curve, so if the bootstrap histogram is very skewed, then this is gonna be not very accurate. A more accurate confidence interval is gonna be based on the quantiles of the bootstrap distribution. So a 95%. This is called a percentile confidence interval. You're gonna take the 2.5 percent lower quantile, and the 97.5 percent upper quantile. So the, this range contains 95 percent of the bootstrap observations. Okay. That would be your estimate of your 95 percent confidence interval. And so again very easy to compute because these, these are just sample quantiles from your Bootstrap values. Now, let us talk little bit about actually doing the Bootstrap in R. So, you have two choices, brute force, that is programming the boot strap by hand. And doing the boot strap by hand by brute force, the coding is exactly like the Monte Carlo simulation. You write a fore loop and, and then you just resample from your observations inside the fore loop. There is a R package called Boot. As you'll find out, there's an R package for doing almost everything imaginable. There are, in fact, almost 4,000 different R packages. Okay. There's almost like 50 new R packages a day, that show up. And one of the, the, the Ya know, sort of, the most difficult things associated with r is just, sort of, keeping up with, you know, what's all that, that's available. The boot package r is very good. It's written by, two of the inventors of the bootstrap. And, it's very reliable and, and easy to use. Alright. So if we do the brute force version of bootstraping, what do we have to do? So, you sample with replacement from the original data. And, there's a nice arch command called sample that allows you to random sample from a data matrix, or anything, any object. So, if you have your data in the data frame, or your data in the matrix, then you can create random samples from that matrix using sample. And you do this capital B times. And then, you wanna compute a statistic from the bootstrap, from each of these bootstrap samples. Now, typically, we, we, capital B is set = to 999. So usually, you know? And the reason why you'd use 999 is, Well, one of the reasons is, when you calculate quantiles in order to get, you know, sort of an exact quantile, you need an odd number of observations. So if you want the median of a sample, you can get a median, exact, but, you can get exactly the center of the distribution if you have nine observations or some odd number of observations. If you have an even number of observations, you can't actually get halfway. So very often when you're doing bootstrapping, that's why you see an odd number of values, 99, 999, something like that.