So, the first kind of data descriptive Statistics is going to be a graphical statistics. It's called the histogram. In the histogram, the idea is we want to describe the shape of the distribution of the underlying data. And so, how do we construct a histogram? So, we order the data from largest to smallest. Then, we divide the range of the data into, say, in equally spaced bins. And then, we count the number of observations in each of the bins and then we create a bar graph of those counts. And we might normalize the area so that it's equal to one. And so, we're going to get a, you know, kind of a bar chart of the distribution of data and we can think of this histogram as a very crude estimate for the underlying probability curve, alright? So, in R, one of the reasons why R is very nice is that it has all of these data descriptive statistics built into the language, right? So, if we want to create a histogram, we just use a function called hist and it produces our histogram plot for us. And we can change the number of bends and, you know, do all sorts of things as an option to the hist function. And so, I encourage you, you know, just to play around with it and, you know, and, and see what you get. So, let's look at some histograms of the underlying data and try to get an idea of what the distribution looks like, okay? Now, before I do that I want to use Gaussian white noise as a benchmark model for the returns on Microsoft and the returns on the S and P 500. And we want to think about whether or not the returns on Microsoft and the returns on S and P 500 can be considered as a realization from a Gaussian white noise process, okay? So, to help you do this thought experiment I've plotted up here, the monthly continuously compounded returns on Microsoft. This is the actual data. Now here, what I've done is, I've created a computer simulation of, of a Gaussian white noise that has the same mean value as Microsoft and it has the same volatility, the same standard deviation as Microsoft. And I just created the 120 random draws from that you know, Gaussian distribution, that's calibrated to have the same mean and standard deviation as Microsoft. And so, here's the computer simulation and here's Microsoft. Now again, the differences between these simulations are, you know, here it's the volatility changing in the actual data is not reflected in the Gaussian white noise data because this always has the same volatility all the way throughout the process, okay? So, that's one thing that's clearly different between the simulated data and the actual data. And then, the question is well are the things that are similar, the mean values of these two series are the same by construction? And what about the empirical distribution? Let's calculate the histogram of the actual data and calculate the histogram of the normal data and see if they look similar. So, here's a histogram plot of the monthly returns on Microsoft, okay? So this is done using the hist command in R and I don't use any special options. I use just all the defaults okay. So, this is an estimate of the underlying probability curve. So, right away, one of things that you should notice about this is that it looks very similar to a bell-shaped distribution, okay? And so, this is one of the reasons why the normal distribution is often used as a benchmark because when you look at histograms of returns they, they resemble a normal distribution, okay? Now, it's blocky because, you know, we're, we're, we're dividing these things into bins. Now again, when we look at a histogram, you can say, where is the histogram centered, that's the mean return, the volatility is the standard deviation in the histogram, the skewness is the asymmetry in the histogram, and the excess kurtosis is telling us about how thick the tails are relative to a normal distribution, okay? So, here, you know, if we think there might be a slight long left tail. So, if we're going to guess, it might have a negative skewness. And the tails, you know, kind of are, you know, well it's hard to say. Are they fatter than the normal distribution? It's hard to tell from the histogram by itself. But the fact that there is this kind of big, outlying value, is a clue that perhaps the kurtosis is bigger than three. Okay. Now, here is a histogram of the computer simulated Gaussian data, okay? This thing should look like a normal curve, because it actually generated from a normal distribution, okay? It has the same mean as the returns on Microsoft and the volatility is the same volatility of Microsoft and notice the, the range of the data. So, it goes from, you know, -two to +two. And, if we look at the actual Microsoft data, this is going out to -0.4 to 0.4. So, it's looking like the Microsoft returns have fatter tails than, than the Gaussian data, because of the, how far they're going out in the actual data, okay? Now, if you look at the S and P 500. So, here's the again, we use the hist command on the S and P 500 monthly returns and so it's, it's pretty, I mean, the histogram is pretty blocky here, you know, the default I could have probably used a few more bins. But I think the striking feature of the histogram is the fact that it has a pretty clear, long, left tail. So, there's a pretty clear negative skewness in this distribution and so, you know, if you think, is the normal distribution a good model for the S and P 500 returns, probably not, you know, because of this, this big negative skewness. Now if you want to compare Microsoft to the S and P 500, we would like to use the same x-axis scale for the two return series. And so, what I did was I first computed the histogram for Microsoft, and then I used the bin ranges for Microsoft to compute the histogram for the S and P 500 and the reason for doing that is you can see now that the S and P 500, the data or much more concentrated around zero than for Microsoft. Microsoft has a much larger spread above the average than the S and P 500. And so, that's just another way of saying that the volatility of Microsoft is higher than the volatility than, than the S and P 500. I made the comment that I thought the kurtosis of this was, was bigger than three, that the kurtosis is bigger than the standard normal. And the reason is because the, in the actual data, notice that this is, the left tail is -40%, the right tail is +40%. When I did the computer simulated Gaussian data, the right tail is only going to -twenty percent and +twenty%. So, the Microsoft returns are out here in the left, left tail, and are out here in the right tail. So, the actual data produce more extreme values than the computer simulation from a normal distribution. So, that's why I'm saying it's probably the case that the kurtosis of the actual Microsoft data is going to be bigger than three. Alright. So, when we look at histograms, his histograms are blocky, okay? Now, if we want to eliminate the blockiness of a histogram, we can compute what's called a smoothed histogram. And so, this is a graph, of a smoothed, of the smoothed histogram and it's created with the R function called, R function called density. And another name for a smoothed histogram is a kernel density estimate, okay? And literally, what you're doing, you know, if you think about it, think about drawing a smooth curve over the data, that's what the kernel density algorithm does. And how do you get a smooth curve throughout the data? Well, at a given point, what you do is you take an kind of a weighted average of values around a particular point where the weights you know, kind of decline going out in this direction and decline going out in that direction. And then essentially, what you do is you, you take that kind of smooth and you pass it all the way across the data and that allows you to draw this kind of smooth curve over the histogram. The, the algorithm for computing the smooth histogram, this kernel density estimate, this is taken from Rupert is of the following form. So, if you want to estimate the, the probability curve at a single point, then what we do is we take a weighted average where the weights are determined by what's called a kernel function. And a kernel function is very often just a symmetric probability distribution. And usually, the kernel function is the standard normal curve. And essentially, what you're doing is you're, you're taking a weighted average of, of the values in a histogram where you're weighting the points with a normal distribution before and after the given point. And that allows you to smooth the blockiness of, of, of the histogram, okay? The, when you use this kind of function there's a, a parameter B that's called a bandwidth parameter and that controls the degree of smoothing. So, if B is big, then you get a lot of smoothing and if B is small then you don't get very much smoothing, okay? So, anyway, for our purposes, I just want you to know that sometimes, it's, it's easier to look at the smooth version of the histogram to get a, a general view of the shape, instead of the histogram by itself. And so, notice you get, when you do the smooth, you get some sort of wiggles out here, you get like little wiggle over here. And so, you know, does this look like a normal curve? Well, not exactly, you know, it's a, a, a nice normal curve would be perfect symmetric and bell-shaped, whereas you know, this is a little bit less, less perfect, okay? So, we can overlay, for example, the smooth histogram on top of the, the histogram from Microsoft. And so, we see how the smooth, you know, sort of picks up this little bump, picks up this little bump here, you know, captures the basic shape of the distribution in the middle, gets that little bump here from these two things, and so on. So, we often, you know, will report a histogram with the little smooth line on top of it. And again, you just want to get an idea of what the shape of the distribution looks like.