So the box plot gives you these kind of robust measures of a distribution. So here's an illustration of the box plot for the monthly returns on Microsoft. And so this black line in the middle represents the, the median. This is roughly the interquartile range. It's not exactly it and then if you look at the online help for a box plot, you'll see why. Notice that this is not quite symmetric. So by definition, the median is in the middle of the interquartile range, but This is so, so that these, these fences are, the, the, the edges of the box are not quite the, the. The first and the third quartile, and then you have, You know, essentially these fences that, illustrate, you know, when you have extreme outliers in the data. So these points out over here, essentially are representing those values that are, you know, what, what we characterize as an extreme outlier. And this would be in the right tail, and this would be in the left tail. So you have a, the bulk of the distribution is roughly symmetric. You have two big negative outliers and you have a few big positive outliers. And again, if you look at the online help for the box plot, ya know, they'll tell you exactly how this fence and this fence is determined. From the point of view of interpretation, the big power of box plots comes when you compare multiple series. Because it allows you to, you know, very quickly look at the distribution of multiple series in the same plot. So, here's a, a box plot where I have the box plot of the Gaussian white noise. This is the simulated data that comes from a normal distribution that has the same mean invariance as the Microsoft returns. These are the, this is the box plot for the actual Microsoft returns and this is the box plot for the SMP 500 returns. And so we see that they're on the same scale which is percent. So the average return for all of these things is pretty close to zero and a little bit positive. So, that's the median return. We see that you know, again, Microsoft and Guassian White Noise, the middle of the distribution is, is roughly the same. And, what we see with the Microsoft data is, we have some outliers here. With the Gaussian white noise data we don't have the outliers because normal distribution doesn't really produce outliers. And with the S and P 500 data, we see the smaller spread than Microsoft, and we see the small two negative outliers. So again, we can get sort of an idea of the shape of the distribution. So it's similar to the histogram, but again it's based on robust measures, and it's quite useful for doing quick comparisons of many series at once. Now, I've put together a little four graph summary for looking at the distribution of asset returns. And this four graph summary is borrowed from a very nice book by Rene Carmona called The Statistical Analysis of Financial Data. And in this four graph summary there's going to be a histogram, a box plot, a smooth histogram, and then a QQ plot relative to a normal. So we have four pictures and it looks like this. So, say for Microsoft, here we have the histogram, here we have the smooth density down here so you can get a rough shape of the distribution. And then, we have the box plot over here, and the QQ plot relative to the normal. And all of these, these three plots are giving you similar information. We see kind of a long left tail, a negative skewness, so we see a little long left tail over here. We see the negative outliers in the box plot, which is corresponding to these observations here. And in the Q-Q plot relative to normal, we're seeing this drop-down relative to the normal graph, and this dipping up a little bit on the right hand side. So if you look at, you know, a, a p-, a summary graph like this, you can get an idea that Microsoft returns are. Kind of approximately normally distributed in terms of the bulk of the distribution looks normal. But there is some negative skewness and fatter tales relative to the normal distribution. Okay. Alright. So when we have two or more random variables, then, we have, we might want to look at some descriptive statistics that tell us the relationship between two or more variables. Now, when we studied probability theory, when we were looking at the dependence between two random variables, we looked at covariance and correlation. Now, from the point of view of descriptive statistics, we can, we have, say we have two random variables X and Y. And then we observe the sample X1, Y1, X2, Y2. So, we can think of X as the return on Microsoft, Y as the return on the S and P 500; and then our sample is the data that we download from Yahoo. Okay? We assume that these returns, you know, follow a multi-, a bi-variant normal distribution. For example, as, as being kind of the benchmark. And then, say we wanna measure the dependence between the Microsoft returns, and the S and P 500 returns. We'd wanna compute the, sample covariance and the sample correlation, to get, an idea of what the data say the relationship between these returns are, okay. Now, in, looking at pairwise relationships, there's a graphical, diagnost-, a graphical descriptive statistic called a scatter plot. And a scatter plot is just an x-y plot of your bi-variant data. So you can plot the returns on the SMP500 on one axis and returns of Microsoft on the other. And then you can see what the data looks like. So for example. Here's a scatter plot of the monthly returns on Microsoft versus the S and P 500. So here, I, I put Microsoft on the X axis. The S and P 500 on the Y axis. And the black lines here represent. This is the mean for Microsoft, this is the mean for the S and P 500. Okay? So remember when we, we studied co-variance and correlation, we, we looked at these probability scatter plots. And you know? Essentially, when you look at this data, you know? What would you say, are, is there a positive or negative relationship between Microsoft and the S and P 500? Positive, right? Cuz as Microsoft returns go up, S and P 500 returns tend to go up as well. As the Microsoft returns go down, the S and P 500 returns are going down. Okay? So here we see, we see a negative relationship. And later on, we'll compute the sample co-variance and correlation. And we, there's a positive sample co-variance and the sample correlation is.6. So again, there's a reasonably strong positive linear association between these two returns. Okay? Now a. When you have more than two returns, there's a nifty function in R called pairs, P, A, I, R, S. And what the pairs function does, it creates all pair y scatter plots. So here I have three data series, my computer-simulated Gaussian white noise, the returns on Microsoft, and the returns in S and P 500. And so now I have what's called a scatter plot matrix. And, so this graph right here, this is the scatter plot with Gaussian white noise on this axis and Microsoft on this axis. And we see the scatter plot looks like a shock and blast. There appears to be no linear association between them. So we would assume that the covariance is close to zero and the correlation is close to zero, based on this plot. This plot represents Gaussian White Noise on this axis, and the S and P 500 on this axis. Okay? And, again, this plot also kinda looks like a s-, a shotgun blast, where it doesn't appear to be any systematic positive or negative relationship in the data. If anything, you know, there might be what might, what looks, maybe, to be a slight negative relationship in, in the data. But it certainly isn't very strong. So one might expect this sample covariance to be slightly negative, and the correlation to be a negative number that's, but pretty close to zero. This plot over here is just, we put Microsoft on this axis and Gaussian white noise on this axis. So this plot is a, just slipping the axis from this plot, okay? And similarly, this, this plot down here is flipping this plot. And then finally our last plot shows Microsoft on this axis, SMP 500 on this axis. So that's what I showed you before. And we see positive linear relationships. So, these two series are positively correlated positive covariance and a positive correlation. So again, this is nice if you have ten assets. You can do this and then you can get a very quick summary of what appears to be the linear association between the variables. And you can also I mean again, cause these are plots. Even if there's a non linear association you know, you would see, you could possibly see that in, in the plot itself. So, We often summarize the sample variances and covariances in a matrix. And actually before I get there let me define the sample statistics. So if you wanna compute the sample covariance, then you would take your sample of your first data series, the sample of your second data series. You compute this sample average of x minus its mean, times y minus its mean. And that's sometimes this is called at, little s with a subscript xy, or sigma hat xy. So this is the sample covariance and then the sample correlation is the sample covariance divided by the product of the sample standard deviations. So if we're working in R, the bar function computes sample covariance matrix. So if you so if we have three assets, and we have a three by three matrix, then have the variances along the diagonals and the covariances on the off-diagonals. That's what the bar function computes. Notice that there isn't a, actually, there is a co function, but since with the co function, does it same thing as, as a bar. The core function COR, gives you the sample correlation matrix. So, here. So if I take my three data series, my Gaussian white noise, my Microsoft, and my S and P 500. I use the bar command. And then this is the variance/covariance matrix. The sample variances are along the diagonals. And the covariances are on the off diagonals. So, so notice that we see a negative sample covariance between Microsoft and Gaussian white noise and a negative covariance between Microsoft, sorry, between the S and P 500 and the Gaussian white noise. And we have a positive covariance between the S and P 500 and Microsoft. Now, covariance is just direction of association, correlation of strength. So when you look at the correlation matrix, you have ones along the diagonal because the correlation with each series with itself by definition is one. And so here's the correlation between the Gaussian white noise and the Microsoft, that's -.19 Not very strong. Correlation between Gaussian white noise and S and P 500 is -.24. Now this is a completely spurious correlation, because the Gaussian white noise was computer generated. Right? It has nothing to do with the actual S and P 500 data, but we still had a number that's, you know, kinda large and negative, and that's just by chance. Right? If I simulated another Gaussian white noise, then this correlation, probably close to zero or could be a little positive, or something like that. The Microsoft and S and P 500 data has a reasonably strong correlation, that's .6. The last descriptive statistic I want to talk about is, has to do with time series dependence. So again we were covering probability theory, we went through the time series concepts section, and we were trying to think about time dependence in data. And so we defined what are called autocorrelations. We defined the covariance between y(t) and y(t-1), the correlation between y(t) and y(t-1) as well. Well, we can compute sample versions of these quantities. So given, the time series of data. If, and, and if we're interested in determining if there's linear dependence over time. Then we can compute the sample auto covariance and the sample auto correlation, and see if these things are different from zero. So the sample autocovariance is just at lag J, is just the sample covariance between XT and XT minus J. And the sample auto correlation is just this sample auto co-variance divided by the sample variance, and these are measures of linear dependence between a variable and its lags, and then we can do a graphical plot called the sample auto correlation function, where we just plot this sample auto correlation against the lag. So, notice that when you calculate this sum. We start the summation at J plus one, right, because we're looking at the summation. So, the first index in this sum is J plus one, so we cut, we look at here. That's going to be X, J plus one minus the mean. And then we have XJ + one - J. So that's X1. Right? Minus the mean. So we're looking at the relationship between XJ and J lags from XJ. So from the first observation to the Jth observation. Then the next term in the sum is XJ + one. And then this would be X2, and so on. So we're looking at the relationship between, you know, X at time T and its lag. So if I guess if you draw a picture. Right? So we have you know, one, two, three, four, you know, up to J, and so this would be XJ, and then we have X1, and then we have XJ plus one X2. So when we're calculating this sum, so we're looking at the relationship between this variable and this variable, this variable and that variable, and in terms of the computation of this. If we started this at zero. Then this would be XO. Or say, if we started this at one, then this would be X1. But this would be X1-J. And we don't have data, for that. So, this, this notation is to emphasize that we have to do this computation based on the data sample that we actually observed. Okay, so here's an example of sample auto correlations. So this is, these are plots that are measuring the estimated time dependence in the data. And the top graph is for the Gaussian white noise. And the middle graph is for Microsoft. The bottom graph is for the S and P 500. And this vehicular is being plotted here. So, this is I guess, lag one, lag two, lag three, lag four. So for Gaussian white noise, the estimated correlation between xt and xt minus one is, is actually zero. It seems, you're just, because you're not seeing here. The estimated correlation between xt and xt minus two is a small negative number. So the scale here this is minus 0.15. So this value is you know like minus 0.05. So this is a very small number. Now on these graphs are blue dotted lines, okay. These blue dotted lines are thresholds to determine whether or not these values are statistically different from zero. Alright. After the midterm I'll explain where these blue dotted lines come from. They're essentially based upon 95 percent confidence interval for these estimates. So when you look at this graph, how you're supposed to read it, is if any of the estimated auto correlations extend beyond the blue line, then those auto correlations are thought to be statistically different from zero at the at with, with 95 percent confidence. And so for the Gaussian white noise, we see none of these auto- correlations lie outside of the blue dotted lines. So we see no evidence of time dependence in the data. And again, that should be the case. Because the Gaussian white noise is computer simulated with no time dependence in the data. On the other hand, if we look at the Microsoft returns, we see that the first return is negative, and it's about -.2. So this is the correlation between the return in month T and the return in month T -one. And notice that it's negative. And it extends beyond, beyond the blue dotted line. So we would view this as being statistically significant. So there appears to be a negative correlation between the return this month and the return last month. Okay? Now, the other auto-correlations. So this is the, return, the autocorrelation between the return at month T and the return at month T minus two. So this is the lag two autocorrelation, the lag three autocorrelation and so on. Here, notice that the lag two is not beyond the blue dotted line but the lag three is slightly positive, 'kay. So if we run through the Microsoft data and we view this graph we would say there appears to be some evidence for time dependence in the data, you know. But if we, then, if you look at the S and P 500 data. We see none of the sample auto-correlations are outside of the blue dotted line. So for the S and P 500 data, there appears to be no evidence for time-dependence in, in the date. In statistical analysis you should never say there is, right, because you're, in statistics you're never certain, you're never 100 percent certain of anything. Right? All you can say, and you want to think of yourself as, like, sitting on the jury evaluating evidence, right? Is the evidence in favor of, or is the evidence against, something, right? So you would interpret this as, there's data evidence in favor of some time dependence in the Microsoft returns. When we look at the Microsoft data, just how do you interpret this negative correlation? And it's literally saying that, if the return this month is positive. Then, there's a tendency for the return next month to be negative. Right, so there's kinda of a reversal that's going on here. >> And vice versa. >> And vice versa, and if the return is negative this month there is a tendency for the return to be positive next month. Right, so, one could ask the question, what could be causing such a reversal? Now, in finance there is some explanation for the reversal effect in asset returns that's known as the bid-ass bounce. I'll come back to that a little bit later on. But that's, that's one explanation. It's, it's not a very good explanation for monthly data. It's, it's a better explanation if, if you have say, asset returns computed every minute. But for monthly data, it's a bit of a mystery. You know, why is this negative correlation here? Okay, so to summarize, I like to call these stylized facts. So in, in this section we looked at three assets. We looked at Microsoft, the S and P five, actually we only looked at two assets. We looked at Microsoft and we looked at the S and P 500. And of course, one is always tempted to try to make sweeping generalizations based upon, or, maybe I should rephrase this. You should be careful not to make sweeping generalizations based on the analysis of two assets. But, you know, I've, I've analyzed a lot more than two assets, and others have analyzed a lot more than two assets. And, so when you look at monthly continuously compounded returns, these are the results that you tend to find when you look at, you know, lots and lots and lots of different assets. At the monthly basis, returns appear to be approximately normally distributed, okay? They don't follow the normal distribution exactly, But the normal distribution is not a horrible mistake. All right? There's some noticeable negative skewness and excess kurtosis in typical assets, right? If you look at bivariate relationships, many assets are contemporaneously correlated, that is, and they tend to be positively correlated. And so, and that's one of the things as well, if you look at temporal dependence in the data. At the monthly level there is not a whole lot of evidence for strong temporal dependence, so assets are approximately uncorrelated over time. Then again when you look at many assets, the Microsoft data showed there was some negative dependence on one lag, but when you looked at, you know, other lags there was not much that was there. So these are broad, what I call stylized facts and you want to keep these sort of in the back of your mind. Stylized facts are useful at the model building phase. So if you want to build a probability model for asset returns, we want our probability model to capture the basic stylized facts of the data, okay. If our model doesn't capture the basic stylized facts, it's not a good model. And we're always gonna keep that in mind as we look at particular models in this class. We always wanna ask ourselves, you know, if we you know, simulate data from our model, does it look like real data?