So at the end of last question, last lecture, we were talking about outliers in data. And an outlier in the data is something that's unusual, doesn't follow the, the normal pattern of the data. So this is a little cartoon that shows, you know, most of the data points fall in this nice curvilinear relationship, but there's one data point that's off. And so this would typically be called an outlier, because it's, it's unusual. And outliers could happen because of data mistakes. Or outliers could happen because there's really something special about this point, such that, it's, its, you know? It's not following the, the usual relationship. And perhaps we ought to understand better what's, what's happening at this point. In financial data, very often outliers are, observations that are associated with market crashes, you know like the 1987 stock market crash or something like that, that's the day when the returns dropped twenty percent in a day, and then you know, you see something, you know, like, like observations in the extreme tales of the distribution or if you're looking at an individual stock you might see all of a sudden, you know one day you know the return is big and negative and there might have been a very bad earnings announcement or something like that and that could be driving the observation down. So sometimes we can come up with an explanation for an outlier, sometimes we can't. From the point of view of descriptive statistics, outliers do have a impact on sample statistics like the mean, the standard deviation, the kurtosis, the skewness. And, and so when you look at statistics computed with and without outliers you can actually get very, very different results. So let me illustrate this a little bit. So here I took the monthly returns on Microsoft and I created an outlier. So I created a negative return that was you know, a continuously compounded return space, was -80%. And so notice that, you know, this observation is very, very different from, from the others. It's sort of like, you know, very bad news happened on, on this date. And if we look at the histogram, we see that, that outlier showing up in this, this big, you know, tail. And, and one question that you know, we often have about outliers is what is the impact of the outlier on, on our descriptive statistics. We see that the histogram now has a bump way out over here. And then another question is. What happens to you know, the sample mean, the sample standard deviation, skewness and kurtosis? You know, when the outlier is included in the data and when it is not. So, I did some simple calculations, where, essentially, in blue here, I have sample statistics that do not include the outlier. So the mean of Microsoft returns when the outlier was not included is, you know, about .67%. In black, on the right hand side, I recompute the sample mean where I include that big, negative outlier. And now when I compute the sample mean the, it turns out, goes from a positive number to a negative number. And, and it's you know, so we see that the one big negative outlier has pulled the mean from positive .67 percent to, to a negative number. You know, and again this is, you know, if we're looking using the mean as a, our best guess for the expected monthly return, you know, here we're actually making money, now we're losing money on average, 'kay? So we see that the big impact of, of one observation on the interpretation of the desirability of this asset. Similarly we can look at the sample standard deviation which measures the volatility. When you don't have the outliers ten%, when you do have the outliers almost fourteen%, 13.7%. So we see that one observation has inflated the standard deviation, you know, 30%. And so that's a, that's a big impact of, of one observation. And then skewness without the outlier is -.07, with the outlier is -2.3. So the big negative return obviously has given us a big negative skewness. And then our kurtosis, or actually this is excess kurtosis. Without the outlier it's 1.8, with the outlier is fourteen. So we see that, you know, the impact of this outlier to, you know, it can greatly change these, these sample statistics, and so then again, you know, the question is do you keep the outlier in or not. Kind of depends upon the purpose at, at, of, of what you're doing. If you're worried about tail risk, and if that big negative outlier really represents something that, you know, could happen again, then you would probably want to keep it in to get a, a better estimate for, you know, you know, a loss that, that could happen due to a big negative return. On the other hand, if you know, you're, you're interested in perhaps in, in estimating the expected return, Most of the time, Microsoft returns is positive. Every now and then, you get a negative return, but on average, you know, you wanna know what the expected gain is. Then maybe if we're estimating the mean, you, you might want to down weight that outlier in a computation. Now, there's a nice graphical, summary statistic for a distribution that also highlights outliers in the data, and this is called a box plot. And a box plot is a, you know, again, something that's created, oh, I should, actually, in. And a lot of the scripts in statistics have their origins from a statistician whose name is John Tukey. And he was very much about looking at data and trying to create ways of, of summarizing data, and in, in very nice and nifty ways. And so, he's the father of the box plot. And the idea of the box plot is you, you show the basic features of a distribution of one-dimensional data, and You, You illustrate features of the data using sample statistics that are robust to outliers. So for showing the center of the data instead of using the mean you'll use the median, because the median is less sensitive to outliers than, than is the mean. And similarly to show the spread in the data instead of using the standard deviation, you would use the interquartile range, the difference essentially between the first quartile and the third quartile, as a measure of the middle of the distribution, and again the quartiles are less sensitive to outliers than the standard deviation so that gives you what they say is a more robust measure of spread. And then there are outer fences that are shown in the box plot that illustrate a moderate outlier and an extreme outlier. And last time I sort of said there was a, these definitions of outliers that are based on, on percentiles and a moderate outlier or something where you look at the, the 75th percentile, and you look at an observation that's beyond the 75th percentile, but less than the 75th percentile plus three times the interquartile range. And the interquartile range is just different here. So, something out here is a moderate outlier. And then an extreme outlier is something that's even farther out like this.