So we talk about outliers, this, this, I think this graphic sort of gives the, the right idea. You know, here most of the data points are falling in a nice little relationship except for one. Right? And so, you know, it's, it's, it's just like one data point to disprove a nice theory, right? And so, very often, you know, people see something like that, it's sort of, like, you know, it's, like, just, just get rid of it, you know. [laugh]. And then things will look good. And every time you see a, an outlyer, you know, the question is. You know, what's causing this? I mean, sometimes this could be just a data mistake, you know somebody entering the data, you know transposed numbers or something like that. And, or it could be that something really unusual happened at this point. So most of the time things are operating, you know, according to some well-defined rule. But, you know, for whatever reason something unusual created this thing to happen. And it wasn't a mistake, but it's, you know, it's actually showing that there's kind of maybe two different structures that are, that are happening. And so when you see outliers you shouldn't necessarily throw them away, because very often outliers, you know, exhibit some very important information. And for us, like risk, when we see crashes in the stock market, we just don't wanna throw those observations away. You know, if our goal is to try to estimate probability of loss, actually the outliers are extremely important in, in telling us what, what these losses can be. The problem with outliers is that they can cause havoc with sample statistics, that is to say, sample statistics are not robust to outliers. So, we can take a sample statistic and we can pollute data with an outlier and we can make that sample statistic any value we want depending upon what the value of what the outlier is. And so in statistics, there's a whole branch of statistics that's called Robust Statistics. And Robust Statistics are about creating sample statistics that are not heavily influenced by outliers. So if we want a measure of the spread of the data. Which is, you know, the typical deviation from the average. We don't want that typical deviation to be influenced by one or two outliers in the tail, because those aren't representing sort of what's happening, you know, typically around the mean. Unfortunately, if you use the sample variance in the sample standard deviation, that's not robust to outliers. You have one outlier that can greatly influence the, the, inflate the standard deviation. And, and to make it a, a misleading statistic. So very often we would like to use some measures of characteristics of the data that are less susceptible to outliers than others. That is, sometimes we would like to use robust measures. Now, outliers greatly influence the sample mean, the variance, the standard deviation, skewness, and kurtosis. These sample statistics are not robust to outliers. On the other hand, when you use percentile measures, for example, if you want to measure the center of the distribution, instead of using the mean, you can compute the median. The median is a must, much more robust measure of the center than the mean, okay. It's not gonna be so heavily influenced by outliers, big and small. Similarly, if you want to measure the spread, about the average, a robust measure of spread is the interquartile range. It's the, you know, the, the third quartile minus the first quartile. Whereas the standard deviation is, is not a robust measure of, of spread. So in, for each one of these sample statistics, there's often a robust version of the statistic that's less susceptible to outliers. Okay. So there's a lot of debate in statistics about what is an outlier. So, you know the picture that I showed you is sort of you know, I mean, outliers of, are, are sort of like what, what, you know, what's the phrase I wanna say? I'll be a little bit polemic. So like, so I wanna say, I wanna say like, outliers are like pornography. Right? It's hard to define what it is, but you know it when you see it. Right? [laugh] And so, so in statistics there, there's a lot of debate. People, when they see outliers, they know what it is but, there's a lot of debate about how to exactly define what an outlier is. And that's because outliers, you know, sometimes you say that an outlier is something that's more than three standard deviations from the mean. Now, the problem with that definition is, the standard deviation itself is influenced by the outlier. So the outlier inflates the standard deviation. So when you say something that's three standard deviations from the mean, if you have outliers in the data, your standard deviation is big and you actually might miss some outliers by using that definition. So, a definition, of an outlier that is, I don't want to say universally accepted but one that is a bit, that's robust to outliers themselves are the following. So, a data point that's often called a moderate outlier is a data point that is smaller than the, so if you think of a, of a distribution, right, going through here. And so we have the, the center of the distribution. So, say we have this as the median, and then we have the twenty-fifth percentile, and we have the 75th percentile here. And this distance here is the interquartile range. The difference between the, 75th percentile and the twenty-fifth percentile. So this is the measure of the, a robust measure of the spread of the distribution. So a, a moderate outlier on the right, on the right tail, is gonna be a data point that is between the 75th percentile and, and you go out one and a half times the in, interquartile range, and then you look at, but it's less than the, the 75th percentile plus three times the interquartile range. So there's a point that is a, a, sort of like out here which is Q.75 plus 1.5 times the IQR. And then there's another point out here that is the seventy-fifth percentile, plus three times the inter quartile range. And if a data point lives in, in, in this area here, it's called a moderate outlier. So it's, it's lying pretty far in the tail. But if a data point is out over here, that if it satisfies this, then it's called an extreme outlier. And similarly on the left hand side you have the same kind of definition.