1
00:00:00,000 --> 00:00:04,071
So we talk about outliers, this, this, I
think this graphic sort of gives the, the

2
00:00:04,071 --> 00:00:07,088
right idea.
You know, here most of the data points are

3
00:00:07,088 --> 00:00:11,005
falling in a nice little relationship
except for one.

4
00:00:11,005 --> 00:00:13,074
Right?
And so, you know, it's, it's, it's just

5
00:00:13,074 --> 00:00:16,085
like one data point to disprove a nice
theory, right?

6
00:00:16,085 --> 00:00:21,057
And so, very often, you know, people see
something like that, it's sort of, like,

7
00:00:21,057 --> 00:00:24,067
you know, it's, like, just, just get rid
of it, you know.

8
00:00:24,067 --> 00:00:27,001
[laugh].
And then things will look good.

9
00:00:27,001 --> 00:00:31,071
And every time you see a, an outlyer, you
know, the question is.

10
00:00:31,071 --> 00:00:35,034
You know, what's causing this?
I mean, sometimes this could be just a

11
00:00:35,034 --> 00:00:39,067
data mistake, you know somebody entering
the data, you know transposed numbers or

12
00:00:39,067 --> 00:00:43,022
something like that.
And, or it could be that something really

13
00:00:43,022 --> 00:00:47,028
unusual happened at this point.
So most of the time things are operating,

14
00:00:47,028 --> 00:00:49,081
you know, according to some well-defined
rule.

15
00:00:49,081 --> 00:00:53,092
But, you know, for whatever reason
something unusual created this thing to

16
00:00:53,092 --> 00:00:56,039
happen.
And it wasn't a mistake, but it's, you

17
00:00:56,039 --> 00:01:00,084
know, it's actually showing that there's
kind of maybe two different structures

18
00:01:00,084 --> 00:01:04,084
that are, that are happening.
And so when you see outliers you shouldn't

19
00:01:04,084 --> 00:01:08,881
necessarily throw them away, because very
often outliers, you know, exhibit some

20
00:01:08,881 --> 00:01:13,067
very important information.
And for us, like risk, when we see crashes

21
00:01:13,067 --> 00:01:17,086
in the stock market, we just don't wanna
throw those observations away.

22
00:01:17,086 --> 00:01:22,042
You know, if our goal is to try to
estimate probability of loss, actually the

23
00:01:22,042 --> 00:01:27,040
outliers are extremely important in, in
telling us what, what these losses can be.

24
00:01:27,059 --> 00:01:33,021
The problem with outliers is that they can
cause havoc with sample statistics, that

25
00:01:33,021 --> 00:01:36,080
is to say, sample statistics are not
robust to outliers.

26
00:01:36,080 --> 00:01:42,009
So, we can take a sample statistic and we
can pollute data with an outlier and we

27
00:01:42,009 --> 00:01:47,045
can make that sample statistic any value
we want depending upon what the value of

28
00:01:47,045 --> 00:01:54,996
what the outlier is.
And so in statistics, there's a whole

29
00:01:54,996 --> 00:01:58,093
branch of statistics that's called Robust
Statistics.

30
00:01:58,093 --> 00:02:03,094
And Robust Statistics are about creating
sample statistics that are not heavily

31
00:02:03,094 --> 00:02:07,095
influenced by outliers.
So if we want a measure of the spread of

32
00:02:07,095 --> 00:02:11,025
the data.
Which is, you know, the typical deviation

33
00:02:11,025 --> 00:02:15,006
from the average.
We don't want that typical deviation to be

34
00:02:15,006 --> 00:02:20,015
influenced by one or two outliers in the
tail, because those aren't representing

35
00:02:20,015 --> 00:02:23,096
sort of what's happening, you know,
typically around the mean.

36
00:02:24,015 --> 00:02:29,011
Unfortunately, if you use the sample
variance in the sample standard deviation,

37
00:02:29,011 --> 00:02:33,043
that's not robust to outliers.
You have one outlier that can greatly

38
00:02:33,043 --> 00:02:36,055
influence the, the, inflate the standard
deviation.

39
00:02:36,055 --> 00:02:39,053
And, and to make it a, a misleading
statistic.

40
00:02:39,053 --> 00:02:45,907
So very often we would like to use some
measures of characteristics of the data

41
00:02:45,907 --> 00:02:49,497
that are less susceptible to outliers than
others.

42
00:02:49,497 --> 00:02:53,185
That is, sometimes we would like to use
robust measures.

43
00:02:53,185 --> 00:02:58,054
Now, outliers greatly influence the sample
mean, the variance, the standard

44
00:02:58,054 --> 00:03:03,331
deviation, skewness, and kurtosis.
These sample statistics are not robust to

45
00:03:03,331 --> 00:03:06,721
outliers.
On the other hand, when you use percentile

46
00:03:06,721 --> 00:03:11,451
measures, for example, if you want to
measure the center of the distribution,

47
00:03:11,451 --> 00:03:14,733
instead of using the mean, you can compute
the median.

48
00:03:14,733 --> 00:03:19,959
The median is a must, much more robust
measure of the center than the mean, okay.

49
00:03:19,959 --> 00:03:24,043
It's not gonna be so heavily influenced by
outliers, big and small.

50
00:03:24,043 --> 00:03:28,405
Similarly, if you want to measure the
spread, about the average, a robust

51
00:03:28,405 --> 00:03:31,536
measure of spread is the interquartile
range.

52
00:03:31,536 --> 00:03:36,283
It's the, you know, the, the third
quartile minus the first quartile.

53
00:03:36,283 --> 00:03:40,881
Whereas the standard deviation is, is not
a robust measure of, of spread.

54
00:03:40,881 --> 00:03:45,603
So in, for each one of these sample
statistics, there's often a robust version

55
00:03:45,603 --> 00:03:50,179
of the statistic that's less susceptible
to outliers.

56
00:03:50,179 --> 00:03:54,141
Okay.
So there's a lot of debate in statistics

57
00:03:54,141 --> 00:03:59,137
about what is an outlier.
So, you know the picture that I showed you

58
00:03:59,137 --> 00:04:04,509
is sort of you know, I mean, outliers of,
are, are sort of like what, what, you

59
00:04:04,509 --> 00:04:09,153
know, what's the phrase I wanna say?
I'll be a little bit polemic.

60
00:04:09,153 --> 00:04:13,757
So like, so I wanna say, I wanna say like,
outliers are like pornography.

61
00:04:13,757 --> 00:04:17,154
Right?
It's hard to define what it is, but you

62
00:04:17,154 --> 00:04:20,121
know it when you see it.
Right?

63
00:04:20,121 --> 00:04:28,150
[laugh] And so, so in statistics there,
there's a lot of debate.

64
00:04:28,150 --> 00:04:32,145
People, when they see outliers, they know
what it is but, there's a lot of debate

65
00:04:32,145 --> 00:04:34,750
about how to exactly define what an
outlier is.

66
00:04:34,750 --> 00:04:38,920
And that's because outliers, you know,
sometimes you say that an outlier is

67
00:04:38,920 --> 00:04:42,754
something that's more than three standard
deviations from the mean.

68
00:04:42,754 --> 00:04:47,167
Now, the problem with that definition is,
the standard deviation itself is

69
00:04:47,167 --> 00:04:50,093
influenced by the outlier.
So the outlier inflates the standard

70
00:04:50,093 --> 00:04:53,078
deviation.
So when you say something that's three

71
00:04:53,078 --> 00:04:58,067
standard deviations from the mean, if you
have outliers in the data, your standard

72
00:04:58,067 --> 00:05:02,092
deviation is big and you actually might
miss some outliers by using that

73
00:05:02,092 --> 00:05:07,071
definition.
So, a definition, of an outlier that is, I

74
00:05:07,071 --> 00:05:15,056
don't want to say universally accepted but
one that is a bit, that's robust to

75
00:05:15,056 --> 00:05:21,835
outliers themselves are the following.
So, a data point that's often called a

76
00:05:21,835 --> 00:05:28,589
moderate outlier is a data point that is
smaller than the, so if you think of a, of

77
00:05:28,589 --> 00:05:34,976
a distribution, right, going through here.
And so we have the, the center of the

78
00:05:34,976 --> 00:05:40,339
distribution.
So, say we have this as the median, and

79
00:05:40,339 --> 00:05:48,353
then we have the twenty-fifth percentile,
and we have the 75th percentile here.

80
00:05:48,353 --> 00:05:52,611
And this distance here is the
interquartile range.

81
00:05:52,611 --> 00:05:57,275
The difference between the, 75th
percentile and the twenty-fifth

82
00:05:57,275 --> 00:06:00,465
percentile.
So this is the measure of the, a robust

83
00:06:00,465 --> 00:06:06,544
measure of the spread of the distribution.
So a, a moderate outlier on the right, on

84
00:06:06,544 --> 00:06:13,139
the right tail, is gonna be a data point
that is between the 75th percentile and,

85
00:06:13,139 --> 00:06:19,399
and you go out one and a half times the
in, interquartile range, and then you look

86
00:06:19,399 --> 00:06:25,219
at, but it's less than the, the 75th
percentile plus three times the

87
00:06:25,219 --> 00:06:32,498
interquartile range.
So there's a point that is a, a, sort of

88
00:06:32,498 --> 00:06:39,419
like out here which is Q.75 plus 1.5 times
the IQR.

89
00:06:39,419 --> 00:06:47,815
And then there's another point out here
that is the seventy-fifth percentile, plus

90
00:06:47,815 --> 00:06:55,970
three times the inter quartile range.
And if a data point lives in, in, in this

91
00:06:55,970 --> 00:07:02,659
area here, it's called a moderate outlier.
So it's, it's lying pretty far in the

92
00:07:02,659 --> 00:07:06,051
tail.
But if a data point is out over here, that

93
00:07:06,051 --> 00:07:10,157
if it satisfies this, then it's called an
extreme outlier.

94
00:07:10,157 --> 00:07:15,054
And similarly on the left hand side you
have the same kind of definition.
