1
00:00:00,000 --> 00:00:04,066
So at the end of last question, last
lecture, we were talking about outliers in

2
00:00:04,066 --> 00:00:07,042
data.
And an outlier in the data is something

3
00:00:07,042 --> 00:00:11,034
that's unusual, doesn't follow the, the
normal pattern of the data.

4
00:00:11,034 --> 00:00:16,012
So this is a little cartoon that shows,
you know, most of the data points fall in

5
00:00:16,012 --> 00:00:20,050
this nice curvilinear relationship, but
there's one data point that's off.

6
00:00:20,050 --> 00:00:25,007
And so this would typically be called an
outlier, because it's, it's unusual.

7
00:00:25,007 --> 00:00:28,005
And outliers could happen because of data
mistakes.

8
00:00:28,005 --> 00:00:32,092
Or outliers could happen because there's
really something special about this point,

9
00:00:32,092 --> 00:00:36,049
such that, it's, its, you know?
It's not following the, the usual

10
00:00:36,049 --> 00:00:39,071
relationship.
And perhaps we ought to understand better

11
00:00:39,071 --> 00:00:44,004
what's, what's happening at this point.
In financial data, very often outliers

12
00:00:44,004 --> 00:00:48,098
are, observations that are associated with
market crashes, you know like the 1987

13
00:00:48,098 --> 00:00:53,086
stock market crash or something like that,
that's the day when the returns dropped

14
00:00:53,086 --> 00:00:58,060
twenty percent in a day, and then you
know, you see something, you know, like,

15
00:00:58,060 --> 00:01:03,019
like observations in the extreme tales of
the distribution or if you're looking at

16
00:01:03,019 --> 00:01:07,061
an individual stock you might see all of a
sudden, you know one day you know the

17
00:01:07,061 --> 00:01:12,036
return is big and negative and there might
have been a very bad earnings announcement

18
00:01:12,036 --> 00:01:16,023
or something like that and that could be
driving the observation down.

19
00:01:16,023 --> 00:01:20,026
So sometimes we can come up with an
explanation for an outlier, sometimes we

20
00:01:20,026 --> 00:01:23,094
can't.
From the point of view of descriptive

21
00:01:23,094 --> 00:01:30,016
statistics, outliers do have a impact on
sample statistics like the mean, the

22
00:01:30,016 --> 00:01:33,057
standard deviation, the kurtosis, the
skewness.

23
00:01:33,080 --> 00:01:40,047
And, and so when you look at statistics
computed with and without outliers you can

24
00:01:40,047 --> 00:01:46,076
actually get very, very different results.
So let me illustrate this a little bit.

25
00:01:46,076 --> 00:01:51,056
So here I took the monthly returns on
Microsoft and I created an outlier.

26
00:01:51,056 --> 00:01:56,092
So I created a negative return that was
you know, a continuously compounded return

27
00:01:56,092 --> 00:02:00,001
space, was -80%.
And so notice that, you know, this

28
00:02:00,001 --> 00:02:03,048
observation is very, very different from,
from the others.

29
00:02:03,048 --> 00:02:07,064
It's sort of like, you know, very bad news
happened on, on this date.

30
00:02:07,064 --> 00:02:12,063
And if we look at the histogram, we see
that, that outlier showing up in this,

31
00:02:12,063 --> 00:02:15,957
this big, you know, tail.
And, and one question that you know, we

32
00:02:15,957 --> 00:02:19,825
often have about outliers is what is the
impact of the outlier on, on our

33
00:02:19,825 --> 00:02:23,832
descriptive statistics.
We see that the histogram now has a bump

34
00:02:23,832 --> 00:02:27,104
way out over here.
And then another question is.

35
00:02:27,104 --> 00:02:32,315
What happens to you know, the sample mean,
the sample standard deviation, skewness

36
00:02:32,315 --> 00:02:36,047
and kurtosis?
You know, when the outlier is included in

37
00:02:36,047 --> 00:02:43,392
the data and when it is not.
So, I did some simple calculations, where,

38
00:02:43,392 --> 00:02:50,226
essentially, in blue here, I have sample
statistics that do not include the

39
00:02:50,226 --> 00:02:54,613
outlier.
So the mean of Microsoft returns when the

40
00:02:54,613 --> 00:03:01,770
outlier was not included is, you know,
about .67%.

41
00:03:01,770 --> 00:03:09,336
In black, on the right hand side, I
recompute the sample mean where I include

42
00:03:09,336 --> 00:03:14,070
that big, negative outlier.
And now when I compute the sample mean

43
00:03:14,070 --> 00:03:18,062
the, it turns out, goes from a positive
number to a negative number.

44
00:03:18,062 --> 00:03:24,085
And, and it's you know, so we see that the
one big negative outlier has pulled the

45
00:03:24,085 --> 00:03:28,003
mean from positive .67 percent to, to a
negative number.

46
00:03:28,003 --> 00:03:33,035
You know, and again this is, you know, if
we're looking using the mean as a, our

47
00:03:33,035 --> 00:03:38,033
best guess for the expected monthly
return, you know, here we're actually

48
00:03:38,033 --> 00:03:41,086
making money, now we're losing money on
average, 'kay?

49
00:03:41,086 --> 00:03:47,046
So we see that the big impact of, of one
observation on the interpretation of the

50
00:03:47,067 --> 00:03:51,079
desirability of this asset.
Similarly we can look at the sample

51
00:03:51,079 --> 00:03:54,087
standard deviation which measures the
volatility.

52
00:03:54,087 --> 00:04:00,066
When you don't have the outliers ten%,
when you do have the outliers almost

53
00:04:00,066 --> 00:04:04,006
fourteen%, 13.7%.
So we see that one observation has

54
00:04:04,006 --> 00:04:07,821
inflated the standard deviation, you know,
30%.

55
00:04:07,821 --> 00:04:12,591
And so that's a, that's a big impact of,
of one observation.

56
00:04:12,591 --> 00:04:18,053
And then skewness without the outlier is
-.07, with the outlier is -2.3.

57
00:04:18,053 --> 00:04:22,827
So the big negative return obviously has
given us a big negative skewness.

58
00:04:22,827 --> 00:04:26,539
And then our kurtosis, or actually this is
excess kurtosis.

59
00:04:26,539 --> 00:04:31,410
Without the outlier it's 1.8, with the
outlier is fourteen.

60
00:04:31,410 --> 00:04:37,040
So we see that, you know, the impact of
this outlier to, you know, it can greatly

61
00:04:37,040 --> 00:04:42,732
change these, these sample statistics, and
so then again, you know, the question is

62
00:04:42,732 --> 00:04:47,504
do you keep the outlier in or not.
Kind of depends upon the purpose at, at,

63
00:04:47,504 --> 00:04:51,254
of, of what you're doing.
If you're worried about tail risk, and if

64
00:04:51,254 --> 00:04:55,525
that big negative outlier really
represents something that, you know, could

65
00:04:55,525 --> 00:05:00,623
happen again, then you would probably want
to keep it in to get a, a better estimate

66
00:05:00,623 --> 00:05:06,756
for, you know, you know, a loss that, that
could happen due to a big negative return.

67
00:05:06,756 --> 00:05:11,623
On the other hand, if you know, you're,
you're interested in perhaps in, in

68
00:05:11,623 --> 00:05:16,007
estimating the expected return, Most of
the time, Microsoft returns is positive.

69
00:05:16,007 --> 00:05:19,094
Every now and then, you get a negative
return, but on average, you know, you

70
00:05:19,094 --> 00:05:24,004
wanna know what the expected gain is.
Then maybe if we're estimating the mean,

71
00:05:24,004 --> 00:05:27,043
you, you might want to down weight that
outlier in a computation.

72
00:05:30,092 --> 00:05:36,097
Now, there's a nice graphical, summary
statistic for a distribution that also

73
00:05:36,097 --> 00:05:41,072
highlights outliers in the data, and this
is called a box plot.

74
00:05:41,072 --> 00:05:47,024
And a box plot is a, you know, again,
something that's created, oh, I should,

75
00:05:47,024 --> 00:05:50,098
actually, in.
And a lot of the scripts in statistics

76
00:05:50,098 --> 00:05:55,062
have their origins from a statistician
whose name is John Tukey.

77
00:05:55,062 --> 00:06:00,087
And he was very much about looking at data
and trying to create ways of, of

78
00:06:01,007 --> 00:06:04,053
summarizing data, and in, in very nice and
nifty ways.

79
00:06:04,053 --> 00:06:09,077
And so, he's the father of the box plot.
And the idea of the box plot is you, you

80
00:06:09,077 --> 00:06:16,032
show the basic features of a distribution
of one-dimensional data, and You, You

81
00:06:16,032 --> 00:06:21,022
illustrate features of the data using
sample statistics that are robust to

82
00:06:21,022 --> 00:06:24,017
outliers.
So for showing the center of the data

83
00:06:24,017 --> 00:06:29,006
instead of using the mean you'll use the
median, because the median is less

84
00:06:29,006 --> 00:06:31,082
sensitive to outliers than, than is the
mean.

85
00:06:31,082 --> 00:06:36,022
And similarly to show the spread in the
data instead of using the standard

86
00:06:36,022 --> 00:06:41,009
deviation, you would use the interquartile
range, the difference essentially between

87
00:06:41,009 --> 00:06:45,055
the first quartile and the third quartile,
as a measure of the middle of the

88
00:06:45,055 --> 00:06:50,006
distribution, and again the quartiles are
less sensitive to outliers than the

89
00:06:50,006 --> 00:06:54,087
standard deviation so that gives you what
they say is a more robust measure of

90
00:06:54,087 --> 00:06:59,019
spread.
And then there are outer fences that are

91
00:06:59,019 --> 00:07:05,056
shown in the box plot that illustrate a
moderate outlier and an extreme outlier.

92
00:07:05,056 --> 00:07:11,068
And last time I sort of said there was a,
these definitions of outliers that are

93
00:07:11,068 --> 00:07:17,065
based on, on percentiles and a moderate
outlier or something where you look at

94
00:07:17,065 --> 00:07:22,201
the, the 75th percentile, and you look at
an observation that's beyond the 75th

95
00:07:22,201 --> 00:07:26,835
percentile, but less than the 75th
percentile plus three times the

96
00:07:26,835 --> 00:07:31,013
interquartile range.
And the interquartile range is just

97
00:07:31,013 --> 00:07:34,040
different here.
So, something out here is a moderate

98
00:07:34,040 --> 00:07:37,054
outlier.
And then an extreme outlier is something

99
00:07:37,054 --> 00:07:39,072
that's even farther out like this.
