1
00:00:00,000 --> 00:00:04,074
So, the first kind of data descriptive
Statistics is going to be a graphical

2
00:00:04,074 --> 00:00:06,096
statistics.
It's called the histogram.

3
00:00:06,096 --> 00:00:11,089
In the histogram, the idea is we want to
describe the shape of the distribution of

4
00:00:11,089 --> 00:00:15,043
the underlying data.
And so, how do we construct a histogram?

5
00:00:15,043 --> 00:00:18,066
So, we order the data from largest to
smallest.

6
00:00:18,066 --> 00:00:23,043
Then, we divide the range of the data
into, say, in equally spaced bins.

7
00:00:23,043 --> 00:00:27,850
And then, we count the number of
observations in each of the bins and then

8
00:00:27,850 --> 00:00:33,767
we create a bar graph of those counts.
And we might normalize the area so that

9
00:00:33,767 --> 00:00:36,772
it's equal to one.
And so, we're going to get a, you know,

10
00:00:36,772 --> 00:00:41,451
kind of a bar chart of the distribution of
data and we can think of this histogram as

11
00:00:41,451 --> 00:00:48,027
a very crude estimate for the underlying
probability curve, alright?

12
00:00:48,027 --> 00:00:52,359
So, in R, one of the reasons why R is very
nice is that it has all of these data

13
00:00:52,359 --> 00:00:55,476
descriptive statistics built into the
language, right?

14
00:00:55,476 --> 00:01:00,180
So, if we want to create a histogram, we
just use a function called hist and it

15
00:01:00,180 --> 00:01:04,706
produces our histogram plot for us.
And we can change the number of bends and,

16
00:01:04,706 --> 00:01:08,722
you know, do all sorts of things as an
option to the hist function.

17
00:01:08,722 --> 00:01:13,582
And so, I encourage you, you know, just to
play around with it and, you know, and,

18
00:01:13,582 --> 00:01:17,545
and see what you get.
So, let's look at some histograms of the

19
00:01:17,545 --> 00:01:25,023
underlying data and try to get an idea of
what the distribution looks like, okay?

20
00:01:25,023 --> 00:01:33,863
Now, before I do that I want to use
Gaussian white noise as a benchmark model

21
00:01:33,863 --> 00:01:38,987
for the returns on Microsoft and the
returns on the S and P 500.

22
00:01:38,987 --> 00:01:45,670
And we want to think about whether or not
the returns on Microsoft and the returns

23
00:01:45,670 --> 00:01:52,344
on S and P 500 can be considered as a
realization from a Gaussian white noise

24
00:01:52,344 --> 00:01:58,179
process, okay?
So, to help you do this thought experiment

25
00:01:58,477 --> 00:02:05,141
I've plotted up here, the monthly
continuously compounded returns on

26
00:02:05,141 --> 00:02:07,851
Microsoft.
This is the actual data.

27
00:02:07,851 --> 00:02:14,230
Now here, what I've done is, I've created
a computer simulation of, of a Gaussian

28
00:02:14,230 --> 00:02:20,514
white noise that has the same mean value
as Microsoft and it has the same

29
00:02:20,514 --> 00:02:24,422
volatility, the same standard deviation as
Microsoft.

30
00:02:24,422 --> 00:02:31,190
And I just created the 120 random draws
from that you know, Gaussian distribution,

31
00:02:31,190 --> 00:02:36,318
that's calibrated to have the same mean
and standard deviation as Microsoft.

32
00:02:36,318 --> 00:02:39,912
And so, here's the computer simulation and
here's Microsoft.

33
00:02:39,912 --> 00:02:44,425
Now again, the differences between these
simulations are, you know, here it's the

34
00:02:44,425 --> 00:02:49,840
volatility changing in the actual data is
not reflected in the Gaussian white noise

35
00:02:49,840 --> 00:02:54,333
data because this always has the same
volatility all the way throughout the

36
00:02:54,333 --> 00:02:57,362
process, okay?
So, that's one thing that's clearly

37
00:02:57,362 --> 00:03:01,109
different between the simulated data and
the actual data.

38
00:03:01,109 --> 00:03:06,671
And then, the question is well are the
things that are similar, the mean values

39
00:03:06,671 --> 00:03:09,936
of these two series are the same by
construction?

40
00:03:09,936 --> 00:03:15,356
And what about the empirical distribution?
Let's calculate the histogram of the

41
00:03:15,356 --> 00:03:20,528
actual data and calculate the histogram of
the normal data and see if they look

42
00:03:20,528 --> 00:03:26,801
similar.
So, here's a histogram plot of the monthly

43
00:03:26,801 --> 00:03:32,040
returns on Microsoft, okay?
So this is done using the hist command in

44
00:03:32,040 --> 00:03:37,028
R and I don't use any special options.
I use just all the defaults okay.

45
00:03:37,028 --> 00:03:41,027
So, this is an estimate of the underlying
probability curve.

46
00:03:41,027 --> 00:03:45,094
So, right away, one of things that you
should notice about this is that it looks

47
00:03:45,094 --> 00:03:48,071
very similar to a bell-shaped
distribution, okay?

48
00:03:48,071 --> 00:03:53,050
And so, this is one of the reasons why the
normal distribution is often used as a

49
00:03:53,050 --> 00:03:58,000
benchmark because when you look at
histograms of returns they, they resemble

50
00:03:58,000 --> 00:04:01,084
a normal distribution, okay?
Now, it's blocky because, you know, we're,

51
00:04:01,084 --> 00:04:04,044
we're, we're dividing these things into
bins.

52
00:04:04,044 --> 00:04:09,362
Now again, when we look at a histogram,
you can say, where is the histogram

53
00:04:09,362 --> 00:04:14,966
centered, that's the mean return, the
volatility is the standard deviation in

54
00:04:14,966 --> 00:04:20,220
the histogram, the skewness is the
asymmetry in the histogram, and the excess

55
00:04:20,220 --> 00:04:25,667
kurtosis is telling us about how thick the
tails are relative to a normal

56
00:04:25,667 --> 00:04:30,716
distribution, okay?
So, here, you know, if we think there

57
00:04:30,716 --> 00:04:35,761
might be a slight long left tail.
So, if we're going to guess, it might have

58
00:04:35,761 --> 00:04:40,492
a negative skewness.
And the tails, you know, kind of are, you

59
00:04:40,492 --> 00:04:44,888
know, well it's hard to say.
Are they fatter than the normal

60
00:04:44,888 --> 00:04:48,619
distribution?
It's hard to tell from the histogram by

61
00:04:48,619 --> 00:04:52,287
itself.
But the fact that there is this kind of

62
00:04:52,287 --> 00:04:58,299
big, outlying value, is a clue that
perhaps the kurtosis is bigger than three.

63
00:04:58,533 --> 00:05:02,047
Okay.
Now, here is a histogram of the computer

64
00:05:02,047 --> 00:05:06,892
simulated Gaussian data, okay?
This thing should look like a normal

65
00:05:06,892 --> 00:05:12,030
curve, because it actually generated from
a normal distribution, okay?

66
00:05:12,030 --> 00:05:18,017
It has the same mean as the returns on
Microsoft and the volatility is the same

67
00:05:18,017 --> 00:05:22,065
volatility of Microsoft and notice the,
the range of the data.

68
00:05:22,065 --> 00:05:29,127
So, it goes from, you know, -two to +two.
And, if we look at the actual Microsoft

69
00:05:29,127 --> 00:05:34,562
data, this is going out to -0.4 to 0.4.
So, it's looking like the Microsoft

70
00:05:34,562 --> 00:05:40,521
returns have fatter tails than, than the
Gaussian data, because of the, how far

71
00:05:40,521 --> 00:05:44,507
they're going out in the actual data,
okay?

72
00:05:44,507 --> 00:05:51,196
Now, if you look at the S and P 500.
So, here's the again, we use the hist

73
00:05:51,196 --> 00:05:56,855
command on the S and P 500 monthly returns
and so it's, it's pretty, I mean, the

74
00:05:56,855 --> 00:06:02,196
histogram is pretty blocky here, you know,
the default I could have probably used a

75
00:06:02,196 --> 00:06:06,055
few more bins.
But I think the striking feature of the

76
00:06:06,055 --> 00:06:10,163
histogram is the fact that it has a pretty
clear, long, left tail.

77
00:06:10,163 --> 00:06:15,220
So, there's a pretty clear negative
skewness in this distribution and so, you

78
00:06:15,220 --> 00:06:19,463
know, if you think, is the normal
distribution a good model for the S and P

79
00:06:19,463 --> 00:06:24,018
500 returns, probably not, you know,
because of this, this big negative

80
00:06:24,018 --> 00:06:32,567
skewness.
Now if you want to compare Microsoft to

81
00:06:32,567 --> 00:06:38,516
the S and P 500, we would like to use the
same x-axis scale for the two return

82
00:06:38,516 --> 00:06:42,001
series.
And so, what I did was I first computed

83
00:06:42,001 --> 00:06:48,004
the histogram for Microsoft, and then I
used the bin ranges for Microsoft to

84
00:06:48,004 --> 00:06:53,751
compute the histogram for the S and P 500
and the reason for doing that is you can

85
00:06:53,751 --> 00:06:59,735
see now that the S and P 500, the data or
much more concentrated around zero than

86
00:06:59,735 --> 00:07:03,202
for Microsoft.
Microsoft has a much larger spread above

87
00:07:03,202 --> 00:07:07,520
the average than the S and P 500.
And so, that's just another way of saying

88
00:07:07,520 --> 00:07:12,801
that the volatility of Microsoft is higher
than the volatility than, than the S and P

89
00:07:12,801 --> 00:07:17,054
500.
I made the comment that I thought the

90
00:07:17,054 --> 00:07:21,732
kurtosis of this was, was bigger than
three, that the kurtosis is bigger than

91
00:07:21,732 --> 00:07:25,022
the standard normal.
And the reason is because the, in the

92
00:07:25,022 --> 00:07:29,335
actual data, notice that this is, the left
tail is -40%, the right tail is +40%.

93
00:07:29,335 --> 00:07:36,968
When I did the computer simulated Gaussian
data, the right tail is only going to

94
00:07:36,968 --> 00:07:40,697
-twenty percent and +twenty%.
So, the Microsoft returns are out here in

95
00:07:40,697 --> 00:07:44,288
the left, left tail, and are out here in
the right tail.

96
00:07:44,288 --> 00:07:49,789
So, the actual data produce more extreme
values than the computer simulation from a

97
00:07:49,789 --> 00:07:54,086
normal distribution.
So, that's why I'm saying it's probably

98
00:07:54,086 --> 00:07:59,661
the case that the kurtosis of the actual
Microsoft data is going to be bigger than

99
00:07:59,661 --> 00:08:04,129
three.
Alright.

100
00:08:04,129 --> 00:08:12,078
So, when we look at histograms, his
histograms are blocky, okay?

101
00:08:12,078 --> 00:08:19,221
Now, if we want to eliminate the
blockiness of a histogram, we can compute

102
00:08:19,221 --> 00:08:24,526
what's called a smoothed histogram.
And so, this is a graph, of a smoothed, of

103
00:08:24,526 --> 00:08:29,908
the smoothed histogram and it's created
with the R function called, R function

104
00:08:29,908 --> 00:08:33,837
called density.
And another name for a smoothed histogram

105
00:08:33,837 --> 00:08:39,026
is a kernel density estimate, okay?
And literally, what you're doing, you

106
00:08:39,026 --> 00:08:46,579
know, if you think about it, think about
drawing a smooth curve over the data,

107
00:08:46,579 --> 00:08:49,441
that's what the kernel density algorithm
does.

108
00:08:49,441 --> 00:08:53,018
And how do you get a smooth curve
throughout the data?

109
00:08:53,018 --> 00:08:57,929
Well, at a given point, what you do is you
take an kind of a weighted average of

110
00:08:57,929 --> 00:09:03,097
values around a particular point where the
weights you know, kind of decline going

111
00:09:03,097 --> 00:09:07,113
out in this direction and decline going
out in that direction.

112
00:09:07,113 --> 00:09:12,345
And then essentially, what you do is you,
you take that kind of smooth and you pass

113
00:09:12,345 --> 00:09:17,327
it all the way across the data and that
allows you to draw this kind of smooth

114
00:09:17,327 --> 00:09:21,861
curve over the histogram.
The, the algorithm for computing the

115
00:09:21,861 --> 00:09:29,412
smooth histogram, this kernel density
estimate, this is taken from Rupert is of

116
00:09:29,412 --> 00:09:33,086
the following form.
So, if you want to estimate the, the

117
00:09:33,086 --> 00:09:40,493
probability curve at a single point, then
what we do is we take a weighted average

118
00:09:40,493 --> 00:09:44,577
where the weights are determined by what's
called a kernel function.

119
00:09:44,577 --> 00:09:49,511
And a kernel function is very often just a
symmetric probability distribution.

120
00:09:49,511 --> 00:09:53,542
And usually, the kernel function is the
standard normal curve.

121
00:09:53,542 --> 00:09:58,615
And essentially, what you're doing is
you're, you're taking a weighted average

122
00:09:58,615 --> 00:10:03,637
of, of the values in a histogram where
you're weighting the points with a normal

123
00:10:03,637 --> 00:10:06,500
distribution before and after the given
point.

124
00:10:06,500 --> 00:10:12,005
And that allows you to smooth the
blockiness of, of, of the histogram, okay?

125
00:10:12,005 --> 00:10:17,007
The, when you use this kind of function
there's a, a parameter B that's called a

126
00:10:17,007 --> 00:10:21,001
bandwidth parameter and that controls the
degree of smoothing.

127
00:10:21,001 --> 00:10:26,009
So, if B is big, then you get a lot of
smoothing and if B is small then you don't

128
00:10:26,009 --> 00:10:32,073
get very much smoothing, okay?
So, anyway, for our purposes, I just want

129
00:10:32,073 --> 00:10:38,073
you to know that sometimes, it's, it's
easier to look at the smooth version of

130
00:10:38,073 --> 00:10:43,043
the histogram to get a, a general view of
the shape, instead of the histogram by

131
00:10:43,043 --> 00:10:45,620
itself.
And so, notice you get, when you do the

132
00:10:45,620 --> 00:10:49,799
smooth, you get some sort of wiggles out
here, you get like little wiggle over

133
00:10:49,799 --> 00:10:52,306
here.
And so, you know, does this look like a

134
00:10:52,306 --> 00:10:55,604
normal curve?
Well, not exactly, you know, it's a, a, a

135
00:10:55,604 --> 00:11:00,086
nice normal curve would be perfect
symmetric and bell-shaped, whereas you

136
00:11:00,086 --> 00:11:03,048
know, this is a little bit less, less
perfect, okay?

137
00:11:03,048 --> 00:11:07,074
So, we can overlay, for example, the
smooth histogram on top of the, the

138
00:11:07,074 --> 00:11:11,047
histogram from Microsoft.
And so, we see how the smooth, you know,

139
00:11:11,047 --> 00:11:15,084
sort of picks up this little bump, picks
up this little bump here, you know,

140
00:11:15,084 --> 00:11:19,549
captures the basic shape of the
distribution in the middle, gets that

141
00:11:19,549 --> 00:11:22,082
little bump here from these two things,
and so on.

142
00:11:22,082 --> 00:11:27,032
So, we often, you know, will report a
histogram with the little smooth line on

143
00:11:27,032 --> 00:11:30,039
top of it.
And again, you just want to get an idea of

144
00:11:30,039 --> 00:11:33,012
what the shape of the distribution looks
like.
