1
00:00:00,000 --> 00:00:07,140
So the box plot gives you these kind of
robust measures of a distribution.

2
00:00:07,140 --> 00:00:13,037
So here's an illustration of the box plot
for the monthly returns on Microsoft.

3
00:00:13,037 --> 00:00:17,609
And so this black line in the middle
represents the, the median.

4
00:00:17,609 --> 00:00:23,064
This is roughly the interquartile range.
It's not exactly it and then if you look

5
00:00:23,064 --> 00:00:26,550
at the online help for a box plot, you'll
see why.

6
00:00:26,550 --> 00:00:31,589
Notice that this is not quite symmetric.
So by definition, the median is in the

7
00:00:31,589 --> 00:00:37,565
middle of the interquartile range, but
This is so, so that these, these fences

8
00:00:37,565 --> 00:00:41,262
are, the, the, the edges of the box are
not quite the, the.

9
00:00:41,262 --> 00:00:47,021
The first and the third quartile, and then
you have, You know, essentially these

10
00:00:47,021 --> 00:00:52,434
fences that, illustrate, you know, when
you have extreme outliers in the data.

11
00:00:52,434 --> 00:00:57,979
So these points out over here, essentially
are representing those values that are,

12
00:00:57,979 --> 00:01:02,158
you know, what, what we characterize as an
extreme outlier.

13
00:01:02,158 --> 00:01:06,572
And this would be in the right tail, and
this would be in the left tail.

14
00:01:06,572 --> 00:01:10,282
So you have a, the bulk of the
distribution is roughly symmetric.

15
00:01:10,282 --> 00:01:14,837
You have two big negative outliers and you
have a few big positive outliers.

16
00:01:14,837 --> 00:01:19,405
And again, if you look at the online help
for the box plot, ya know, they'll tell

17
00:01:19,405 --> 00:01:22,764
you exactly how this fence and this fence
is determined.

18
00:01:22,764 --> 00:01:30,595
From the point of view of interpretation,
the big power of box plots comes when you

19
00:01:30,595 --> 00:01:34,664
compare multiple series.
Because it allows you to, you know, very

20
00:01:34,664 --> 00:01:39,326
quickly look at the distribution of
multiple series in the same plot.

21
00:01:39,326 --> 00:01:44,384
So, here's a, a box plot where I have the
box plot of the Gaussian white noise.

22
00:01:44,384 --> 00:01:50,009
This is the simulated data that comes from
a normal distribution that has the same

23
00:01:50,009 --> 00:01:55,445
mean invariance as the Microsoft returns.
These are the, this is the box plot for

24
00:01:55,445 --> 00:01:59,986
the actual Microsoft returns and this is
the box plot for the SMP 500 returns.

25
00:01:59,986 --> 00:02:03,552
And so we see that they're on the same
scale which is percent.

26
00:02:03,552 --> 00:02:08,623
So the average return for all of these
things is pretty close to zero and a

27
00:02:08,623 --> 00:02:12,048
little bit positive.
So, that's the median return.

28
00:02:12,048 --> 00:02:18,151
We see that you know, again, Microsoft and
Guassian White Noise, the middle of the

29
00:02:18,151 --> 00:02:23,874
distribution is, is roughly the same.
And, what we see with the Microsoft data

30
00:02:23,874 --> 00:02:28,488
is, we have some outliers here.
With the Gaussian white noise data we

31
00:02:28,488 --> 00:02:33,441
don't have the outliers because normal
distribution doesn't really produce

32
00:02:33,441 --> 00:02:37,552
outliers.
And with the S and P 500 data, we see the

33
00:02:37,552 --> 00:02:43,774
smaller spread than Microsoft, and we see
the small two negative outliers.

34
00:02:43,774 --> 00:02:49,762
So again, we can get sort of an idea of
the shape of the distribution.

35
00:02:49,762 --> 00:02:56,824
So it's similar to the histogram, but
again it's based on robust measures, and

36
00:02:56,824 --> 00:03:05,050
it's quite useful for doing quick
comparisons of many series at once.

37
00:03:05,050 --> 00:03:14,802
Now, I've put together a little four graph
summary for looking at the distribution of

38
00:03:14,802 --> 00:03:18,619
asset returns.
And this four graph summary is borrowed

39
00:03:18,619 --> 00:03:23,863
from a very nice book by Rene Carmona
called The Statistical Analysis of

40
00:03:23,863 --> 00:03:27,629
Financial Data.
And in this four graph summary there's

41
00:03:27,629 --> 00:03:32,604
going to be a histogram, a box plot, a
smooth histogram, and then a QQ plot

42
00:03:32,604 --> 00:03:37,621
relative to a normal.
So we have four pictures and it looks like

43
00:03:37,621 --> 00:03:40,403
this.
So, say for Microsoft, here we have the

44
00:03:40,403 --> 00:03:45,201
histogram, here we have the smooth density
down here so you can get a rough shape of

45
00:03:45,201 --> 00:03:48,605
the distribution.
And then, we have the box plot over here,

46
00:03:48,605 --> 00:03:53,817
and the QQ plot relative to the normal.
And all of these, these three plots are

47
00:03:53,817 --> 00:03:57,831
giving you similar information.
We see kind of a long left tail, a

48
00:03:57,831 --> 00:04:01,412
negative skewness, so we see a little long
left tail over here.

49
00:04:01,412 --> 00:04:06,720
We see the negative outliers in the box
plot, which is corresponding to these

50
00:04:06,720 --> 00:04:10,479
observations here.
And in the Q-Q plot relative to normal,

51
00:04:10,479 --> 00:04:15,189
we're seeing this drop-down relative to
the normal graph, and this dipping up a

52
00:04:15,189 --> 00:04:19,477
little bit on the right hand side.
So if you look at, you know, a, a p-, a

53
00:04:19,477 --> 00:04:24,170
summary graph like this, you can get an
idea that Microsoft returns are.

54
00:04:24,170 --> 00:04:29,278
Kind of approximately normally distributed
in terms of the bulk of the distribution

55
00:04:29,278 --> 00:04:32,115
looks normal.
But there is some negative skewness and

56
00:04:32,115 --> 00:04:36,937
fatter tales relative to the normal
distribution.

57
00:04:36,937 --> 00:04:44,787
Okay.
Alright.

58
00:04:44,787 --> 00:04:50,509
So when we have two or more random
variables, then, we have, we might want to

59
00:04:50,509 --> 00:04:55,214
look at some descriptive statistics that
tell us the relationship between two or

60
00:04:55,214 --> 00:04:58,047
more variables.
Now, when we studied probability theory,

61
00:04:58,047 --> 00:05:02,485
when we were looking at the dependence
between two random variables, we looked at

62
00:05:02,485 --> 00:05:07,171
covariance and correlation.
Now, from the point of view of descriptive

63
00:05:07,171 --> 00:05:11,710
statistics, we can, we have, say we have
two random variables X and Y.

64
00:05:11,710 --> 00:05:14,284
And then we observe the sample X1, Y1, X2,
Y2.

65
00:05:14,284 --> 00:05:19,569
So, we can think of X as the return on
Microsoft, Y as the return on the S and P

66
00:05:19,569 --> 00:05:23,728
500; and then our sample is the data that
we download from Yahoo.

67
00:05:23,728 --> 00:05:26,651
Okay?
We assume that these returns, you know,

68
00:05:26,651 --> 00:05:30,125
follow a multi-, a bi-variant normal
distribution.

69
00:05:30,125 --> 00:05:33,702
For example, as, as being kind of the
benchmark.

70
00:05:33,702 --> 00:05:39,012
And then, say we wanna measure the
dependence between the Microsoft returns,

71
00:05:39,012 --> 00:05:44,204
and the S and P 500 returns.
We'd wanna compute the, sample covariance

72
00:05:44,204 --> 00:05:49,962
and the sample correlation, to get, an
idea of what the data say the relationship

73
00:05:49,962 --> 00:05:54,039
between these returns are, okay.
Now, in, looking at pairwise

74
00:05:54,039 --> 00:05:59,465
relationships, there's a graphical,
diagnost-, a graphical descriptive

75
00:05:59,465 --> 00:06:04,069
statistic called a scatter plot.
And a scatter plot is just an x-y plot of

76
00:06:04,069 --> 00:06:08,040
your bi-variant data.
So you can plot the returns on the SMP500

77
00:06:08,040 --> 00:06:11,055
on one axis and returns of Microsoft on
the other.

78
00:06:11,055 --> 00:06:14,081
And then you can see what the data looks
like.

79
00:06:14,081 --> 00:06:19,008
So for example.
Here's a scatter plot of the monthly

80
00:06:19,008 --> 00:06:21,086
returns on Microsoft versus the S and P
500.

81
00:06:21,086 --> 00:06:26,029
So here, I, I put Microsoft on the X axis.
The S and P 500 on the Y axis.

82
00:06:26,029 --> 00:06:30,096
And the black lines here represent.
This is the mean for Microsoft, this is

83
00:06:30,096 --> 00:06:33,011
the mean for the S and P 500.
Okay?

84
00:06:33,011 --> 00:06:37,066
So remember when we, we studied
co-variance and correlation, we, we looked

85
00:06:37,066 --> 00:06:40,088
at these probability scatter plots.
And you know?

86
00:06:40,088 --> 00:06:43,092
Essentially, when you look at this data,
you know?

87
00:06:43,092 --> 00:06:48,072
What would you say, are, is there a
positive or negative relationship between

88
00:06:48,072 --> 00:06:51,676
Microsoft and the S and P 500?
Positive, right?

89
00:06:51,676 --> 00:06:57,040
Cuz as Microsoft returns go up, S and P
500 returns tend to go up as well.

90
00:06:57,040 --> 00:07:02,057
As the Microsoft returns go down, the S
and P 500 returns are going down.

91
00:07:02,057 --> 00:07:05,051
Okay?
So here we see, we see a negative

92
00:07:05,051 --> 00:07:09,017
relationship.
And later on, we'll compute the sample

93
00:07:09,017 --> 00:07:13,055
co-variance and correlation.
And we, there's a positive sample

94
00:07:13,055 --> 00:07:16,071
co-variance and the sample correlation
is.6.

95
00:07:16,071 --> 00:07:22,037
So again, there's a reasonably strong
positive linear association between these

96
00:07:22,037 --> 00:07:23,060
two returns.
Okay?

97
00:07:26,025 --> 00:07:31,039
Now a.
When you have more than two returns,

98
00:07:31,039 --> 00:07:35,010
there's a nifty function in R called
pairs, P, A, I, R, S.

99
00:07:35,010 --> 00:07:39,093
And what the pairs function does, it
creates all pair y scatter plots.

100
00:07:39,093 --> 00:07:45,025
So here I have three data series, my
computer-simulated Gaussian white noise,

101
00:07:45,025 --> 00:07:49,017
the returns on Microsoft, and the returns
in S and P 500.

102
00:07:49,017 --> 00:07:52,095
And so now I have what's called a scatter
plot matrix.

103
00:07:52,095 --> 00:07:58,097
And, so this graph right here, this is the
scatter plot with Gaussian white noise on

104
00:07:58,097 --> 00:08:03,087
this axis and Microsoft on this axis.
And we see the scatter plot looks like a

105
00:08:03,087 --> 00:08:07,014
shock and blast.
There appears to be no linear association

106
00:08:07,014 --> 00:08:10,025
between them.
So we would assume that the covariance is

107
00:08:10,025 --> 00:08:14,019
close to zero and the correlation is close
to zero, based on this plot.

108
00:08:14,019 --> 00:08:21,033
This plot represents Gaussian White Noise
on this axis, and the S and P 500 on this

109
00:08:21,033 --> 00:08:22,019
axis.
Okay?

110
00:08:22,019 --> 00:08:25,687
And, again, this plot also kinda looks
like a s-, a shotgun blast, where it

111
00:08:25,687 --> 00:08:30,710
doesn't appear to be any systematic
positive or negative relationship in the

112
00:08:30,710 --> 00:08:33,630
data.
If anything, you know, there might be what

113
00:08:33,630 --> 00:08:38,102
might, what looks, maybe, to be a slight
negative relationship in, in the data.

114
00:08:38,102 --> 00:08:42,907
But it certainly isn't very strong.
So one might expect this sample covariance

115
00:08:42,907 --> 00:08:47,147
to be slightly negative, and the
correlation to be a negative number

116
00:08:47,147 --> 00:08:52,469
that's, but pretty close to zero.
This plot over here is just, we put

117
00:08:52,469 --> 00:08:57,928
Microsoft on this axis and Gaussian white
noise on this axis.

118
00:08:57,928 --> 00:09:03,109
So this plot is a, just slipping the axis
from this plot, okay?

119
00:09:03,109 --> 00:09:08,895
And similarly, this, this plot down here
is flipping this plot.

120
00:09:08,895 --> 00:09:14,674
And then finally our last plot shows
Microsoft on this axis, SMP 500 on this

121
00:09:14,674 --> 00:09:17,337
axis.
So that's what I showed you before.

122
00:09:17,337 --> 00:09:22,948
And we see positive linear relationships.
So, these two series are positively

123
00:09:22,948 --> 00:09:27,099
correlated positive covariance and a
positive correlation.

124
00:09:27,099 --> 00:09:29,524
So again, this is nice if you have ten
assets.

125
00:09:29,524 --> 00:09:34,213
You can do this and then you can get a
very quick summary of what appears to be

126
00:09:34,213 --> 00:09:36,647
the linear association between the
variables.

127
00:09:36,647 --> 00:09:39,655
And you can also I mean again, cause these
are plots.

128
00:09:39,655 --> 00:09:43,875
Even if there's a non linear association
you know, you would see, you could

129
00:09:43,875 --> 00:09:52,764
possibly see that in, in the plot itself.
So, We often summarize the sample

130
00:09:52,764 --> 00:09:59,916
variances and covariances in a matrix.
And actually before I get there let me

131
00:09:59,916 --> 00:10:04,488
define the sample statistics.
So if you wanna compute the sample

132
00:10:04,488 --> 00:10:09,615
covariance, then you would take your
sample of your first data series, the

133
00:10:09,615 --> 00:10:14,876
sample of your second data series.
You compute this sample average of x minus

134
00:10:14,876 --> 00:10:19,449
its mean, times y minus its mean.
And that's sometimes this is called at,

135
00:10:19,449 --> 00:10:22,276
little s with a subscript xy, or sigma hat
xy.

136
00:10:22,276 --> 00:10:27,906
So this is the sample covariance and then
the sample correlation is the sample

137
00:10:27,906 --> 00:10:32,989
covariance divided by the product of the
sample standard deviations.

138
00:10:32,989 --> 00:10:38,762
So if we're working in R, the bar function
computes sample covariance matrix.

139
00:10:38,762 --> 00:10:44,556
So if you so if we have three assets, and
we have a three by three matrix, then have

140
00:10:44,556 --> 00:10:49,394
the variances along the diagonals and the
covariances on the off-diagonals.

141
00:10:49,394 --> 00:10:55,113
That's what the bar function computes.
Notice that there isn't a, actually, there

142
00:10:55,113 --> 00:11:00,655
is a co function, but since with the co
function, does it same thing as, as a bar.

143
00:11:00,655 --> 00:11:06,866
The core function COR, gives you the
sample correlation matrix.

144
00:11:06,866 --> 00:11:10,429
So, here.
So if I take my three data series, my

145
00:11:10,429 --> 00:11:13,783
Gaussian white noise, my Microsoft, and my
S and P 500.

146
00:11:13,783 --> 00:11:18,442
I use the bar command.
And then this is the variance/covariance

147
00:11:18,442 --> 00:11:21,347
matrix.
The sample variances are along the

148
00:11:21,347 --> 00:11:24,346
diagonals.
And the covariances are on the off

149
00:11:24,346 --> 00:11:27,430
diagonals.
So, so notice that we see a negative

150
00:11:27,430 --> 00:11:32,563
sample covariance between Microsoft and
Gaussian white noise and a negative

151
00:11:32,563 --> 00:11:37,343
covariance between Microsoft, sorry,
between the S and P 500 and the Gaussian

152
00:11:37,343 --> 00:11:40,391
white noise.
And we have a positive covariance between

153
00:11:40,391 --> 00:11:44,553
the S and P 500 and Microsoft.
Now, covariance is just direction of

154
00:11:44,553 --> 00:11:49,391
association, correlation of strength.
So when you look at the correlation

155
00:11:49,391 --> 00:11:54,269
matrix, you have ones along the diagonal
because the correlation with each series

156
00:11:54,269 --> 00:11:58,676
with itself by definition is one.
And so here's the correlation between the

157
00:11:58,676 --> 00:12:06,135
Gaussian white noise and the Microsoft,
that's -.19 Not very strong.

158
00:12:06,135 --> 00:12:10,238
Correlation between Gaussian white noise
and S and P 500 is -.24.

159
00:12:10,238 --> 00:12:14,347
Now this is a completely spurious
correlation, because the Gaussian white

160
00:12:14,347 --> 00:12:16,421
noise was computer generated.
Right?

161
00:12:16,421 --> 00:12:21,196
It has nothing to do with the actual S and
P 500 data, but we still had a number

162
00:12:21,196 --> 00:12:24,800
that's, you know, kinda large and
negative, and that's just by chance.

163
00:12:24,800 --> 00:12:27,604
Right?
If I simulated another Gaussian white

164
00:12:27,604 --> 00:12:32,132
noise, then this correlation, probably
close to zero or could be a little

165
00:12:32,132 --> 00:12:36,566
positive, or something like that.
The Microsoft and S and P 500 data has a

166
00:12:36,566 --> 00:12:42,032
reasonably strong correlation, that's .6.
The last descriptive statistic I want to

167
00:12:42,032 --> 00:12:45,722
talk about is, has to do with time series
dependence.

168
00:12:45,722 --> 00:12:51,085
So again we were covering probability
theory, we went through the time series

169
00:12:51,085 --> 00:12:56,653
concepts section, and we were trying to
think about time dependence in data.

170
00:12:56,653 --> 00:13:00,557
And so we defined what are called
autocorrelations.

171
00:13:00,557 --> 00:13:06,305
We defined the covariance between y(t) and
y(t-1), the correlation between y(t) and

172
00:13:06,305 --> 00:13:09,798
y(t-1) as well.
Well, we can compute sample versions of

173
00:13:09,798 --> 00:13:13,020
these quantities.
So given, the time series of data.

174
00:13:13,020 --> 00:13:18,270
If, and, and if we're interested in
determining if there's linear dependence

175
00:13:18,270 --> 00:13:21,712
over time.
Then we can compute the sample auto

176
00:13:21,712 --> 00:13:26,686
covariance and the sample auto
correlation, and see if these things are

177
00:13:26,686 --> 00:13:31,196
different from zero.
So the sample autocovariance is just at

178
00:13:31,196 --> 00:13:36,019
lag J, is just the sample covariance
between XT and XT minus J.

179
00:13:36,019 --> 00:13:41,762
And the sample auto correlation is just
this sample auto co-variance divided by

180
00:13:41,762 --> 00:13:47,331
the sample variance, and these are
measures of linear dependence between a

181
00:13:47,331 --> 00:13:53,507
variable and its lags, and then we can do
a graphical plot called the sample auto

182
00:13:53,507 --> 00:13:59,745
correlation function, where we just plot
this sample auto correlation against the

183
00:13:59,745 --> 00:14:02,783
lag.
So, notice that when you calculate this

184
00:14:02,783 --> 00:14:06,465
sum.
We start the summation at J plus one,

185
00:14:06,465 --> 00:14:09,965
right, because we're looking at the
summation.

186
00:14:09,965 --> 00:14:15,550
So, the first index in this sum is J plus
one, so we cut, we look at here.

187
00:14:15,550 --> 00:14:19,341
That's going to be X, J plus one minus the
mean.

188
00:14:19,341 --> 00:14:22,429
And then we have XJ + one - J.
So that's X1.

189
00:14:22,429 --> 00:14:24,205
Right?
Minus the mean.

190
00:14:24,205 --> 00:14:29,586
So we're looking at the relationship
between XJ and J lags from XJ.

191
00:14:29,586 --> 00:14:33,945
So from the first observation to the Jth
observation.

192
00:14:33,945 --> 00:14:40,012
Then the next term in the sum is XJ + one.
And then this would be X2, and so on.

193
00:14:40,012 --> 00:14:45,780
So we're looking at the relationship
between, you know, X at time T and its

194
00:14:45,780 --> 00:14:51,215
lag.
So if I guess if you draw a picture.

195
00:14:51,215 --> 00:14:54,841
Right?
So we have you know, one, two, three,

196
00:14:54,841 --> 00:15:01,372
four, you know, up to J, and so this would
be XJ, and then we have X1, and then we

197
00:15:01,372 --> 00:15:06,404
have XJ plus one X2.
So when we're calculating this sum, so

198
00:15:06,404 --> 00:15:12,958
we're looking at the relationship between
this variable and this variable, this

199
00:15:12,958 --> 00:15:18,558
variable and that variable, and in terms
of the computation of this.

200
00:15:18,558 --> 00:15:22,856
If we started this at zero.
Then this would be XO.

201
00:15:22,856 --> 00:15:27,299
Or say, if we started this at one, then
this would be X1.

202
00:15:27,299 --> 00:15:31,776
But this would be X1-J.
And we don't have data, for that.

203
00:15:31,776 --> 00:15:38,277
So, this, this notation is to emphasize
that we have to do this computation based

204
00:15:38,277 --> 00:15:42,507
on the data sample that we actually
observed.

205
00:15:42,507 --> 00:15:50,115
Okay, so here's an example of sample auto
correlations.

206
00:15:50,115 --> 00:15:56,333
So this is, these are plots that are
measuring the estimated time dependence in

207
00:15:56,333 --> 00:16:00,052
the data.
And the top graph is for the Gaussian

208
00:16:00,052 --> 00:16:04,055
white noise.
And the middle graph is for Microsoft.

209
00:16:04,055 --> 00:16:10,535
The bottom graph is for the S and P 500.
And this vehicular is being plotted here.

210
00:16:10,535 --> 00:16:15,361
So, this is I guess, lag one, lag two, lag
three, lag four.

211
00:16:15,361 --> 00:16:22,485
So for Gaussian white noise, the estimated
correlation between xt and xt minus one

212
00:16:22,485 --> 00:16:27,222
is, is actually zero.
It seems, you're just, because you're not

213
00:16:27,222 --> 00:16:31,603
seeing here.
The estimated correlation between xt and

214
00:16:31,603 --> 00:16:37,866
xt minus two is a small negative number.
So the scale here this is minus 0.15.

215
00:16:37,866 --> 00:16:44,045
So this value is you know like minus 0.05.
So this is a very small number.

216
00:16:44,045 --> 00:16:48,486
Now on these graphs are blue dotted lines,
okay.

217
00:16:48,486 --> 00:16:54,782
These blue dotted lines are thresholds to
determine whether or not these values are

218
00:16:54,782 --> 00:16:57,643
statistically different from zero.
Alright.

219
00:16:57,643 --> 00:17:03,026
After the midterm I'll explain where these
blue dotted lines come from.

220
00:17:03,026 --> 00:17:08,986
They're essentially based upon 95 percent
confidence interval for these estimates.

221
00:17:08,986 --> 00:17:15,129
So when you look at this graph, how you're
supposed to read it, is if any of the

222
00:17:15,129 --> 00:17:20,742
estimated auto correlations extend beyond
the blue line, then those auto

223
00:17:20,742 --> 00:17:25,727
correlations are thought to be
statistically different from zero at the

224
00:17:25,727 --> 00:17:30,016
at with, with 95 percent confidence.
And so for the Gaussian white noise, we

225
00:17:30,016 --> 00:17:34,291
see none of these auto- correlations lie
outside of the blue dotted lines.

226
00:17:34,291 --> 00:17:37,557
So we see no evidence of time dependence
in the data.

227
00:17:37,557 --> 00:17:41,339
And again, that should be the case.
Because the Gaussian white noise is

228
00:17:41,339 --> 00:17:44,787
computer simulated with no time dependence
in the data.

229
00:17:44,787 --> 00:17:49,648
On the other hand, if we look at the
Microsoft returns, we see that the first

230
00:17:49,648 --> 00:17:54,483
return is negative, and it's about -.2.
So this is the correlation between the

231
00:17:54,483 --> 00:17:57,587
return in month T and the return in month
T -one.

232
00:17:57,587 --> 00:18:02,342
And notice that it's negative.
And it extends beyond, beyond the blue

233
00:18:02,342 --> 00:18:05,235
dotted line.
So we would view this as being

234
00:18:05,235 --> 00:18:09,290
statistically significant.
So there appears to be a negative

235
00:18:09,290 --> 00:18:13,740
correlation between the return this month
and the return last month.

236
00:18:13,740 --> 00:18:16,115
Okay?
Now, the other auto-correlations.

237
00:18:16,115 --> 00:18:20,595
So this is the, return, the
autocorrelation between the return at

238
00:18:20,595 --> 00:18:23,658
month T and the return at month T minus
two.

239
00:18:23,658 --> 00:18:29,202
So this is the lag two autocorrelation,
the lag three autocorrelation and so on.

240
00:18:29,202 --> 00:18:34,838
Here, notice that the lag two is not
beyond the blue dotted line but the lag

241
00:18:34,838 --> 00:18:39,943
three is slightly positive, 'kay.
So if we run through the Microsoft data

242
00:18:39,943 --> 00:18:45,158
and we view this graph we would say there
appears to be some evidence for time

243
00:18:45,158 --> 00:18:50,146
dependence in the data, you know.
But if we, then, if you look at the S and

244
00:18:50,146 --> 00:18:52,542
P 500 data.
We see none of the sample

245
00:18:52,542 --> 00:18:55,710
auto-correlations are outside of the blue
dotted line.

246
00:18:55,710 --> 00:19:01,885
So for the S and P 500 data, there appears
to be no evidence for time-dependence in,

247
00:19:01,885 --> 00:19:05,154
in the date.
In statistical analysis you should never

248
00:19:05,154 --> 00:19:09,890
say there is, right, because you're, in
statistics you're never certain, you're

249
00:19:09,890 --> 00:19:12,420
never 100 percent certain of anything.
Right?

250
00:19:12,420 --> 00:19:17,388
All you can say, and you want to think of
yourself as, like, sitting on the jury

251
00:19:17,388 --> 00:19:21,209
evaluating evidence, right?
Is the evidence in favor of, or is the

252
00:19:21,209 --> 00:19:25,675
evidence against, something, right?
So you would interpret this as, there's

253
00:19:25,675 --> 00:19:30,564
data evidence in favor of some time
dependence in the Microsoft returns.

254
00:19:30,564 --> 00:19:35,360
When we look at the Microsoft data, just
how do you interpret this negative

255
00:19:35,360 --> 00:19:38,316
correlation?
And it's literally saying that, if the

256
00:19:38,316 --> 00:19:43,063
return this month is positive.
Then, there's a tendency for the return

257
00:19:43,063 --> 00:19:46,981
next month to be negative.
Right, so there's kinda of a reversal

258
00:19:46,981 --> 00:19:49,007
that's going on here.
>> And vice versa.

259
00:19:49,007 --> 00:19:53,076
>> And vice versa, and if the return is
negative this month there is a tendency

260
00:19:53,076 --> 00:19:57,726
for the return to be positive next month.
Right, so, one could ask the question,

261
00:19:57,726 --> 00:20:03,075
what could be causing such a reversal?
Now, in finance there is some explanation

262
00:20:03,075 --> 00:20:09,032
for the reversal effect in asset returns
that's known as the bid-ass bounce.

263
00:20:09,054 --> 00:20:12,085
I'll come back to that a little bit later
on.

264
00:20:13,007 --> 00:20:18,087
But that's, that's one explanation.
It's, it's not a very good explanation for

265
00:20:18,087 --> 00:20:22,083
monthly data.
It's, it's a better explanation if, if you

266
00:20:22,083 --> 00:20:26,036
have say, asset returns computed every
minute.

267
00:20:26,058 --> 00:20:30,032
But for monthly data, it's a bit of a
mystery.

268
00:20:30,032 --> 00:20:34,000
You know, why is this negative correlation
here?

269
00:20:36,055 --> 00:20:43,063
Okay, so to summarize, I like to call
these stylized facts.

270
00:20:43,063 --> 00:20:47,010
So in, in this section we looked at three
assets.

271
00:20:47,010 --> 00:20:52,073
We looked at Microsoft, the S and P five,
actually we only looked at two assets.

272
00:20:52,073 --> 00:20:56,077
We looked at Microsoft and we looked at
the S and P 500.

273
00:20:56,077 --> 00:21:02,077
And of course, one is always tempted to
try to make sweeping generalizations based

274
00:21:02,077 --> 00:21:07,088
upon, or, maybe I should rephrase this.
You should be careful not to make sweeping

275
00:21:07,088 --> 00:21:11,019
generalizations based on the analysis of
two assets.

276
00:21:11,019 --> 00:21:16,396
But, you know, I've, I've analyzed a lot
more than two assets, and others have

277
00:21:16,396 --> 00:21:21,040
analyzed a lot more than two assets.
And, so when you look at monthly

278
00:21:21,040 --> 00:21:27,002
continuously compounded returns, these are
the results that you tend to find when you

279
00:21:27,002 --> 00:21:30,084
look at, you know, lots and lots and lots
of different assets.

280
00:21:30,084 --> 00:21:36,060
At the monthly basis, returns appear to be
approximately normally distributed, okay?

281
00:21:36,060 --> 00:21:42,307
They don't follow the normal distribution
exactly, But the normal distribution is

282
00:21:42,307 --> 00:21:44,723
not a horrible mistake.
All right?

283
00:21:44,723 --> 00:21:50,187
There's some noticeable negative skewness
and excess kurtosis in typical assets,

284
00:21:50,187 --> 00:21:53,079
right?
If you look at bivariate relationships,

285
00:21:53,079 --> 00:21:58,091
many assets are contemporaneously
correlated, that is, and they tend to be

286
00:21:58,091 --> 00:22:04,002
positively correlated.
And so, and that's one of the things as

287
00:22:04,002 --> 00:22:08,015
well, if you look at temporal dependence
in the data.

288
00:22:08,015 --> 00:22:14,027
At the monthly level there is not a whole
lot of evidence for strong temporal

289
00:22:14,027 --> 00:22:19,018
dependence, so assets are approximately
uncorrelated over time.

290
00:22:19,041 --> 00:22:25,053
Then again when you look at many assets,
the Microsoft data showed there was some

291
00:22:25,053 --> 00:22:31,057
negative dependence on one lag, but when
you looked at, you know, other lags there

292
00:22:31,057 --> 00:22:36,071
was not much that was there.
So these are broad, what I call stylized

293
00:22:36,071 --> 00:22:41,456
facts and you want to keep these sort of
in the back of your mind.

294
00:22:41,456 --> 00:22:45,402
Stylized facts are useful at the model
building phase.

295
00:22:45,402 --> 00:22:50,409
So if you want to build a probability
model for asset returns, we want our

296
00:22:50,409 --> 00:22:54,856
probability model to capture the basic
stylized facts of the data, okay.

297
00:22:54,856 --> 00:23:00,378
If our model doesn't capture the basic
stylized facts, it's not a good model.

298
00:23:00,378 --> 00:23:07,518
And we're always gonna keep that in mind
as we look at particular models in this

299
00:23:07,518 --> 00:23:11,108
class.
We always wanna ask ourselves, you know,

300
00:23:11,108 --> 00:23:16,510
if we you know, simulate data from our
model, does it look like real data?
