1
00:00:00,000 --> 00:00:05,472
So, we'll look at data now finally.
And before we actually start looking at

2
00:00:05,472 --> 00:00:10,011
data, remind yourself a covariance
stationary.

3
00:00:10,011 --> 00:00:14,111
So, we're going to look at data.
And, when we start looking at asset return

4
00:00:14,111 --> 00:00:19,656
data, you want to think yourself that the
asset returns that we observe are

5
00:00:19,656 --> 00:00:23,087
realization of a covariance stationary
time series, okay?

6
00:00:23,087 --> 00:00:29,483
And so, you think of Gaussian White Noise
as the benchmark and, and then you want to

7
00:00:29,483 --> 00:00:34,082
think with a critical eye of may, of, what
a moving average process looks like, so

8
00:00:34,082 --> 00:00:39,154
time dependence of, you know, one period,
or an autoregressive process where you

9
00:00:39,154 --> 00:00:43,470
have decaying time dependence.
And you want to ask yourself, you know,

10
00:00:43,470 --> 00:00:47,192
you know, what properties do you actually
see in the data?

11
00:00:47,401 --> 00:00:52,928
And if the process were covariance
stationary process, what kind of process

12
00:00:52,928 --> 00:00:55,014
would it most likely be?
Okay?

13
00:00:55,014 --> 00:00:58,061
So, that's the kind of game that we want
to play.

14
00:00:58,061 --> 00:01:02,088
So, all of the returns have the same mean,
the same variance.

15
00:01:02,088 --> 00:01:08,081
They have a correlation that depends upon
how far apart they are, but on time, okay?

16
00:01:08,081 --> 00:01:12,050
There should be no trends in the data,
and, and so on.

17
00:01:12,050 --> 00:01:16,045
So, that's the process.
So, now let's look at some data.

18
00:01:20,064 --> 00:01:27,619
So, for this next set of lectures, I'm
going to be looking well, two time series.

19
00:01:27,863 --> 00:01:34,055
I have monthly continuously compounded
returns that I downloaded from Yahoo and

20
00:01:34,275 --> 00:01:40,482
one series is going to be the monthly
returns on Microsoft, which is just an

21
00:01:40,482 --> 00:01:44,849
individual asset.
And the other series we're going to look

22
00:01:44,849 --> 00:01:49,584
at is the, the S and P 500 index, the
Standard and Poor's 500.

23
00:01:49,584 --> 00:01:55,131
The S and P 500 think of this as a
portfolio of 500 assets, okay?

24
00:01:55,131 --> 00:02:01,025
So, it's a, it has, it's a portfolio of,
of 500 stocks, and the weight of each

25
00:02:01,025 --> 00:02:06,071
asset in the portfolio is based on its
market capitalization.

26
00:02:06,071 --> 00:02:10,880
So, for example, if Microsoft were in the
S and P 500, the weight of Microsoft is

27
00:02:10,880 --> 00:02:15,011
equal to the, the market value of
Microsoft which is the price of Microsoft

28
00:02:15,011 --> 00:02:19,448
times the number of shares outstanding,
that's its market capitalization.

29
00:02:19,448 --> 00:02:24,241
So, if you had to go and, if you wanted to
buy Microsoft as a company, what price

30
00:02:24,241 --> 00:02:27,551
would you pay?
It would be price times number of shares

31
00:02:27,551 --> 00:02:31,076
outstanding, which for Microsoft is about
$35 billion.

32
00:02:31,356 --> 00:02:37,463
And then, so, you take the market
capitalization of Microsoft and then you

33
00:02:37,463 --> 00:02:42,220
divide it by the sum of the market
capitalizations of all of the other 500

34
00:02:42,220 --> 00:02:46,093
stocks, right?
So, the weights of, of all the assets in

35
00:02:46,093 --> 00:02:51,860
the S and P 500 add to one, and assets
that have high market capitalizations have

36
00:02:51,860 --> 00:02:57,426
higher weight in the S and P 500.
It's most of the, the 500 stocks that are

37
00:02:57,426 --> 00:03:04,494
in the S and P 500 tend to be very large
capitalization stocks, so they're big

38
00:03:04,494 --> 00:03:09,430
well-established companies.
And, and, and the weights, you know, of

39
00:03:09,430 --> 00:03:14,018
these assets, it's, I don't want to say
it's equally weighted but it's fairly

40
00:03:14,018 --> 00:03:17,205
equally weighted.
So, there isn't a, a huge concentration on

41
00:03:17,205 --> 00:03:20,513
one or two stocks.
It's, it's a fairly diversified spread-out

42
00:03:20,513 --> 00:03:24,980
portfolio.
So, the key thing to keep in mind is, this

43
00:03:24,980 --> 00:03:30,194
is an individual stock.
This is a highly diversified portfolio of

44
00:03:30,194 --> 00:03:33,271
500 stocks.
And what we want to do is we wanna say,

45
00:03:33,271 --> 00:03:37,709
what are the, the stylized facts about the
returns on an individual security and a

46
00:03:37,709 --> 00:03:42,070
highly diversified portfolio.
And we're going to see that they're very

47
00:03:42,070 --> 00:03:47,023
distinct differences in the properties of
the returns of these two series, okay?

48
00:03:47,230 --> 00:03:52,000
Now, here are the plots of the
continuously compounded return, and the

49
00:03:52,000 --> 00:03:56,040
data goes from 1998 to 2008.
So, we don't have the financial crisis in

50
00:03:56,040 --> 00:04:01,052
this period and so we'll look, we'll do a
homework assignment you know, you'll be

51
00:04:01,052 --> 00:04:04,767
looking at more updated data.
And so, when you put the financial crisis

52
00:04:04,767 --> 00:04:07,066
in here, you'll see a big negative
returns, okay?

53
00:04:07,066 --> 00:04:10,066
So, that's you know, that's something to
be considered.

54
00:04:10,066 --> 00:04:15,746
Now, when we look at these two graphs, you
want to ask yourself, what properties do

55
00:04:15,746 --> 00:04:20,000
you see, and what similarities do you see
between the two series?

56
00:04:20,000 --> 00:04:22,487
Okay?
So, first of all, do they look like

57
00:04:22,487 --> 00:04:27,053
Gaussian White Noise?
Notice in, in the, from 1998 to 2002, the

58
00:04:27,053 --> 00:04:33,013
fluctuations of Microsoft were pretty big,
so we call that the volatility.

59
00:04:33,013 --> 00:04:38,058
But, after about 2002, notice the
magnitude of the fluctuations decrease,

60
00:04:38,058 --> 00:04:41,088
right?
So, is that a property of a covariance

61
00:04:41,088 --> 00:04:44,042
stationary time series?
Yes or no?

62
00:04:44,042 --> 00:04:47,015
No.
It's because the variance appears to

63
00:04:47,015 --> 00:04:51,006
depend upon time.
During this period the variance was big.

64
00:04:51,006 --> 00:04:54,055
During the, this, this period the variance
was small.

65
00:04:54,055 --> 00:04:59,083
So, that's an indication, perhaps, that
the volatility is, is changing over time,

66
00:04:59,083 --> 00:05:02,084
okay?
So, that's one, one future to see.

67
00:05:02,084 --> 00:05:08,032
What about the mean in the series?
There doesn't appear to be any obvious

68
00:05:08,032 --> 00:05:11,041
trend, right?
So, the, the data appears to fluctuate

69
00:05:11,041 --> 00:05:16,039
pretty close to zero in, in this plot.
And so, so if we think about covariance

70
00:05:16,039 --> 00:05:21,158
stationary in terms of the mean being
constant, we don't see like the mean is up

71
00:05:21,158 --> 00:05:26,237
here, and then the mean is down below, and
then the mean is up here, and, and things

72
00:05:26,237 --> 00:05:30,068
like that.
So, now one of the things about, and, and

73
00:05:30,068 --> 00:05:36,052
so we, I, I want to follow up on this,
this volatility comment that was made

74
00:05:36,052 --> 00:05:40,020
before.
Microsoft has high volatility here than

75
00:05:40,020 --> 00:05:44,017
low volatility, we also that in the S and
P 500 index.

76
00:05:44,017 --> 00:05:48,695
So, very often, the volatility of the S
and P 500 index is called market

77
00:05:48,695 --> 00:05:53,327
volatility because the S and P 500 is
thought to be a proxy for the overall

78
00:05:53,327 --> 00:05:57,522
movements in the stock market.
And so here, market volatility was high,

79
00:05:57,522 --> 00:06:00,357
and then down here market volatility was
low.

80
00:06:00,357 --> 00:06:05,810
So, notice that Microsoft's volatility is
tracking the market volatility as well.

81
00:06:05,810 --> 00:06:10,658
So, we see some common behavior between
the series and in this regard.

82
00:06:10,658 --> 00:06:15,347
Now, these two graphs are a little bit
misleading because the scales are

83
00:06:15,347 --> 00:06:18,651
different.
Notice that here, this, this is -40%, this

84
00:06:18,651 --> 00:06:23,262
is twenty percent for Microsoft.
For the S and P 500, this is -fifteen%,

85
00:06:23,262 --> 00:06:26,728
this is six%.
So, the volatility of the market is much,

86
00:06:26,728 --> 00:06:31,439
much lower than the volatility of
Microsoft, but that's not apparent because

87
00:06:31,439 --> 00:06:37,097
of the way that the graphs are drawn.
If we put these two asset, assets on the

88
00:06:37,097 --> 00:06:41,862
same scale.
So, here, I have Microsoft in red, and the

89
00:06:41,862 --> 00:06:47,032
S and P 500 in blue.
Then, we can see I think, more clearly,

90
00:06:47,032 --> 00:06:51,040
some similarities and differences between
the series, okay?

91
00:06:51,040 --> 00:06:56,027
So, very clearly, the volatility of
Microsoft is much greater than the

92
00:06:56,027 --> 00:07:00,085
volatility in the S and P 500, okay?
So, this one fact about assets.

93
00:07:00,085 --> 00:07:06,006
When you have a highly diversified
portfolio like the S and P 500, you, you

94
00:07:06,006 --> 00:07:10,349
do, you have a smaller volatility than an
individual security, okay?

95
00:07:10,349 --> 00:07:14,406
And remember, volatility is very often
used as a proxy for risk.

96
00:07:14,406 --> 00:07:17,810
Volatility is telling you uncertainty
about the mean.

97
00:07:17,810 --> 00:07:22,713
So, the uncertainty about the mean for the
S and P 500 is a lot smaller than the

98
00:07:22,713 --> 00:07:25,585
uncertainty associated with, with
Microsoft.

99
00:07:25,585 --> 00:07:30,551
So, we might think there's less risk in
the market than there is in Microsoft by

100
00:07:30,551 --> 00:07:35,054
itself, okay?
Now, what's another feature that we can

101
00:07:35,054 --> 00:07:42,041
see, or, or another commonality between
Microsoft and the S and P 500, that we can

102
00:07:42,041 --> 00:07:48,050
see in this graph, that we couldn't see
very clearly in the other graph?

103
00:07:48,050 --> 00:07:52,044
Are Microsoft and the S and P 500
correlated with each other?

104
00:07:52,044 --> 00:07:57,092
So, Microsoft and the S and P 500 are kind
of moving up and down together at the same

105
00:07:57,092 --> 00:08:00,050
time.
Well, that means they're positively

106
00:08:00,050 --> 00:08:03,019
correlated, right?
When one, when the market goes up,

107
00:08:03,019 --> 00:08:07,226
Microsoft goes up.
When the market goes down, the Microsoft

108
00:08:07,226 --> 00:08:10,411
goes down.
Now, I'm not saying that there is a causal

109
00:08:10,411 --> 00:08:14,461
relationship between the two.
They're, they're just moving at the same

110
00:08:14,461 --> 00:08:15,612
time, okay?
Right.

111
00:08:15,612 --> 00:08:21,306
So, those are some important properties
and these are the kinds of things that we

112
00:08:21,306 --> 00:08:24,515
want to do when we look at data.
We want to see, what do we see?

113
00:08:24,515 --> 00:08:26,746
What are commonalities?
What are features?

114
00:08:26,965 --> 00:08:29,514
Is it covariance stationary?
If so, why?

115
00:08:29,514 --> 00:08:31,701
If not, why not?
Those kinds of things.

116
00:08:31,701 --> 00:08:36,576
Is there a relationship between the
volatility of the series and the expected

117
00:08:36,576 --> 00:08:38,818
return?
So, if you invest, if you take a stock

118
00:08:38,818 --> 00:08:43,220
that has a high volatility, do you tend to
get a high average return as well?

119
00:08:43,220 --> 00:08:47,141
And the answer is, is the stylized fact
is, is generally yes.

120
00:08:47,141 --> 00:08:51,345
When we look at stocks with high
volatility, they tend to have higher

121
00:08:51,345 --> 00:08:54,438
expected returns than lower volatility
stocks.

122
00:08:54,438 --> 00:08:58,340
Not always.
This is the idea that risk and return go

123
00:08:58,340 --> 00:09:01,173
together.
So, and we'll be, we'll be doing some

124
00:09:01,173 --> 00:09:05,668
plots later on while we're say, put
expected average return or expected return

125
00:09:05,668 --> 00:09:11,373
on this axis and volatility on this axis,
and we'll see that you know, things tend

126
00:09:11,373 --> 00:09:16,618
to move up in a direction like this and
it's, it's tends to be a somewhat stylized

127
00:09:16,618 --> 00:09:22,580
fact but it's not always true.
Okay.

128
00:09:22,580 --> 00:09:29,035
So, our observed sample is a realization
of this covariance stationary time series.

129
00:09:29,035 --> 00:09:34,078
And, now, what we want to do is we want to
compute some of descriptive statistics,

130
00:09:34,078 --> 00:09:36,950
okay?
Now, in, in a very simple term, you know,

131
00:09:36,950 --> 00:09:42,092
a statistic is just a data summary.
So, we compute descriptive statistics,

132
00:09:42,092 --> 00:09:48,091
that means we're creating data summaries
and we're trying to capture particular

133
00:09:48,091 --> 00:09:52,001
features.
Now, when we studied random variables and

134
00:09:52,001 --> 00:09:56,010
looked at probability distributions, I
said there were certain shape

135
00:09:56,010 --> 00:09:58,088
characteristics of a probability
distribution.

136
00:09:58,088 --> 00:10:03,034
There was the center of the distribution,
which was the expected value, the

137
00:10:03,034 --> 00:10:07,008
variance, the spread, the skewness,
asymmetry, kurtosis, fat tails.

138
00:10:07,008 --> 00:10:11,054
Well, all of these properties that we've
talked about for random variables, we can

139
00:10:11,054 --> 00:10:14,025
create a descriptive statistics from the
sample.

140
00:10:14,025 --> 00:10:18,055
So, if we want to know what is the
probability distribution look like?

141
00:10:18,055 --> 00:10:22,878
We can compute a histogram which is a
graphical descript, description of the of

142
00:10:22,878 --> 00:10:27,094
the distribution of the data.
And then, of that histogram, we can think

143
00:10:27,094 --> 00:10:31,098
where's the histogram centered?
You know, what's the spread in the

144
00:10:31,098 --> 00:10:34,065
histogram?
What's it's skewness, or kurtosis?

145
00:10:34,065 --> 00:10:39,007
So, for all of the probability model
concepts, there's going to be a sample

146
00:10:39,007 --> 00:10:43,250
analogue, a descriptive statistics that
gives us, you know, the same kind of

147
00:10:43,250 --> 00:10:46,362
information, okay?
So, we're going to create data summaries

148
00:10:46,362 --> 00:10:51,433
to describe certain features of the data,
to learn about the unknown probability

149
00:10:51,433 --> 00:10:55,773
distribution that we think is actually
generating the data.

150
00:10:55,773 --> 00:11:01,535
So, the, the thing about, you know,
probability modeling and statistics in the

151
00:11:01,535 --> 00:11:07,005
real world is, we don't know what the true
model is that generated the data.

152
00:11:07,005 --> 00:11:09,073
I mean, we'll never know what it is,
right?

153
00:11:09,073 --> 00:11:14,306
And so, what we're trying to do is we're
trying to kinda guess at what we think is

154
00:11:14,306 --> 00:11:19,013
the underlying model.
And, we're going to use these descriptive

155
00:11:19,013 --> 00:11:24,611
statistics to try to help us pick good
models for describing the data versus you

156
00:11:24,611 --> 00:11:30,084
know, bad models, and, and so on.
Alright.
