1
00:00:01,017 --> 00:00:05,537
Alright. So, we're going to talk about
random variables and probability review.

2
00:00:05,537 --> 00:00:10,618
And so, again, this is a course about
modeling financial data. And the primary

3
00:00:10,618 --> 00:00:15,023
thing that we're going to be looking at
are asset returns. And, and when you think

4
00:00:15,023 --> 00:00:19,837
about an asset return, You know, we went
over return calculations before, you know?

5
00:00:19,837 --> 00:00:23,788
Say, we're investing in Microsoft stock,
we buy it today, we hold it for a month,

6
00:00:23,788 --> 00:00:28,316
one month from now, we sell it, we can
calculate the percentage change in price

7
00:00:28,316 --> 00:00:32,374
that's the rate of return. But from the
point of view of today, the rate of return

8
00:00:32,374 --> 00:00:36,406
on this one-month investment is not known,
because we need to know what the price

9
00:00:36,406 --> 00:00:41,016
next month is in order to be able to
calculate that return. So, we can think of

10
00:00:41,016 --> 00:00:46,372
the rate of a return as a random variable
because the future is not known and the

11
00:00:46,372 --> 00:00:52,282
return depends upon the future price,
which is not known. And so you know, the

12
00:00:52,282 --> 00:00:57,598
outcome is uncertain. And so we can
characterize returns as a random variable,

13
00:00:57,598 --> 00:01:02,832
which is a variable that can take on a, a
possible set of values called the sample

14
00:01:02,832 --> 00:01:07,797
space. So, the future price could go up,
it could go down. And so, there is a

15
00:01:07,797 --> 00:01:13,079
potential list of values of the future
price and then, we attach a probability to

16
00:01:13,275 --> 00:01:18,563
each of those prices and that gives us the
probability distribution over those

17
00:01:18,563 --> 00:01:23,905
potential values. So, what we are going to
review today is various mathematical ways

18
00:01:23,905 --> 00:01:29,136
of describing random variables and
distributions and with several examples

19
00:01:29,136 --> 00:01:33,685
towards thinking about asset returns and
the properties of probability

20
00:01:33,685 --> 00:01:40,255
distributions for asset returns that look
reasonable. So some examples typically, an

21
00:01:40,255 --> 00:01:44,852
uppercase letter denotes a random
variable. So, uppercase X is, in this

22
00:01:44,852 --> 00:01:50,096
case, the price of Microsoft stock next
month. It's a random variable cuz we don't

23
00:01:50,096 --> 00:01:54,436
know what the future price is going to be.
What's the sample space of the future

24
00:01:54,436 --> 00:01:58,362
prices? Well, the sample space, this is
the possible values that we can think,

25
00:01:58,362 --> 00:02:03,441
that we think the prices can take on. Now,
prices can't become negative so we know

26
00:02:03,441 --> 00:02:07,858
that the sample space is going to be t
hose say positive values. And you know, if

27
00:02:07,858 --> 00:02:12,841
the stock is trading, if the stock doesn't
go out of business, then the price should

28
00:02:12,841 --> 00:02:18,053
be positive. Now, prices can't go up to
infinity, so there's some realistic upper

29
00:02:18,053 --> 00:02:23,779
bound on prices and that upper bound might
be a $1,000 or something like that. So, we

30
00:02:23,779 --> 00:02:29,399
can think future prices can lie anywhere,
any real number, say, between zero and

31
00:02:29,399 --> 00:02:35,351
some big number M like a 1,000. So, that
would be a, a characterization of a random

32
00:02:35,351 --> 00:02:40,545
variable, the future price and its sample
space. Another random variable could be

33
00:02:40,545 --> 00:02:45,369
the rate of return on the investment and
the rate of return, which is the

34
00:02:45,369 --> 00:02:50,866
percentage change in price. Now what is
this appropriate example space for a rate

35
00:02:50,866 --> 00:02:56,077
of return. Now, what's the smallest value
that a rate of return can take on? The

36
00:02:56,077 --> 00:03:02,414
worst you can do is lose all your money,
right? So, if you, you buy Microsoft today

37
00:03:02,414 --> 00:03:07,644
for 30, and if Microsoft goes bankrupt,
and its price goes to zero in the future,

38
00:03:07,644 --> 00:03:12,078
then the percentage change in price is
minus a 100%. So, the returns are bounded

39
00:03:12,078 --> 00:03:17,088
from below by -one, so it can't be any
more than that. And the upper bound,

40
00:03:17,088 --> 00:03:22,454
again, is some big positive number. You
know, say, Microsoft, you know, I don't

41
00:03:22,454 --> 00:03:27,017
know, in, invent some product that
revolutionizes the world and its price is

42
00:03:27,017 --> 00:03:30,512
going to shoot up. And but, you know, it's
not going to go off to infinity. So,

43
00:03:30,512 --> 00:03:35,834
there's some reasonable upper bound
associated with that. Another random

44
00:03:35,834 --> 00:03:41,740
variable we can think of, and that's used
a lot in probability modelling in finance,

45
00:03:41,740 --> 00:03:46,559
is a just a, a discrete random variable
that just takes on two values. So, we're

46
00:03:46,559 --> 00:03:51,465
going to set x to be equal to one if the
stock price goes up and we'll say, x is

47
00:03:51,465 --> 00:03:55,479
equal to zero if the stock price goes
down. This is sort of like a coin flipping

48
00:03:55,479 --> 00:04:00,261
example. You flip the coin and it lands
heads, that's like the stock prices going

49
00:04:00,261 --> 00:04:04,445
up. You flip the coin, it lands tails,
it's like the stock prices going down, and

50
00:04:04,445 --> 00:04:09,031
then we're just coding the random variable
to be 1,0 based upon those events. And

51
00:04:09,031 --> 00:04:12,620
here, the sample space is very simple with
just, just two va lues zero and one. So,

52
00:04:12,620 --> 00:04:17,081
those are examples of, of random variables
that we'll be looking at. So, once we have

53
00:04:17,081 --> 00:04:21,085
these random variables, we need to
characterize their probability

54
00:04:21,085 --> 00:04:26,488
distribution. Now we in, in probability
models, we usually distinguish between

55
00:04:26,488 --> 00:04:31,014
what are call discrete random variables.
And the discrete random variables is a

56
00:04:31,014 --> 00:04:35,700
random variables that can only take on
finite set of values. So, in the last

57
00:04:35,700 --> 00:04:40,737
example, the up-down indicator, we took on
two values, zero and one, that's the

58
00:04:40,737 --> 00:04:45,954
discrete random variables and its sample
space is two discrete points, okay? The

59
00:04:45,954 --> 00:04:51,609
probability distribution of a discrete
random variable is a function, say, P(x).

60
00:04:51,609 --> 00:04:58,481
Such that P(x) is the probability that the
random variable is equal to little x. So

61
00:04:58,481 --> 00:05:04,525
typically, in notation capital letter
denotes the random variable and a lower

62
00:05:04,525 --> 00:05:10,668
case letter denotes a value in the sample
space that the random variable can take

63
00:05:10,668 --> 00:05:16,819
on. Now, this probability function must
satisfy certain conditions in order to be

64
00:05:16,819 --> 00:05:21,685
a valid probability function. So,
probabilities are greater and equal to

65
00:05:21,685 --> 00:05:27,795
zero for all values in the sample space.
Probabilities are equal to zero for values

66
00:05:27,795 --> 00:05:33,961
outside of the samples space. You know, we
assume that all the values in the sample

67
00:05:33,961 --> 00:05:38,306
space, you know, again are, are discrete,
distinct points. The sum of the

68
00:05:38,306 --> 00:05:42,635
probabilities of the values in the sample
space is equal to a 100%. And

69
00:05:42,822 --> 00:05:48,730
probabilities have to be less than one, as
well. So, as an example of a simple

70
00:05:48,730 --> 00:05:56,804
discrete random variable and a probability
distribution, here's a case where we can

71
00:05:56,804 --> 00:06:09,692
have a random variable x here and X is
going to denote the annual rate of return

72
00:06:09,692 --> 00:06:19,340
on Microsoft stock. Alright. So, we have a
random variable X that's going to

73
00:06:19,340 --> 00:06:24,310
represent the annual rate of return on
Microsoft stock. And this is an example

74
00:06:24,310 --> 00:06:28,957
that's kind of like the case where you
might have a stock analyst who's working

75
00:06:28,957 --> 00:06:33,814
for an investment bank, and they're doing
some fundamental an analysis and they need

76
00:06:33,978 --> 00:06:38,255
you know, make some forecast of what the
rate of the return might be over the next

77
00:06:38,255 --> 00:06:42,837
year. And the analysis might have a very
simplified way of viewing the world that

78
00:06:43,004 --> 00:06:47,797
you know, over the next year the price of
Microsoft stock is primarily contingent

79
00:06:47,797 --> 00:06:51,927
upon what's going to happen to the state
of the economy. And so, the analyst says,

80
00:06:51,927 --> 00:06:56,859
well, I think there are really one, two,
three, four, five potential states of the

81
00:06:56,859 --> 00:07:02,431
economy, you know, depression, you know,
it's a very bad state. And if the economy

82
00:07:02,431 --> 00:07:07,163
is in a depression, then Microsoft is
going to lose 30%. And the analyst puts a

83
00:07:07,163 --> 00:07:11,966
probability of five% on that event
occurring. And then, the other state of

84
00:07:11,966 --> 00:07:17,399
the world could be, say, a recession. And
if a recession happens, then Microsoft has

85
00:07:17,399 --> 00:07:22,682
an annual return of zero%, and the analyst
puts a probability of twenty% on that. And

86
00:07:22,682 --> 00:07:27,068
then similarly, normal mild boom, major
booms, these are other states of the

87
00:07:27,068 --> 00:07:32,030
world, and then we see as the economy gets
better, the rate of return goes up, and,

88
00:07:32,190 --> 00:07:36,881
and then we have these probabilities. So
notice that the, the normal state of the

89
00:07:36,881 --> 00:07:42,049
world is the state that gets the highest
probability associated with it. Now this

90
00:07:42,049 --> 00:07:47,791
is a valid probability distribution, all
the probabilities are between zero and one

91
00:07:47,791 --> 00:07:53,897
and the sum of all the probabilities add
up to one. Now, in this case, you know,

92
00:07:53,897 --> 00:07:58,059
where do these probabilities come from?
You now, in, in this case, the

93
00:07:58,059 --> 00:08:04,018
probabilities are the subjective beliefs
of the analyst. They may have nothing to

94
00:08:04,018 --> 00:08:08,090
do with the real world, you know, in
quotations. But there are just, you know,

95
00:08:08,090 --> 00:08:12,610
views or opinions associated with, with
the analyst. In, in probability theory,

96
00:08:12,610 --> 00:08:17,523
there are generally two types of ways that
we view probability. One way, the

97
00:08:17,523 --> 00:08:22,788
subjective approach is probabilities are
opinions or degrees of belief. And this is

98
00:08:22,788 --> 00:08:28,439
associated with what is often referred to
as Basian statistics. The other approach

99
00:08:28,439 --> 00:08:33,876
views probabilities of, actually,
actually, real physical objective things .

100
00:08:33,876 --> 00:08:39,706
So, the, the, the objective view of
probability. Think of the coin flipping

101
00:08:39,706 --> 00:08:45,060
example. So, if you have a fair coin that
is, it's weighted such that the

102
00:08:45,060 --> 00:08:50,701
probability that the coin lands heads or
tails is 50%. So, where doe s that

103
00:08:50,701 --> 00:08:56,354
probability of 50% come from? Well, it's
an objective feature of the coin, and the

104
00:08:56,354 --> 00:09:01,459
experiment of flipping the coin and the
idea is that, you know, when you flip a

105
00:09:01,459 --> 00:09:06,530
fair coin, you know, and you, you do this
in an experimental setting, say, a million

106
00:09:06,530 --> 00:09:11,502
times, the probability that the coin lands
heads is a fraction of the times in your

107
00:09:11,502 --> 00:09:16,237
experiment that you actually observe the
coin landing heads. And the probability

108
00:09:16,237 --> 00:09:21,033
that lands tails is the fraction of times
you actually observe the coin landing

109
00:09:21,033 --> 00:09:25,480
tails. So there, you think of probability
as being a property of the coin, the

110
00:09:25,480 --> 00:09:29,641
experimental setting, and something that
you can repeat over and over again and,

111
00:09:29,641 --> 00:09:34,054
and reproduce. Whereas, the opinion
approach is, there's no experiment going

112
00:09:34,054 --> 00:09:39,049
on here. You can't repeat, you know, the
idea of the views going over and over

113
00:09:39,049 --> 00:09:44,283
again. It's just a degree of belief. Both
views of probability theory are, are, are

114
00:09:44,283 --> 00:09:48,213
equally valid, and but, you know, within
the statistics literature, you know,

115
00:09:48,213 --> 00:09:53,575
there's often very strong opinions about
what is the right way or the superior way

116
00:09:53,575 --> 00:09:58,998
to view probability. In this class, we
won't take a, make a judgement on that,

117
00:09:58,998 --> 00:10:06,733
but we'll, you know, essentially just take
probability as given and do computations.

118
00:10:06,733 --> 00:10:15,663
So one of the things I want to, to show
you is an Excel spreadsheet that I have

119
00:10:15,663 --> 00:10:26,063
and now, Excel is not necessarily a good
tool for doing probability calculations.

120
00:10:26,063 --> 00:10:32,065
And so, the point here is just to
illustrate how to do certain computations

121
00:10:32,065 --> 00:10:37,833
in Excel and, and, and show some, and show
graphics and, and, and things like that. R

122
00:10:37,833 --> 00:10:43,153
is a much better environment for doing a
probability calculations and, and so on.

123
00:10:43,153 --> 00:10:49,037
But in, for my case, you know, I have my
example I just put out an Excel

124
00:10:49,037 --> 00:10:54,423
spreadsheet, I have my returns, I have my
probabilities, and then I can do a nice

125
00:10:54,423 --> 00:10:59,059
simple bar chart to represent the
probability distribution. So, in the

126
00:10:59,059 --> 00:11:03,076
graphical representation of the
distribution, I put the values in the

127
00:11:03,076 --> 00:11:08,046
sample space on this axis. And then, the
heights of the bars just represent the

128
00:11:08,046 --> 00:11:13,023
probabilities. And so, from this graphic
al representation, we can clearly see the

129
00:11:13,023 --> 00:11:17,893
most likely value is around one%. The
shape of the distribution is symmetric and

130
00:11:17,893 --> 00:11:22,294
that this is the middle of the
distribution. The shape to the left of the

131
00:11:22,294 --> 00:11:27,712
middle, and the shape to the right of the
middle is the same and, and so on. We can

132
00:11:27,712 --> 00:11:34,814
do the same computation in, in R. So, here
is my this is new, this is again, just

133
00:11:34,814 --> 00:11:41,405
based on the examples in my lecture notes.
So, if I wanted to plot the probability

134
00:11:41,405 --> 00:11:49,229
distribution in R I would just create a
vector of values that represent the values

135
00:11:49,229 --> 00:11:58,229
in the sample space, create a vector of
probabilities, and then do a bar plot. And

136
00:11:58,698 --> 00:12:09,074
that gives me the, the values here. Okay.
A more mathematical model for a discreet

137
00:12:09,074 --> 00:12:15,016
random variable is based on what's called
the Bernoulli distribution And the

138
00:12:15,016 --> 00:12:20,142
Bernoulli distribution is a probability
model that essentially describes the

139
00:12:20,142 --> 00:12:26,006
coin-flipping experiment. So, we have two
mutually exclusive events generically

140
00:12:26,006 --> 00:12:31,522
called success and failure, right? So, in
modeling stock prices, a success event is

141
00:12:31,522 --> 00:12:36,554
the stock price goes up, a fail event is
the stock price goes down. That's assuming

142
00:12:36,554 --> 00:12:41,464
you have a long position in, in, in, in
the asset. If you have a short position,

143
00:12:41,464 --> 00:12:46,754
then stock price going up is the failure,
and the stock price going down is the

144
00:12:46,754 --> 00:12:51,615
success. What, so, in investments, a long
position means you buy something today.

145
00:12:51,615 --> 00:12:56,564
You hold it, you sell it in the future. A
short position means you sell it today,

146
00:12:56,564 --> 00:13:01,153
you, and then you buy it back in the
future. Okay, and so, when your long

147
00:13:01,153 --> 00:13:05,871
something, you're hoping the price will go
up. When you're short something, you hope

148
00:13:05,871 --> 00:13:10,569
the price is going to go down. Typically
shorting works, I just, just as an aside,

149
00:13:10,569 --> 00:13:15,891
because this is going to come up lat er
on. When you short a stock, typically, you

150
00:13:15,891 --> 00:13:20,921
open a brokerage account in, you know,
like at Fidelity or E-trade or something

151
00:13:20,921 --> 00:13:25,156
like that. Say, you want to short
Microsoft. Well, you are going to sell

152
00:13:25,156 --> 00:13:29,330
something you don't own. So, how do you do
that? Well, if you have a brokerage

153
00:13:29,330 --> 00:13:33,498
account, you can borrow the stock from
somebody who owns it and you borrow it,

154
00:13:33,498 --> 00:13:37,564
and you sell it, you get the proceeds that
you hold on to it. But because you

155
00:13:37,564 --> 00:13:42,127
borrowed it, you have to give it back at
some point. So, when you close out the

156
00:13:42,127 --> 00:13:46,608
short position, you go back in the market,
you buy it back, and then you return the

157
00:13:46,608 --> 00:13:51,293
stock to who you borrowed it from. So, you
want to think of that transaction is

158
00:13:51,293 --> 00:13:57,670
taking place when you do a short sale.
Alright. So we're going to calling this in

159
00:13:57,670 --> 00:14:03,379
this, go back to the Bernoulli example.
Let's say X = one, if a success occurs and

160
00:14:03,379 --> 00:14:08,075
X = zero, if a failure occurs, okay? So,
that's, that's coin lands head, success,

161
00:14:08,075 --> 00:14:13,809
coin lands tail failure. Now, the
probability that we have a success, that X

162
00:14:13,809 --> 00:14:18,537
= one is were, is been equal to pi and pi
is some number between zero and one. And

163
00:14:18,537 --> 00:14:23,963
then the probability of a failure that X =
zero is then one - pi. Alright, cuz we

164
00:14:23,963 --> 00:14:28,609
have two events. The sum of the pi will
always have to add to one, so pi + one -

165
00:14:28,609 --> 00:14:36,994
pi = one. Now, a simple mathematical model
for this probability distribution are

166
00:14:36,994 --> 00:14:44,008
P(x), we can write as pi^x one - pi^1 - x
where x only takes two values, zero and

167
00:14:44,008 --> 00:14:49,185
one. So, this P(x) function gives us our
probabilities. Notice that when X is zero,

168
00:14:49,185 --> 00:14:57,298
P of zero is pi^0 one - pi^1 - zero. So,
that's one - pi. So then, we have pi^0 is

169
00:14:57,298 --> 00:15:05,918
one. One - pi^1 is one - pi. And then,
when X = one, my P of one is pi^1 one -

170
00:15:05,918 --> 00:15:10,005
pi^1 - one so that's equal to -pi. So,
this very simple mathematical

171
00:15:10,005 --> 00:15:15,664
representation gives us our probability
function for the Bernoulli distribution.

172
00:15:15,664 --> 00:15:24,512
Now, the other type of random variables
that we work at, look at, are called

173
00:15:24,512 --> 00:15:29,489
continuous random variables. A continuous
random variable is one, is a random

174
00:15:29,489 --> 00:15:34,725
variable that can take on any real value,
alright? And so now, we talk about the

175
00:15:34,725 --> 00:15:39,956
probab ility density function of a
continuous random variable. We're going to

176
00:15:39,956 --> 00:15:44,924
have a probability function that we're
going to denote as f(x) to distinguish it

177
00:15:44,924 --> 00:15:49,536
from P(x) for the discrete random
variable. And this f(x) represents what's

178
00:15:49,536 --> 00:15:55,460
often referred to as the probability
curve, okay? Now, the probability curve

179
00:15:55,460 --> 00:16:01,642
satisfies such that, if A is any inte rval
on the real line, the probability that the

180
00:16:01,642 --> 00:16:07,064
continuous random variable is in this
interval is equal to the interval of the

181
00:16:07,064 --> 00:16:13,076
probability curve over that interval. So,
in other words, the probability that X is,

182
00:16:13,076 --> 00:16:20,001
is in this interval, is the area under the
probability curve over this particular

183
00:16:20,001 --> 00:16:26,004
interval. So, we have a continuous random
variable, we have a probability curve, and

184
00:16:26,004 --> 00:16:31,077
probabilities are associated with areas
under the curve. Now, this probability

185
00:16:31,077 --> 00:16:35,951
curve must satisfy, it always, is always
positive cuz we want to compute areas

186
00:16:35,951 --> 00:16:41,553
under a curve. And the total area under
the probability curve is equal to a 100%.

187
00:16:41,553 --> 00:16:48,491
So, that's the ideas that the, all the
probabilities add to one. So, if we think

188
00:16:48,491 --> 00:16:54,052
of a probability curve here, let's say, X
is continuous random variable, here, its

189
00:16:54,052 --> 00:16:59,261
the probability curve, and I've witnessed,
suggestively like a, a bell shape curve

190
00:16:59,261 --> 00:17:04,636
like a normal distribution which we will
talk about later. And we want to say, what

191
00:17:04,636 --> 00:17:09,660
is the probability that this random
variables between -two and one. So, our

192
00:17:09,660 --> 00:17:15,356
intervals between -two and one and the
probability of this event is equal to the

193
00:17:15,356 --> 00:17:20,862
area under the probability curve over this
interval. So, we see that one of the

194
00:17:20,862 --> 00:17:25,698
reasons why probability theory with
Calculus is useful because in order to

195
00:17:25,698 --> 00:17:30,581
calculate probabilities, we have to find
area under a curve. In order to find area

196
00:17:30,581 --> 00:17:35,544
under the curve, we have to integrate the
probability function. So, when you take a

197
00:17:35,544 --> 00:17:40,228
more mathematically-oriented probability
theory course, you do a lot of C`alculus

198
00:17:40,228 --> 00:17:44,910
to do these types of calculations. In this
class, we are not going to do the

199
00:17:44,910 --> 00:17:49,496
Calculus, calculations, right? We need to
know the concept of doing this and then

200
00:17:49,496 --> 00:17:53,651
if, even if you have to integrate
something, we can do it in r numerically.

201
00:17:53,651 --> 00:17:58,085
So, we have the function, one in front of
the area under the curve, we can write a

202
00:17:58,085 --> 00:18:03,504
function in r to represent the probability
curve and then, we can use the function

203
00:18:03,504 --> 00:18:08,506
called Integrate to numerically calculate
the area under the curve for us. So, we're

204
00:18:08,506 --> 00:18:13,726
going to be using tools that will do
Calculus for us and, but we need to know

205
00:18:13,726 --> 00:18:21,144
the concept of, of what's going on,
alright. A very simple example of a

206
00:18:21,144 --> 00:18:27,590
continuous random variable in a
distribution is, is the so-called uniform

207
00:18:27,590 --> 00:18:34,510
distribution over the interval ab. So we
say, x is distributed uniform. So, this is

208
00:18:34,510 --> 00:18:40,903
a bit of notation in a probability theory.
X represents a random variable. This

209
00:18:40,903 --> 00:18:47,622
little squiggle character, character
represent, is to be read as is distributed

210
00:18:47,622 --> 00:18:54,000
a. So, x is distributed as uniform over
the interval ab, okay? The probability

211
00:18:54,000 --> 00:19:01,624
curve of the uniform distribution is a
rectangle. So, if we want the think about

212
00:19:01,624 --> 00:19:10,459
the uniform distribution, we have some
interval a to b, and we know that the area

213
00:19:10,459 --> 00:19:18,006
under the probability curve has to equal a
100%. And the idea of a uniform

214
00:19:18,006 --> 00:19:23,634
distribution is that, you know,
probability over any interval of the same

215
00:19:23,634 --> 00:19:28,347
length is the same, okay? So, it's a, it's
a way of thinking of like, equal

216
00:19:28,347 --> 00:19:34,054
probability for events of the same size,
so to say. So, if this total area has to

217
00:19:34,054 --> 00:19:39,946
be 100%, then we know that length times
width is equal to one, so the height of

218
00:19:39,946 --> 00:19:49,219
the probability curve is one / b - a. So,
that represents the probability curve for

219
00:19:49,219 --> 00:19:56,795
a uniform end variable. And we know that
this probability curve is greater or equal

220
00:19:56,795 --> 00:20:03,001
to zero provided the, the right end point
is bigger than the left end point. And we

221
00:20:03,001 --> 00:20:08,071
know that the total area under this curve,
length times width, if we do the

222
00:20:08,071 --> 00:20:11,001
integration it's equal to one.
