1
00:00:00,421 --> 00:00:03,457
Hi.  In this lecture we're going to talk about

2
00:00:03,457 --> 00:00:05,505
a really simple model of aggregation.

3
00:00:05,505 --> 00:00:07,095
So, here's the thing I want to model

4
00:00:07,095 --> 00:00:09,088
I want to model a situation where I've got a group of people

5
00:00:09,088 --> 00:00:10,855
-- it could be 100, it could be 1000 --

6
00:00:10,855 --> 00:00:12,560
and each one is independently

7
00:00:12,560 --> 00:00:14,448
going to make a decision to do something.

8
00:00:14,448 --> 00:00:16,545
It could be to, y'know, go to the gym.

9
00:00:16,545 --> 00:00:18,232
It could be to go to the beach.

10
00:00:18,232 --> 00:00:20,009
It could be to go to the grocery store.

11
00:00:20,009 --> 00:00:22,007
What I want to try and understand is that

12
00:00:22,007 --> 00:00:22,764
we've got a whole bunch of people

13
00:00:22,764 --> 00:00:24,639
each one that is making these independent decisions

14
00:00:24,639 --> 00:00:27,667
What's the number of people that shows up?

15
00:00:27,667 --> 00:00:29,844
Now, to characterize that I'm going to use an idea

16
00:00:29,844 --> 00:00:31,824
called the probability distribution.

17
00:00:31,824 --> 00:00:34,713
So, to make this simple, let's suppose that there is

18
00:00:34,713 --> 00:00:37,149
a small group of people, like my family,

19
00:00:37,149 --> 00:00:38,080
which has four people in it.

20
00:00:38,080 --> 00:00:40,374
And I want to know "What's the distribution of number

21
00:00:40,374 --> 00:00:44,104
of four people who go for a walk on a given Saturday?"

22
00:00:44,104 --> 00:00:46,409
Well, if I think about the numbers could be --

23
00:00:46,409 --> 00:00:48,563
there could be 0 people that go, there could be 1,

24
00:00:48,563 --> 00:00:51,150
there could be 2, there could be 3, or it could be

25
00:00:51,150 --> 00:00:53,763
that all 4 of us decide to go for the walk, right?

26
00:00:53,763 --> 00:00:55,287
The dog would prefer if all four of us went,

27
00:00:55,287 --> 00:00:57,249
but, y'know, there's going be some number that goes.

28
00:00:57,249 --> 00:00:59,263
So, I could take -- I could keep track of data.

29
00:00:59,263 --> 00:01:01,139
I could, y'know, chart this on, like, my wall somewhere.

30
00:01:01,139 --> 00:01:03,418
I doubt we could, right?  And you can ask,

31
00:01:03,418 --> 00:01:05,475
"What's the likelihood that nobody went for a walk?"

32
00:01:05,475 --> 00:01:07,880
Maybe that's 10%.

33
00:01:07,880 --> 00:01:10,347
Now, what's the likelihood that 1 person went for a walk?

34
00:01:10,347 --> 00:01:12,742
Well that might be 15%.

35
00:01:12,742 --> 00:01:18,101
What about 2 people?  That might be 40%.

36
00:01:18,101 --> 00:01:23,014
And what about 3 people?  That might also be 15%.

37
00:01:23,014 --> 00:01:25,784
And then, what's the likelihood that 4 of us went for a walk?

38
00:01:25,784 --> 00:01:27,812
That might be, let's say, 20%.

39
00:01:27,812 --> 00:01:30,066
Now the thing to know about a probability distribution is that

40
00:01:30,066 --> 00:01:31,803
each one of these probabilities is less that one, right?

41
00:01:31,803 --> 00:01:35,411
And if we sum them up we get 25 plus 40 is 65

42
00:01:35,411 --> 00:01:37,959
plus 15 is 80; plus 20 is 100.

43
00:01:37,959 --> 00:01:39,729
So we get a total of 100%.

44
00:01:39,729 --> 00:01:41,208
So a probability distribution tells us is

45
00:01:41,208 --> 00:01:42,980
what are the different things that could happen

46
00:01:42,980 --> 00:01:45,155
-- 0,1,2,3 and 4 --

47
00:01:45,155 --> 00:01:47,655
and then it tells us the likelihood of each of those things.

48
00:01:47,655 --> 00:01:50,993
OK, so here's sorta the huge result that we're going to

49
00:01:50,993 --> 00:01:52,971
leverage to understand how things add up.

50
00:01:52,971 --> 00:01:55,291
There's a theorem called the Central Limit Theorem.

51
00:01:55,291 --> 00:01:57,409
And what the Central Limit Theorem tells us is that

52
00:01:57,409 --> 00:02:01,199
if I add up all the whole bunch of individual, independent events

53
00:02:01,199 --> 00:02:03,213
So what does 'independent' mean?

54
00:02:03,213 --> 00:02:05,401
It means my decision to go to the beach

55
00:02:05,401 --> 00:02:07,389
is independent of your decision to go to the beach,

56
00:02:07,389 --> 00:02:09,518
which is independent of your cousin Mary's

57
00:02:09,518 --> 00:02:10,750
decision to go to the beach.

58
00:02:10,750 --> 00:02:13,533
So, by independent, I mean not influenced.

59
00:02:13,533 --> 00:02:15,371
So, I don't care whether you're going to the beach or not.

60
00:02:15,371 --> 00:02:17,002
I'm  going to make my decision on my own,

61
00:02:17,002 --> 00:02:19,326
completely independent of what you decide to do

62
00:02:19,326 --> 00:02:20,972
... or your cousin Mary.

63
00:02:20,972 --> 00:02:23,268
So, what the Central Limit Theorem tells us is that

64
00:02:23,268 --> 00:02:24,983
if a whole bunch of people make a whole bunch

65
00:02:24,983 --> 00:02:27,394
of independent decisions, the distribution that we get

66
00:02:27,394 --> 00:02:29,382
has this nice bell-shaped curve.

67
00:02:29,382 --> 00:02:31,893
And this bell-shaped curve means that like

68
00:02:31,893 --> 00:02:34,591
the most likely outcome is the one right in the middle.

69
00:02:34,591 --> 00:02:36,755
So, there's a lot of structure to what happens.

70
00:02:36,755 --> 00:02:38,342
And that means that we can predict a lof ot things.

71
00:02:38,342 --> 00:02:40,077
We can tell a lot about what's going on in the world

72
00:02:40,077 --> 00:02:41,793
And that's what we're going to learn about in this lecture.

73
00:02:41,793 --> 00:02:43,015
It's going to be a lot of fun.

74
00:02:43,015 --> 00:02:45,256
To get an understanding of where these distributions

75
00:02:45,256 --> 00:02:48,236
come from, let's start really simple.

76
00:02:48,236 --> 00:02:50,114
Suppose I flip a coin twice.

77
00:02:50,114 --> 00:02:53,257
And I want to know "What are the odds of getting a head?"

78
00:02:53,257 --> 00:02:55,760
What's the probability distribution over heads.

79
00:02:55,760 --> 00:02:57,570
Well, what could I get?

80
00:02:57,570 --> 00:02:59,981
I could get  tails-tails, and that would be 0 heads.

81
00:02:59,981 --> 00:03:03,742
I could get tails-heads, or heads-tails

82
00:03:03,742 --> 00:03:07,215
both of these would be 1 head.

83
00:03:07,215 --> 00:03:10,033
Or I could get heads-heads.

84
00:03:10,033 --> 00:03:12,079
And that would be 2 heads.

85
00:03:12,079 --> 00:03:13,650
So, what's the probability of each of these?

86
00:03:13,650 --> 00:03:15,709
The probability of getting tails-tails is just 1/4.

87
00:03:15,709 --> 00:03:18,562
The probability of getting 1 head is 1/2.

88
00:03:18,562 --> 00:03:21,194
And the probability of getting 2 heads is 1/4.

89
00:03:21,194 --> 00:03:22,755
So, I'm going to get a probability distribution,

90
00:03:22,755 --> 00:03:25,009
if I do it out like this 0, 1, 2

91
00:03:25,009 --> 00:03:28,662
There's a 1/4 chance of that and a 1/2 chance of that

92
00:03:28,662 --> 00:03:29,697
and a 1/4 chance of that

93
00:03:29,697 --> 00:03:31,601
You notice, it sorta looks like a little bell curve.

94
00:03:31,601 --> 00:03:34,866
OK.  Let's suppose I flip it 4 times.

95
00:03:34,866 --> 00:03:36,446
Well, it gets harder.

96
00:03:36,446 --> 00:03:38,863
I could think, OK, what are the odds of getting no heads?

97
00:03:38,863 --> 00:03:43,085
I could get tails-tails-tails-tails

98
00:03:43,085 --> 00:03:45,496
well, how do I figure out the probability of that?

99
00:03:45,496 --> 00:03:49,539
Well, 1/2 time 1/2 times 1/2 times 1/2 -- 4 one halves --

100
00:03:49,539 --> 00:03:52,759
that's 2 times 2 times 2 times 2 ... so that's 1/16.

101
00:03:52,759 --> 00:03:55,480
What are the odds of getting one head?

102
00:03:55,480 --> 00:04:00,176
Well, I could get the head first, and then 3 tails...

103
00:04:00,176 --> 00:04:01,650
I could get it second,

104
00:04:01,650 --> 00:04:04,564
I could get it third,

105
00:04:04,564 --> 00:04:07,026
and it could come last.

106
00:04:07,026 --> 00:04:10,759
So, there's four places that it could show up.

107
00:04:10,759 --> 00:04:12,699
So that means there's a 4/16 chance.

108
00:04:12,699 --> 00:04:15,339
Well, I could do all sorts of math again for

109
00:04:15,339 --> 00:04:17,152
what are the odds of getting two heads?

110
00:04:17,152 --> 00:04:19,314
And I'd actually get 6/16.

111
00:04:19,314 --> 00:04:22,557
And 3 heads, well, that's the same as getting 1 head

112
00:04:22,557 --> 00:04:24,631
really, right, because tails and heads are interchangable.

113
00:04:24,631 --> 00:04:27,705
So what I'd get, I'd get this again --

114
00:04:27,705 --> 00:04:31,234
If I drew this distribution out, I'd get a peak at 2 heads, right?

115
00:04:31,234 --> 00:04:33,110
I'm going to get a nice bell curve, right?

116
00:04:33,110 --> 00:04:34,434
So, I'm going to get this thing where

117
00:04:34,434 --> 00:04:36,372
there's very little chance of getting no heads

118
00:04:36,372 --> 00:04:38,386
not that much chance of getting 4 heads

119
00:04:38,386 --> 00:04:40,816
but the most likely thing is getting 2 heads.

120
00:04:40,816 --> 00:04:42,753
So I can count all this stuff and it's fun...

121
00:04:42,753 --> 00:04:48,949
big data, lots of data,  and we want to try

122
00:04:48,949 --> 00:04:52,415
But, here is the problem:

123
00:04:52,415 --> 00:04:53,729
Remember we have talk about it. [inaudible]

124
00:04:53,729 --> 00:04:54,562
and understand it.

125
00:04:54,562 --> 00:04:55,934
Often we have more than 2 or 4, we have 'n'

126
00:04:55,934 --> 00:04:57,228
and that is a huge number.

127
00:04:57,228 --> 00:04:58,656
So if we're talking about New York City that can be 10 million people.

128
00:04:58,656 --> 00:05:00,508
If we're talking about Ann Harbour, where I live, that's still like a hundred thousand people.

129
00:05:00,508 --> 00:05:04,148
So I don't want to be sitting there writing tails, tails, tails, tails, tails a hundred thousand times.

130
00:05:04,148 --> 00:05:07,082
I want to have a model that will help me explain it.

131
00:05:07,082 --> 00:05:15,019
So what you can do is if you have n things, the mean, the expected number should be

132
00:05:15,019 --> 00:05:17,582
N over 2, right, should be half of n.

133
00:05:17,582 --> 00:05:21,774
But what we'd like to do is understand sort of what that distribution looks like.

134
00:05:21,774 --> 00:05:26,372
Well, what we know from statistics is that distribution is actually gonna be a nice bell-curve

135
00:05:26,372 --> 00:05:31,752
and the mean, right in the middle of this thing, is gonna be N/2 and this just gonna

136
00:05:31,752 --> 00:05:34,355
sort of flow out nice and symmetrically from each side

137
00:05:34,355 --> 00:05:38,381
Now there's a fancy equation, a formula that tells you what this line looks like.

138
00:05:38,381 --> 00:05:40,969
We're not going to get into that but if you take the Statistics class

139
00:05:40,969 --> 00:05:43,693
which I'd encourage you do - it's a lot of fun - you could learn exactly

140
00:05:43,693 --> 00:05:45,876
what this formula is and how it works, OK?

141
00:05:45,876 --> 00:05:49,358
We just wanna use it as a model for understanding how things aggregate.

142
00:05:49,358 --> 00:05:52,533
So, we're gonna take some leaps ahead in statistics.

143
00:05:52,533 --> 00:05:56,188
Here's the trick though, we gotta be a little bit careful.

144
00:05:56,188 --> 00:06:00,436
Flipping a coin is always equally likely, it's either a head or a tail, each one is 50/50

145
00:06:00,436 --> 00:06:02,937
But if I'm worried about people going to the beach

146
00:06:02,937 --> 00:06:07,315
right, or people going to the supermarket, or people showing up for their flight

147
00:06:07,315 --> 00:06:10,037
that's not a 50/50 proposition, right.

148
00:06:10,037 --> 00:06:12,521
So maybe 90% of people may show up for their flight

149
00:06:12,521 --> 00:06:15,767
and maybe only 10% of people of 15% go out to the beach.

150
00:06:15,767 --> 00:06:19,101
So I'd like to change that 1/2 into something else

151
00:06:19,101 --> 00:06:21,874
Well, I can introduce something called the binomial distribution

152
00:06:21,874 --> 00:06:26,745
where instead of having 1/2, that gives some probability p of doing the thing.

153
00:06:26,745 --> 00:06:29,847
So let's suppose going to the beach happens 15% of the time

154
00:06:29,847 --> 00:06:39,373
Well then, if I had a 1000 people, and p = 15%, then p times N is a 150

155
00:06:39,373 --> 00:06:42,156
so I expect to have a 150 show up.

156
00:06:42,156 --> 00:06:45,130
So that makes sense but then I can ask well, what's the distribution now

157
00:06:45,130 --> 00:06:49,778
I mean 150 is the average but I could have 200, I could have 74

158
00:06:49,778 --> 00:06:53,435
Well again, what the central limit theorem tells us is that we're

159
00:06:53,435 --> 00:06:56,133
gonna get a nice bell-curve, right here you got this nice shape here

160
00:06:56,133 --> 00:06:59,635
[inaudible] with the mean here, which is p times N

161
00:06:59,635 --> 00:07:01,442
well this will be provided if N is big enough, right

162
00:07:01,442 --> 00:07:06,575
but if I can get a pretty large N, you're gonna get this nice bell-curve and the mean is gonna be right at p times N

163
00:07:06,575 --> 00:07:14,134
Okay, there's more of that. Here's where it gets a little bit complicated but also interesting.

164
00:07:14,134 --> 00:07:20,013
There's something called the standard deviation and this is this thing called sigma

165
00:07:20,013 --> 00:07:24,789
which is [inaudible] called the standard deviation. Now, when I draw the normal curve,

166
00:07:24,789 --> 00:07:28,745
there's gonna be a mean, that's this point right here at the center

167
00:07:28,745 --> 00:07:34,255
And then there's gonna be a standard deviation which basically tells us how far spread out

168
00:07:34,255 --> 00:07:35,830
that curve is

169
00:07:35,830 --> 00:07:39,862
And what I mean by that is how far spread out the different outcomes are

170
00:07:39,862 --> 00:07:44,109
So it turns out there's this nice structure to any normal distribution. If you tell me the mean,

171
00:07:44,109 --> 00:07:46,340
and then you tell me the standard deviation,

172
00:07:46,340 --> 00:07:49,997
it's always gonna be the case that 68%

173
00:07:49,997 --> 00:07:56,832
of all outcomes will be between -1 and +1 standard deviation

174
00:07:56,832 --> 00:07:59,413
So, if it's got a big standard deviation

175
00:07:59,413 --> 00:08:00,753
That means that that range could be really wide

176
00:08:00,753 --> 00:08:04,024
If it's got a small standard deviation that means that range should be really tight but

177
00:08:04,024 --> 00:08:06,263
if you tell me the mean and tell me the standard deviation

178
00:08:06,263 --> 00:08:12,136
it's always gonna be the case that 68% of the time I'm between -1 and +1 standard deviation

179
00:08:12,136 --> 00:08:19,187
Now, in fact since that's true for one, it's also gonna be true for 2, true for 3 and true for 4, right

180
00:08:19,187 --> 00:08:24,635
So there's gonna be a 95% chance I'm within 2 standard deviations

181
00:08:24,635 --> 00:08:29,234
So wait, why do we care about this, why do we care about this stuff

182
00:08:29,234 --> 00:08:33,763
Here's why. Now I got this model that says if I add up a bunch of independent events,

183
00:08:33,763 --> 00:08:37,809
here's what the mean is, right.

184
00:08:37,809 --> 00:08:39,962
Now, in a second, I'm going to show you the formula for the standard deviation

185
00:08:39,962 --> 00:08:42,604
so then it'll tell you what sigma is here

186
00:08:42,604 --> 00:08:44,935
Well, if you know the mean and you know sigma,

187
00:08:44,935 --> 00:08:50,210
then I can give you range and I can tell you, you know, that 95% of the time

188
00:08:50,210 --> 00:08:54,176
I'm gonna be between -2 sigma and +2 sigma

189
00:08:54,176 --> 00:08:57,060
So if I said the mean number of people that showed up is a hundred

190
00:08:57,060 --> 00:09:02,812
and that's the mean, right, and the standard deviation is only 2

191
00:09:02,812 --> 00:09:09,012
Well then you'd know 95% of the time, you're gonna be between 96 and 104

192
00:09:09,012 --> 00:09:12,786
So you'd know, okay, I should prepare for pretty much exactly 100 people

193
00:09:12,786 --> 00:09:16,098
If I told you the standard deviation was 15,

194
00:09:16,098 --> 00:09:19,575
then you'd know it can be anywhere between 70 to130

195
00:09:19,575 --> 00:09:22,313
So that's what we want to try and use this model to explain [inaudible]

196
00:09:22,313 --> 00:09:26,727
how wide a range of outcomes we're likely to see in any particular setting

197
00:09:26,727 --> 00:09:31,030
So let's go back to our simple binomial distribution where the probability was 1/2

198
00:09:31,030 --> 00:09:33,733
The mean, remember, is just N over 2

199
00:09:33,733 --> 00:09:38,720
the standard deviation is the square root of N over 2

200
00:09:38,720 --> 00:09:41,132
Well, you can do a little bit of math and show that

201
00:09:41,132 --> 00:09:43,496
So, let's suppose I have N = 100

202
00:09:43,496 --> 00:09:47,593
So if N = 100, that tells me the mean is gonna be 50 so if I flip a coin a hundred times,

203
00:09:47,593 --> 00:09:51,387
guess what, the average is 50, no surprise

204
00:09:51,387 --> 00:09:54,202
But, standard deviation is the square root of N/2

205
00:09:54,202 --> 00:09:56,047
What's the square root of 100, that's 10

206
00:09:56,047 --> 00:09:59,032
So this is 10 over 2 so that gives 5

207
00:09:59,032 --> 00:10:02,359
So what that tells me is if I think in binomial distribution

208
00:10:02,359 --> 00:10:03,758
right, if I draw this thing out

209
00:10:03,758 --> 00:10:07,783
I've got a mean of 50, and then I've got a standard deviation of 5

210
00:10:07,783 --> 00:10:12,827
so that means between 55 and 45 -- 68% of all outcomes

211
00:10:12,827 --> 00:10:16,385
So, if you want, you can do this at home -- it'll take a while -- flip a coin a hundred times

212
00:10:16,385 --> 00:10:19,949
Count how many heads you get. Flip it again, count how many heads again

213
00:10:19,949 --> 00:10:26,069
Do that a whole bunch of times, you'll find that 68% of the time, you get between 45 heads and 55 heads

214
00:10:26,069 --> 00:10:30,945
So what this model gives us is it gives a sense of how strange of outcomes we'll get

215
00:10:30,945 --> 00:10:37,277
So, we know that most of the time, 68% of the time, we'll be between 45 and 55, right

216
00:10:37,277 --> 00:10:46,657
So, our mean is 50, 1 standard deviation is 55 and 45, that means 2 standard deviations is 60 and 40

217
00:10:46,657 --> 00:10:53,316
What that tells us is 95% of the time, you're gonna be between 40 and 60 heads

218
00:10:53,316 --> 00:11:00,059
And, 99% of the time, you're gonna be between 35 and 65

219
00:11:00,059 --> 00:11:07,449
So basically it's, you're almost never gonna throw fewer than 35 heads, and never throw more than 65 heads

220
00:11:07,449 --> 00:11:10,299
And so this is what sort of power that Central Limit Theorem is, right

221
00:11:10,299 --> 00:11:14,671
It gives us a sense of not only the average, but also what the spread will be

222
00:11:14,671 --> 00:11:18,774
Okay, remember this is a simple case. This is the P = 1/2 case

223
00:11:18,774 --> 00:11:23,656
And what we'd like is we want it for the more general case where the probability of something happening can be anything

224
00:11:23,656 --> 00:11:24,989
Right, this is this p over N thing

225
00:11:24,989 --> 00:11:33,866
What turns out here, we're okay because the standard deviation is just p times 1 - p times N then square root the whole thing

226
00:11:33,866 --> 00:11:40,828
So the case where p = 1/2 right, then we have the square root of 1/2 times 1/2 times N

227
00:11:40,828 --> 00:11:47,334
But notice I've got a 1/2 squared here inside so we can just pull that outside so it's just 1/2 the square root of N

228
00:11:47,334 --> 00:11:50,195
so that's where that square root of N over 2 came from

229
00:11:50,195 --> 00:11:53,959
So now, for the binomial distribution, I've got this clean formula as well

230
00:11:53,959 --> 00:12:00,476
And we can use that to model and understand stuff that's a little bit more interesting than just flipping a coin

231
00:12:00,476 --> 00:12:03,647
Let's a real example, let's have some fun

232
00:12:03,647 --> 00:12:07,486
So, how, most of us have probably been bumped off a plane before

233
00:12:07,486 --> 00:12:11,573
You show up at the airport and there's like too many people showed up for the plane

234
00:12:11,573 --> 00:12:15,438
And you think why did they do this, but the reason that they sometimes have to [inaudible] is they oversell

235
00:12:15,438 --> 00:12:19,755
And the reason they oversell tickets is because not everybody shows up

236
00:12:19,755 --> 00:12:26,450
So if you're running an airline and you've got 400 seats, and you know people show up, you know, 90% of the time

237
00:12:26,450 --> 00:12:32,659
You want to sell more than those 400 seats, right, so that your plane is pretty much full

238
00:12:32,659 --> 00:12:37,754
So let's do an example. Let's suppose, make it simple, that our plane [inaudible] got 380 seats

239
00:12:37,754 --> 00:12:41,724
So let's suppose we got a Boeing 747 with 380 seats

240
00:12:41,724 --> 00:12:45,047
Let's suppose that 90% of the time, people show up

241
00:12:45,047 --> 00:12:47,969
So we've gathered, we run an airline, we've gathered lots of data

242
00:12:47,969 --> 00:12:52,549
We pretty much know 90% of the time, people show up and that it's independent

243
00:12:52,549 --> 00:12:56,045
So one person's decision to show up doesn't [inaudible] have anything to do with anybody else's

244
00:12:56,045 --> 00:13:00,661
Now, that might not be true, right. Because if it's snowy, if I'm late, you're likely to be late

245
00:13:00,661 --> 00:13:06,557
But let's just suppose that these things are independent. And let's suppose that we sell 400 tickets

246
00:13:06,557 --> 00:13:10,652
Now we're trying to get some understanding why, what is that mean

247
00:13:10,652 --> 00:13:15,086
What's the likelihood that if we sell 400, that we're going to have more than 380 people show up

248
00:13:15,086 --> 00:13:21,020
Here's where the model can help us. It'll be able to tell us what the mean is, it will also tell us what the standard deviation is

249
00:13:21,020 --> 00:13:31,060
So the mean, right if I sell 400 tickets, and on average 90% of people show up, that means I should sell on average 360 tickets

250
00:13:31,060 --> 00:13:39,092
That's less than 380 seats but it should be fine, but what I care about is more than 380 people show up

251
00:13:39,092 --> 00:13:40,773
'cause they're gonna be like, I paid for this to go to Florida, I want to go to Florida, I don't want to be bumped

252
00:13:40,773 --> 00:13:42,394
So more than 380 show up, guess what, they're gonna be mad, right

253
00:13:42,394 --> 00:13:51,255
So the 360 doesn't tell us enough, we want to know something about the distribution

254
00:13:51,255 --> 00:13:54,627
Okay, well look, we've got a formula, right, remember

255
00:13:54,627 --> 00:14:02,298
So N was 400, and p was .9 so p times N is 360, that's our mean

256
00:14:02,298 --> 00:14:06,729
Now, the standard deviation we can solve for pretty easily. That's just the square root

257
00:14:06,729 --> 00:14:13,058
of p, which is .9 times 1-p, which is .1, times N which is 400

258
00:14:13,058 --> 00:14:16,209
So if we multiply right out that's .9 times .1 times 400

259
00:14:16,209 --> 00:14:23,717
.1 times 400 is 40, times .9 is 36, that gives the squared of 36, which is 6

260
00:14:23,717 --> 00:14:33,180
So 6, is our standard deviation. Now I get a bell-curve with a mean of 360 and a standard deviation of 6

261
00:14:33,180 --> 00:14:38,287
Well, that's useful, that can help us 'cause let's go back and let's look

262
00:14:38,287 --> 00:14:49,435
That means our mean's 360, our standard deviation is 6, so that means 68% of the time, we're gonna between 354 and 366. That's great.

263
00:14:49,435 --> 00:14:56,525
It means that 95% of the time, we'll between 348 and 372, also great

264
00:14:56,525 --> 00:15:03,356
It means 99.75% of the time, we'll be between 378 and 342.

265
00:15:03,356 --> 00:15:12,997
Well, how many seats do we have, we have 380 seats, so this means that 99.75 -- actually more than that, right

266
00:15:12,997 --> 00:15:16,390
More than 99.75% of the time we won't overbook.

267
00:15:16,390 --> 00:15:19,903
So here's the Central Limit Theorem, let's, let's say it formally.

268
00:15:19,903 --> 00:15:23,475
Central Limit Theorem [inaudible] is the following. We got a whole bunch of random variables

269
00:15:23,475 --> 00:15:28,237
so those could be decisions to show up to a flight or not so in most case the random variables are just 1s and 0s

270
00:15:28,237 --> 00:15:34,005
Or they could be, you know, the weight of your bag. Each person's weight of their bag is [inaudible] independent variable.

271
00:15:34,005 --> 00:15:43,194
As long as those things are independent, so that means it, each person's decision doesn't depend on somebody else's or how much stuff I jam in my bag doesn't affect how much stuff you jam in your bag

272
00:15:43,194 --> 00:15:49,888
And that those things have finite variance -- what does that mean -- that means that they're bounded

273
00:15:49,888 --> 00:15:56,068
So we know we can't have super huge values, like so my bag couldn't weigh billions and billions of pounds

274
00:15:56,068 --> 00:16:01,609
So long as there's sort of you know, the possible range of [inaudible] that each one can take is bounded in some way

275
00:16:01,609 --> 00:16:06,209
Or doesn't with some high probability take huge, huge values then when you add those things up

276
00:16:06,209 --> 00:16:12,677
When you sum them up, you're gonna get a normal distribution which means a bell-curve, which means we can predict stuff

277
00:16:12,677 --> 00:16:16,255
We can use that model [inaudible] make sense of how the world works

278
00:16:16,255 --> 00:16:22,506
Now, let's step back for just a second and think about like, why this is so cool

279
00:16:22,506 --> 00:16:25,838
Suppose it weren't true, here's a little thought experiment

280
00:16:25,838 --> 00:16:28,394
Suppose it were the case that when I added up a bunch of independent events

281
00:16:28,394 --> 00:16:35,767
Most of the time, I get something nice, then there were some spiky probability of some huge event

282
00:16:35,767 --> 00:16:39,928
over here. What would this mean. Well, this would mean like sometimes

283
00:16:39,928 --> 00:16:44,266
you go to the grocery store, and there would be like, 1000 people there

284
00:16:44,266 --> 00:16:48,626
or sometimes you'd be like I'm just gonna run to the bathroom and there'd be 300 people in line, right

285
00:16:48,626 --> 00:16:53,527
A lot of the predictability of the world, a lot of the predictability of these sort of daily comings and goings

286
00:16:53,527 --> 00:16:57,841
stem from the fact that this can't happen and that we get these nice bell-curves

287
00:16:57,841 --> 00:17:02,390
Because if individual people, individual firms, individual groups of people

288
00:17:02,390 --> 00:17:05,409
make decisions that don't depend on what other people decide

289
00:17:05,409 --> 00:17:08,229
[inaudible] independent decisions, then what you're gonna get is

290
00:17:08,229 --> 00:17:11,419
you're gonna get sort of nice, regular stuff, according to a bell-curve

291
00:17:11,419 --> 00:17:13,845
Yeah, sure there'll be traffic jams, sure there'll be a lot of people at the mall

292
00:17:13,845 --> 00:17:17,518
There will be days where you get a lot coming on, and there'll be days where nothing much is going on

293
00:17:17,518 --> 00:17:22,627
But most of the time, you're gonna get things in that little region which is gonna be predictable and understandable

294
00:17:22,627 --> 00:17:26,350
Now, is everything normally distributed? No, it's not.

295
00:17:26,350 --> 00:17:29,873
What about stock returns? If you look at stock returns, you'll actually see that

296
00:17:29,873 --> 00:17:35,783
there's far too many days where really nothing happens, and there's far too many days where there's huge gains

297
00:17:35,783 --> 00:17:37,826
and far too many days where there's huge losses

298
00:17:37,826 --> 00:17:41,836
And what's going on there is this is that the actions are no longer independent

299
00:17:41,836 --> 00:17:46,387
For example, prices are going up, a lot of people may buy and that's going to cause prices to go up even further

300
00:17:46,387 --> 00:17:50,956
And if prices start to fall, people may sell and that can cause prices to fall even further

301
00:17:50,956 --> 00:17:55,168
So when events fail to become independent, fail to satisfy independence's assumption,

302
00:17:55,168 --> 00:18:00,400
then we can get more big events than we'd expected and more small events that we'd expected

303
00:18:00,400 --> 00:18:03,661
So let's put a [inaudible] on this, what have we got

304
00:18:03,661 --> 00:18:07,229
If we use the Central Limit Theorem as a model, and we use this model to explain

305
00:18:07,229 --> 00:18:10,531
how if we add up a bunch of independent events, then what we get is

306
00:18:10,531 --> 00:18:12,861
we get a nice normal distribution, right

307
00:18:12,861 --> 00:18:15,891
And we can understand the mean, we can understand the standard deviation

308
00:18:15,891 --> 00:18:19,901
We can use that to predict how likely things are to occur, right

309
00:18:19,901 --> 00:18:24,023
We also learn that like, it's that independence that gives us that normality, right

310
00:18:24,023 --> 00:18:27,236
Without independence, we could get really big events, really small events

311
00:18:27,236 --> 00:18:29,273
We can get all sorts of strange stuff happening

312
00:18:29,273 --> 00:18:32,810
So where we're gonna go next, I'm gonna get take, there's a brief lecture on something called

313
00:18:32,810 --> 00:18:36,278
the Six Sigma that pushes this idea sort of, the predictability of the system a little bit

314
00:18:36,278 --> 00:18:39,348
further than we had before. But then after that, we're gonna start

315
00:18:39,348 --> 00:18:41,717
y'know, we're [inaudible]

316
00:18:41,717 --> 00:18:44,359
having systems where there's interdependent actions

317
00:18:44,359 --> 00:18:48,046
and we have those interdependent actions, we're no longer going to get these sort of nice bell-curves

318
00:18:48,046 --> 00:18:50,133
We're going to get all sorts of really interesting, strange stuff

319
00:18:50,133 --> 00:18:52,300
It's going to be a lot of fun. Alright, thank you.
