1
00:00:01,560 --> 00:00:09,943
One very common kind of repeated structure
occurs when we have multiple objects of

2
00:00:09,943 --> 00:00:17,286
the same type. So. That, where we want to
have all these different copies of the

3
00:00:17,286 --> 00:00:22,372
objects. It's not copies of the objects,
but objects of the same type all have a

4
00:00:22,372 --> 00:00:27,973
similar, or in fact, the same probablistic
model. For reasons that we'll talk about

5
00:00:27,973 --> 00:00:33,187
momentarily, the most, one of the most
common type of, such models is called the

6
00:00:33,187 --> 00:00:39,593
plate model. Let's start by modeling
repetition so in this case imagine that

7
00:00:39,593 --> 00:00:46,068
we're repeatedly tossing the same coin
again and again so we have an outcome

8
00:00:46,068 --> 00:00:52,879
variable. And what we'd like to model is
the repetition of multiple tosses and so

9
00:00:52,879 --> 00:00:59,606
we're going to put a little box around
that outcome variable and this box which

10
00:00:59,606 --> 00:01:07,785
is called a plate. Is a way of denoting
that the outcome variable is indexed,

11
00:01:07,785 --> 00:01:14,518
which we usually don't denote explicitly
by the notion, by different tosses of the

12
00:01:14,518 --> 00:01:20,845
coin T. And the reason for calling it a
plate is because the intuition is that

13
00:01:20,845 --> 00:01:27,092
this is a stack of identical plates.
That's kind of where the idea comes from

14
00:01:27,092 --> 00:01:33,581
for a plate model. And looking at what
that model denotes is, if we have a set of

15
00:01:33,581 --> 00:01:38,854
coins, the coin tosses T<u>1 up to T<u>k. It
basically says that we have a set of</u></u>

16
00:01:38,854 --> 00:01:43,820
random variables, outcome of t<u>1 up to
outcome of t<u>k. So we've just reproduced to</u></u>

17
00:01:43,820 --> 00:01:48,471
the outcome variable in its mulitple
copies. Now what does that explicitly

18
00:01:48,471 --> 00:01:53,060
correspond to? I'm now going to do
something that we're gonna do a lot of

19
00:01:53,060 --> 00:01:58,277
times later on in the course when we talk
about learning, which is I'm going to put

20
00:01:58,277 --> 00:02:06,689
the parameters off the CPD explicitly into
the model. So this random variable theta

21
00:02:06,689 --> 00:02:13,322
is the actual CPD parameterization and I'm
puting it explicilty so that I can show

22
00:02:13,322 --> 00:02:19,794
how different variables depend on that.
And so if we have this, the parameters

23
00:02:19,794 --> 00:02:29,800
here we can see that theta is outside of
the plate. Which means that it's not

24
00:02:29,800 --> 00:02:39,486
indexed, by t. Which means it's the
same for all values of t. So what that

25
00:02:39,486 --> 00:02:47,340
means is that we have this cop-, this
parameters theta over here. And we have all

26
00:02:47,340 --> 00:02:55,194
of these outcomes depend on the exact same
parameterization. And the CPD of the

27
00:02:55,194 --> 00:03:05,761
outcome of T<u>1 is copied from this parameterization theta. Let's look at a slightly more</u>

28
00:03:05,761 --> 00:03:11,299
interesting example. Going back to our
university with multiple students, we now

29
00:03:11,299 --> 00:03:16,638
have a two variable model, where we have
intelligence and grade. And we now index

30
00:03:16,638 --> 00:03:21,909
that by different students s. Which again
indicates that we have a repetition, a

31
00:03:21,909 --> 00:03:27,180
copying of this template model. In this
case, I only made two copies for, one for

32
00:03:27,180 --> 00:03:32,725
student one, and the other one for student
two. And, once again, if we wanted to

33
00:03:32,725 --> 00:03:39,274
encode dependence on the parameters. So we
might have theta I, which represents the

34
00:03:39,274 --> 00:03:45,583
CPD for I. So we might have theta G, which
represents the CPD for G. And we would

35
00:03:45,583 --> 00:03:52,211
have exactly the same idea of theta<u>i and
theta G, where theta I influences the two</u>

36
00:03:52,211 --> 00:03:58,200
I variables and theta<u>s influences the two s
variables. And, again, they're out of the</u>

37
00:03:58,200 --> 00:04:03,666
plate. Sometimes
in many models we will include those

38
00:04:03,666 --> 00:04:09,117
parameters explicitly within the model,
but often when you have a parameter that's

39
00:04:09,117 --> 00:04:17,079
outside of all plates, we won't denote it
explicitly. So we just omit it as we did

40
00:04:17,079 --> 00:04:22,254
in this original diagram before I
annotated it. Now, just repeating the

41
00:04:22,254 --> 00:04:27,393
exact same model multiple times is not
particularly interesting. So now let's

42
00:04:27,393 --> 00:04:32,465
talk about how you can overlap different
plates or in, in, in different words,

43
00:04:32,465 --> 00:04:38,004
think about how different types of objects
in the model overlap with each other. So

44
00:04:38,004 --> 00:04:43,143
in this case we have two kinds of objects
that are universe of discourse

45
00:04:43,143 --> 00:04:48,005
We have courses, and we have
students. And courses we're gonna call

46
00:04:48,005 --> 00:04:52,426
little c and students we're gonna call
little s. And so now, let's think about

47
00:04:52,426 --> 00:04:57,134
how you might replicate variables that
correspond to properties of courses and

48
00:04:57,134 --> 00:05:02,009
variables that correspond to properties of
students. So the difficulty variable,

49
00:05:02,009 --> 00:05:06,896
belongs in the course plate, because it's
a property of a course. So it's going to

50
00:05:06,896 --> 00:05:11,783
be the difficulty of the course. And now
let's think about how we're going to put

51
00:05:11,783 --> 00:05:18,120
students in. One possibility is we're
going to nest. The students plate inside

52
00:05:18,120 --> 00:05:25,621
the course plate. Now, what that means is
that the student of each variable here,

53
00:05:25,621 --> 00:05:34,128
both of these variables are indexed by
both s and c, because when a variable is

54
00:05:34,128 --> 00:05:39,091
nested in a plate, it means that it has
the indicies of all plates that it's

55
00:05:39,091 --> 00:05:44,250
nested in. So, if, if the in, intelligence
variable is both, is in both the S plate

56
00:05:44,250 --> 00:05:49,431
and the C plate, it's going to be indexed
by both. So let's build that model, and

57
00:05:49,431 --> 00:05:54,384
see what it looks like when we, sort of,
unravel the courses and unravel the

58
00:05:54,384 --> 00:05:59,404
students. It's gonna look like that. So
we're going to have the difficulty of,

59
00:05:59,404 --> 00:06:04,291
let's say this is a two course model, and
a two student model. So we have a

60
00:06:04,291 --> 00:06:09,310
difficulty of course one and a difficulty
of course two. And now we have the

61
00:06:09,310 --> 00:06:14,330
variables in the nested plate, I and
G, and we can see that they're both

62
00:06:14,330 --> 00:06:22,224
parameterized by both student and course.
Let's think about the implications of this

63
00:06:22,224 --> 00:06:29,132
model. The implications are that the
intelligence is now a property of both

64
00:06:29,132 --> 00:06:42,148
the student's intelligence of student in course. And that, the intelligence of the

65
00:06:42,148 --> 00:06:51,870
student in a particular course, influences
the grade of the student in that course.

66
00:06:51,870 --> 00:07:01,992
And you can see that by having
this dependency model, over here. Now what

67
00:07:01,992 --> 00:07:07,765
in fact, let's think about the
implications of this. This tells us that

68
00:07:07,765 --> 00:07:13,825
there is a core specific intelligence for
every student. For every student in every

69
00:07:13,825 --> 00:07:18,321
course and they may or may not be what we
want. If you're taking radically different

70
00:07:18,321 --> 00:07:22,870
courses and one is an art class and one is
a math class and you can say that there's

71
00:07:22,870 --> 00:07:27,152
and art intelligence, their representing
skill if you will in art and you have a

72
00:07:27,152 --> 00:07:31,059
math skill that represents math
intelligence then you might want to have

73
00:07:31,059 --> 00:07:35,341
two different kinds of intelligence and
not necessarily assume that they're the

74
00:07:35,341 --> 00:07:40,038
same thing. Of course that kind of
complicates the model, and if you have a

75
00:07:40,038 --> 00:07:45,539
bunch of courses that are, in some ways
similar to each other, and take a similar

76
00:07:45,539 --> 00:07:50,178
set of skills, you might not want to have
a bunch of independent – look –

77
00:07:50,178 --> 00:07:57,728
independent random variables. Sorry.
This is, these ones. We're assuming that

78
00:07:57,728 --> 00:08:02,405
the student has two independent
intelligences representing their, the

79
00:08:02,405 --> 00:08:07,472
intelligence in the two different courses.
And in that case, you don't want the

80
00:08:07,472 --> 00:08:12,668
intelligence variable to be part of the
course plate. And so that gives us

81
00:08:12,668 --> 00:08:17,670
an alternative representation, which is
what's called, plates that are not the

82
00:08:17,670 --> 00:08:22,607
same, but that overlap. That are not
nested that overlap with each other.

83
00:08:22,607 --> 00:08:28,362
So, in this case, we have the course
plate, which is this plate over here and

84
00:08:28,362 --> 00:08:34,780
we have the student plate. Which is this
one over here, and the assumption is, that

85
00:08:34,780 --> 00:08:39,825
this difficulty is the property
only of the course. So, this is the

86
00:08:39,825 --> 00:08:45,492
difficulty. The intelligence is the
property of the student and only the grade

87
00:08:45,492 --> 00:08:51,099
is the property, that depends on both. And when we unravel this one, what

88
00:08:51,099 --> 00:08:56,350
we end up with is a model that looks like this. So

89
00:08:56,350 --> 00:09:01,670
we have, in this case, the, we only have a
single, we have a difficulty for the

90
00:09:01,670 --> 00:09:06,652
course. We have an intelligence for the
student. And over here, let's denote

91
00:09:06,652 --> 00:09:11,769
things in the intersection in green. We
have the grade of the student in the

92
00:09:11,769 --> 00:09:17,021
course depends on the difficulty of the
course, and on the intelligence of the

93
00:09:17,021 --> 00:09:24,532
student. And so now we only have a single
intelligence per student. And that is an alternative

94
00:09:24,532 --> 00:09:28,373
modeling. It's not that one of these is
right and the other is wrong. They're just

95
00:09:28,373 --> 00:09:34,358
different. And once again, just to
demonstrate a explicit parameter sharing,

96
00:09:34,358 --> 00:09:40,280
I just wanted to highlight again that the notion of parameter also applies to models

97
00:09:40,280 --> 00:09:45,869
such as this. So here we have a parameter
theta<u>D. We have a parameter, theta<u>I</u></u>

98
00:09:45,869 --> 00:09:51,325
and we have a parameter,
theta<u>G and which influences</u>

99
00:09:51,325 --> 00:09:56,914
the grade and that's shared
among all the all of the different grade

100
00:09:56,914 --> 00:10:04,259
variables. So why are these kinds
of plate models useful? So let's look at

101
00:10:04,259 --> 00:10:09,159
an example to convince ourselves that by
building these richly structured models

102
00:10:09,159 --> 00:10:13,704
that involve multiple entities, you can
actually get much more interesting

103
00:10:13,704 --> 00:10:18,309
conclusions. So let's look at this example
over here. Imagine that we have this,

104
00:10:18,486 --> 00:10:22,973
first quarter freshman came into our
university. And we'd like to figure out,

105
00:10:23,150 --> 00:10:27,873
what we can determine about him. So let's
say that, in this particular university,

106
00:10:27,873 --> 00:10:32,655
a priori we believe that most students
have high intelligence, and so this is the

107
00:10:32,655 --> 00:10:37,973
intelligence distribution. And 80 percent
are high. Now this student that we're

108
00:10:37,973 --> 00:10:44,190
going to call George took two classes. He
took Geology 101 and got an A. So,

109
00:10:44,190 --> 00:10:49,402
the probability that he is intelligent goes up. He took
CS101, didn't do so well, got a C. Well,

110
00:10:49,402 --> 00:10:53,741
the probability goes down, but it doesn't
go down to a very lower number, and that's

111
00:10:53,741 --> 00:10:57,871
because we know from the CPD for grade
that we've seen previously, that we have,

112
00:10:57,871 --> 00:11:01,844
you know, there may be other, may be
multiple reasons why a student might not

113
00:11:01,844 --> 00:11:05,870
do well in a class. For example, maybe it
was a really hard class. And so maybe

114
00:11:05,870 --> 00:11:09,706
everybody did badly. And some of you
shouldn't take this too seriously. If

115
00:11:09,706 --> 00:11:14,355
these are the only two courses that George
took, we're kinda stuck. But now lets

116
00:11:14,355 --> 00:11:18,710
think about this in a more holistic
context, or collective inference where

117
00:11:18,710 --> 00:11:23,536
we're going to think about, a number of
students taking a number of classes and

118
00:11:23,536 --> 00:11:28,185
let's imagine that we have a bunch of
grades for all of those students. So what

119
00:11:28,185 --> 00:11:33,731
we see here are the green ones are A's...
The, yellow ones are Bs, and the red ones

120
00:11:33,731 --> 00:11:39,489
are Cs. And what you see here is shorthand
for a bunch of observed grade variables.

121
00:11:39,489 --> 00:11:45,106
So I didn't put in all the little dots
that represent, all the little ovals that

122
00:11:45,106 --> 00:11:50,794
represent the grade variables. I just put
in this lines to indicate what they are.

123
00:11:50,794 --> 00:11:56,200
So you can think of this as a, as the
induced Markov network, if you will. Okay.

124
00:11:56,200 --> 00:12:02,010
So now, so now, let's think about what
kind of conclusions we can reach from this

125
00:12:02,010 --> 00:12:07,470
network. And what seems, even by looking.
Even looking at this by eye, we can see

126
00:12:07,470 --> 00:12:12,860
that a bunch of people took CS101, and
they all aced it except for our friend

127
00:12:12,860 --> 00:12:18,670
George. And furthermore, even if you look
at this guy over here who got a C in every

128
00:12:18,670 --> 00:12:24,550
other class that he took, he still managed
to ace CS101. So if we do the probablistic

129
00:12:24,550 --> 00:12:32,004
inference over this holistic model, what
we're going to get is that we are pretty

130
00:12:32,004 --> 00:12:39,731
sure that CS101 is an easy class and if
we're pretty sure about that, we are also

131
00:12:39,731 --> 00:12:46,283
pretty sure in this case where the
intelligence is low. And so we can reach

132
00:12:46,283 --> 00:12:50,624
much more form conclusion in this setting,
than we can by reasoning about

133
00:12:50,624 --> 00:12:56,024
individuals and isolation. Now this is a
toy example. But we'll see later on

134
00:12:56,024 --> 00:13:02,112
examples of collective inference that
where we have multiple interrelated

135
00:13:02,112 --> 00:13:09,465
entities, it could be related pixels in an
image, it can be related web pages in

136
00:13:09,465 --> 00:13:16,028
a website that web pages
point to each other. That if we try and

137
00:13:16,028 --> 00:13:21,325
label each entity in isolation we just 
don't get a very informed conclusion, but by

138
00:13:21,325 --> 00:13:26,583
thinking about how they all relate to each
other we get much stronger results, that

139
00:13:26,583 --> 00:13:33,104
are much more informed. So, just to
summarize the plate dependency model. The

140
00:13:33,104 --> 00:13:39,861
plate dependency model has the following
characteristics. It defines a dependency

141
00:13:39,861 --> 00:13:48,635
model for a template variable that is
indexed by a bunch of object, types. So for

142
00:13:48,635 --> 00:13:54,616
example students and courses, or, or
anything else. And we have for each of

143
00:13:54,616 --> 00:14:01,299
those template variables we have a set of
template parents. And what we have

144
00:14:01,299 --> 00:14:09,261
[inaudible] is that each of these has to
be a subset off this. So what does that

145
00:14:09,261 --> 00:14:16,897
mean? It means for example, that for the
template variable g of s, c. So this

146
00:14:16,897 --> 00:14:24,343
is, g corresponds to variable a. S and c
corresponds to the indicies, in this

147
00:14:24,343 --> 00:14:32,265
case u one and u two. And what we have is
we have two template parentswe have

148
00:14:32,265 --> 00:14:39,991
I of s and D of c. And the, stipulation that u
i is a subset

149
00:14:39,991 --> 00:14:47,512
of the variable of U<u>1 up to U<u>k.  How is this for example that</u></u>

150
00:14:47,512 --> 00:14:54,940
we cannot have, an index in the parent  that doesn't appear in the child, so for

151
00:14:54,940 --> 00:15:02,275
example, we cannot have in this, in this
model, for this reason and I'll describe

152
00:15:02,275 --> 00:15:09,723
in a minute and this for example honors.
For student s depending on the grade, of

153
00:15:09,723 --> 00:15:16,826
the student in multiple courses. And the
reason for that is that this is not a

154
00:15:16,826 --> 00:15:23,382
CPD. You have, the honors variable
depending on a potentially unbounded,

155
00:15:23,382 --> 00:15:31,792
number of parents which are all of, the
grades in which the students

156
00:15:31,792 --> 00:15:36,746
participated, and, it's not to say, that
one can one define, such a dependency

157
00:15:36,746 --> 00:15:41,567
model, in fact, there are richer languages
than plates, for which people have

158
00:15:41,567 --> 00:15:48,788
defined this notion of an aggregate,
aggregator CPD. But it's not within the

159
00:15:48,788 --> 00:15:54,535
standard paradigm of what are
traditionally called plate models. So by

160
00:15:54,535 --> 00:16:00,754
preventing that, we now have, effectively,
a traditional model, where, you know, you

161
00:16:00,754 --> 00:16:07,131
have a random variable with a finite fixed
set of parents. And so we can define a

162
00:16:07,131 --> 00:16:14,241
template CPD which we can then reuse
in a model, for any copy of this template

163
00:16:14,241 --> 00:16:20,789
variable, where a copy is, is obtained for
different instantiations of these indicies U

164
00:16:20,789 --> 00:16:27,956
So for example, so
specifically, if we have this model if we

165
00:16:27,956 --> 00:16:35,180
have this variable A of U<u>1 up to U<u>K then
for any instantiation little U<u>1 up to U<u>K</u></u></u></u>

166
00:16:35,180 --> 00:16:42,317
which are concrete instantiations of the
indicies we would have, the following

167
00:16:42,317 --> 00:16:49,280
model we would have the variable A of
u on up to u<u>k depending on</u>

168
00:16:49,700 --> 00:17:03,060
the specific. Which is potentially
confusing notation, because the sets

169
00:17:03,060 --> 00:17:08,398
are a little bit hard to understand. But
this really, just think concretely of the

170
00:17:08,398 --> 00:17:13,407
example. This exactly says that the
intellig-, that the grade of a particular

171
00:17:13,407 --> 00:17:18,877
student in a particular course depends on
the difficulty of that course, and on the

172
00:17:18,877 --> 00:17:24,150
intelligence of that student, that's all
it says, okay? So it's just a general way

173
00:17:24,150 --> 00:17:30,921
of saying that. And, this
is just the formal version of the

174
00:17:30,921 --> 00:17:36,475
statement that I made earlier that
requires the parents not to have variables

175
00:17:36,475 --> 00:17:42,171
that are not explicitly instantiated in
the child. So that we don't have a free

176
00:17:42,171 --> 00:17:47,798
floating variable that can be instantiated
in, arbitrarily many ways. So to

177
00:17:47,798 --> 00:17:52,824
summarize, plate models
are a language which allows us to define a

178
00:17:52,824 --> 00:17:57,431
template for an infinite set of Bayesian
networks. Why infinite? Because you can

179
00:17:57,431 --> 00:18:02,218
have three students, ten students, 1,000
students, a million students, an unbounded

180
00:18:02,218 --> 00:18:06,646
number of students. So there is an
infinite set of Bayesian networks that we

181
00:18:06,646 --> 00:18:11,672
can use this language to encode and each
of them induced by a different combination

182
00:18:11,672 --> 00:18:16,765
of the way many objects in our example,
for instance students of courses. The

183
00:18:16,765 --> 00:18:23,261
parameters and the structure are reused.
In, those, within, the Bayes net, and,

184
00:18:23,261 --> 00:18:27,976
across, the different Bayes nets, so, for
example, within, our university example, we

185
00:18:27,976 --> 00:18:32,810
reuse the same parameter, and, if we have
a different university, with the different

186
00:18:32,810 --> 00:18:38,670
set of student and courses, we could still
use the same parameters. These models by

187
00:18:38,670 --> 00:18:43,434
allowing us to represent an intricate
network of dependencies allow us to

188
00:18:43,434 --> 00:18:48,648
capture very richly correlated structures
in a concise way, which allows us to do

189
00:18:48,648 --> 00:18:53,928
this kind of collective inference. Which
is potentially a very powerful source for

190
00:18:53,928 --> 00:18:59,736
informed conclusions. Now I've presented
plate models, which are the, perhaps

191
00:18:59,736 --> 00:19:06,001
earliest and one of the simplest of these
languages. Which allow us to represent

192
00:19:06,001 --> 00:19:11,294
template structures. This is a simple one
for example, it has this restriction on

193
00:19:11,294 --> 00:19:16,985
the parents not, having variables that are
not instantiated in the child, and so for

194
00:19:16,985 --> 00:19:21,816
example you can't represent temporal
models here, because x of t-1 is not

195
00:19:21,816 --> 00:19:27,109
instantiated in the variable x<u>t, so
you can't have x<u>t-1 as a parent of x<u>t</u></u></u>

196
00:19:27,109 --> 00:19:32,469
not in a plate model, I mean
obviously we have languages that can do

197
00:19:32,469 --> 00:19:37,444
that, but not this one. Similarly, you
can't have the genotype [of the mother] and the genotype

198
00:19:37,444 --> 00:19:42,408
of the father affect the genotype of the
child because once again the child doesn't

199
00:19:42,408 --> 00:19:47,135
instantiate the mother and the father.
These are separate indices. And so this is

200
00:19:47,135 --> 00:19:52,040
a limited language but there's many other
languages that expand on it in different

201
00:19:52,040 --> 00:19:56,885
ways and they each have different trade
offs in terms of what they express easily

202
00:19:56,885 --> 00:20:01,908
and what they don't. And there's an entire
literature on this that we're not going to

203
00:20:01,908 --> 00:20:07,108
go into but has provided a number of very
useful languages representing these claims

204
00:20:07,108 --> 00:20:08,763
of richly structured models.
