1
00:00:03,620 --> 00:00:09,335
We left off last time, talking about some 
implementation issues around, 

2
00:00:09,335 --> 00:00:15,084
multithreading. 
so a little bit of background on what 

3
00:00:15,084 --> 00:00:19,702
we're doing today is we're going to be 
finishing up multithreading. 

4
00:00:19,702 --> 00:00:24,863
We're going to be using cutlery and 
you'll see why we're using cutlery in a 

5
00:00:24,863 --> 00:00:28,340
minute. 
That's right. A whole box a forks. 

6
00:00:28,340 --> 00:00:34,486
And we'll be talking about 
synchronization and synchronization 

7
00:00:34,486 --> 00:00:38,981
permitives. 
So let's, let's go back to 

8
00:00:38,981 --> 00:00:45,382
multithreading, And look a, actual 
simultaneous multithreading which is 

9
00:00:45,382 --> 00:00:48,745
where we left off. 
So just to recap, we talked about 

10
00:00:48,745 --> 00:00:53,296
different types of multithreading. 
We talked about not multithreading 

11
00:00:53,296 --> 00:00:58,309
processors, so we have a superscalar 
where we're trying to fill in the empty 

12
00:00:58,309 --> 00:01:01,738
slots here. 
So we can do some fine-grain, but not the 

13
00:01:01,738 --> 00:01:06,487
simultaneous multithreading. 
We can do coarse-grain multithreading we 

14
00:01:06,487 --> 00:01:12,180
can even possibly think about trying to 
do this in some fast software layer. 

15
00:01:12,180 --> 00:01:16,797
You could think about just cutting your 
cluster in half and using half your ALUs 

16
00:01:16,797 --> 00:01:19,989
for one thread and half your ALUs for 
another thread. 

17
00:01:19,989 --> 00:01:23,181
That is sort of technically a version of 
multithreading. 

18
00:01:23,181 --> 00:01:27,627
It's just not executing oh, it is 
executing at the same time, but it's not 

19
00:01:27,627 --> 00:01:31,902
necessarily using all the resources. 
and then you could think about actual 

20
00:01:31,902 --> 00:01:36,291
simultaneous multithreading, where you're 
issuing instructions from different 

21
00:01:36,291 --> 00:01:39,540
threads into different functional units 
at the same time. 

22
00:01:41,740 --> 00:01:47,106
So, we talked about this last time and 
we, one of the things that came up was, 

23
00:01:47,106 --> 00:01:50,072
how do you go about getting a 
parallelism? 

24
00:01:50,072 --> 00:01:55,156
So, if you have thread-level parallelism, 
you can mix and match different 

25
00:01:55,156 --> 00:02:00,800
instruction slots from different threads 
and get parallelism that way. 

26
00:02:00,800 --> 00:02:04,930
And one of the great things about 
simultaneous multithreading is that if 

27
00:02:04,930 --> 00:02:09,437
you have a fast superscalar machine and 
you try to use it for simultaneous 

28
00:02:09,437 --> 00:02:12,235
multithreading, 
you could think about trying to reduce 

29
00:02:12,235 --> 00:02:16,100
the number of threads if you don't have 
enough thread-levle parallelism and 

30
00:02:16,100 --> 00:02:18,390
actually go at instruction-level 
parallelism. 

31
00:02:18,390 --> 00:02:22,358
So here, we see this blue crosshatch. 
block zero instructions, 

32
00:02:22,358 --> 00:02:26,510
and this is the same thread running in 
the same sort of time period. 

33
00:02:26,510 --> 00:02:31,517
And, if you have lots of threads running, 
you could actually intermix the different 

34
00:02:31,517 --> 00:02:36,096
threads and run the blue crosshatch 
threads slower, than if you were try to 

35
00:02:36,096 --> 00:02:39,027
use parallel resources to run one thread 
faster. 

36
00:02:39,027 --> 00:02:44,033
And this is one of the big insights here 
that symmetric multithreading lets us go 

37
00:02:44,033 --> 00:02:47,928
after. 
So we briefly talked and flew through 

38
00:02:47,928 --> 00:02:52,485
this at the end of lecture last time 
because we talked about, what is the cost 

39
00:02:52,485 --> 00:02:57,100
of implementing symmetric multithreading. 
So, conveniently, there is actually a 

40
00:02:57,100 --> 00:03:01,361
really good example of this, 
the Power 4 by IBM and the Power 5 by IB, 

41
00:03:01,361 --> 00:03:03,668
IBM. 
They're very similar architectures, 

42
00:03:03,668 --> 00:03:07,396
except the Power 5 has two way 
simultaneous multithreading, 

43
00:03:07,396 --> 00:03:12,308
and the Power 4 is almost the same 
architecture without two way symmetric 

44
00:03:12,308 --> 00:03:15,882
multithreading. 
So you look at these two pictures. 

45
00:03:15,882 --> 00:03:20,025
One of the big things they added is they 
added more fetch bandwidth, so you can 

46
00:03:20,025 --> 00:03:23,119
fetch from two different program counters 
at the same time. 

47
00:03:23,119 --> 00:03:27,052
You can decode more instructions, and 
then, they added an extra pipe staging 

48
00:03:27,052 --> 00:03:30,775
here, which they, they do what we'll 
call, what they call group formation, 

49
00:03:30,775 --> 00:03:34,813
which is basically the scheduling of the 
two different threads instructions, 

50
00:03:34,813 --> 00:03:38,900
together by the time that it reaches the 
execution units. 

51
00:03:38,900 --> 00:03:42,027
Also over here, they, 
as, as we've talked about before, even 

52
00:03:42,027 --> 00:03:46,306
without symmetric multithreading, it's 
pretty useful to have more registers or 

53
00:03:46,306 --> 00:03:49,982
more physical registers, registers and 
more architectural registers. 

54
00:03:49,982 --> 00:03:53,493
You, you're minimally going to need more 
architectural registers, 

55
00:03:53,493 --> 00:03:56,840
more physical registers probably is a 
good performance thing. 

56
00:03:59,880 --> 00:04:02,440
So, we're talking about implementation 
details. 

57
00:04:02,440 --> 00:04:06,836
So, one of the questions that comes up 
is, what is the cost of actually going to 

58
00:04:06,836 --> 00:04:10,120
implement this symmetric multithreading 
on this processor? 

59
00:04:12,180 --> 00:04:16,085
And one of the questions that also comes 
up here is, why two threads instead of 

60
00:04:16,085 --> 00:04:18,360
four threads, or eight threads, or more 
threads? 

61
00:04:18,360 --> 00:04:21,597
It has costs. 
So, that's one of the things that they 

62
00:04:21,597 --> 00:04:25,611
would actually have to start replicating 
more data structures. 

63
00:04:25,611 --> 00:04:30,661
For the pipeline that they had roughly 
left over from the penny, excuse me, the 

64
00:04:30,661 --> 00:04:35,158
Power 4, they had enough compute and 
bandwidth through their execution units 

65
00:04:35,158 --> 00:04:39,439
to be able to handle two threads. But if 
they wanted to try to go to more threads, 

66
00:04:39,439 --> 00:04:43,720
they basically bottleneck somewhere in 
their architectures. I don't completely 

67
00:04:43,720 --> 00:04:45,940
know exactly where they turtlenecked. 
They, 

68
00:04:45,940 --> 00:04:49,847
the, the paper talks about this, it 
basically says that they bottleneck. 

69
00:04:49,847 --> 00:04:53,978
Somewhere in their execution I guess they 
kind of leave it a little vague. 

70
00:04:53,978 --> 00:04:58,276
so they decided that it wasn't worth 
adding more than two threads, because 

71
00:04:58,276 --> 00:05:02,072
it's not going to give you any 
performance increase, unless you have 

72
00:05:02,072 --> 00:05:05,812
threads that don't have or, or, threads 
that have a lot of stalls. 

73
00:05:05,812 --> 00:05:10,277
If you have threads with enough stalls, 
you probably have enough holes to try to 

74
00:05:10,277 --> 00:05:16,663
intermix another thread in there. 
So, some of the changes in the hardware 

75
00:05:16,663 --> 00:05:21,096
they were about doing. 
Well, they actually increased their cache 

76
00:05:21,096 --> 00:05:24,402
size, 
they increased the associativity of their 

77
00:05:24,402 --> 00:05:30,860
caches their last level cache here that 
went to 1.92 megabytes versus 1.44 of the 

78
00:05:30,860 --> 00:05:34,560
L2 and L3 together. Sorry. 
and you can see here one of the 

79
00:05:34,560 --> 00:05:38,877
interesting things is they start to 
actually separate data structures, 

80
00:05:38,877 --> 00:05:43,688
because if you have multiple threads 
running and you try to mix them all into 

81
00:05:43,688 --> 00:05:48,252
one big data structure, one of the 
problems like, or one of the big hardware 

82
00:05:48,252 --> 00:05:53,186
data structure, one of the problems that 
comes up is maybe one of the threads is 

83
00:05:53,186 --> 00:05:58,000
going to hog the data structure. 
So, threading actually introduces a lot 

84
00:05:58,000 --> 00:06:03,292
of complexity, having to do with figuring 
out how to be fair in the usage of the 

85
00:06:03,292 --> 00:06:06,579
structures. 
So one solution is you just replicate. 

86
00:06:06,579 --> 00:06:11,811
So anywhere there was, it was hard to get 
it right or hard to have they have a lot 

87
00:06:11,811 --> 00:06:16,270
of conflicts you could just replicate. 
So they added a separate instruction 

88
00:06:16,270 --> 00:06:20,729
fetch, prefetch and buffering. 
and then they added more physical they, 

89
00:06:20,729 --> 00:06:24,950
they call them virtual registers here, 
but it's more physical registers 

90
00:06:24,950 --> 00:06:27,566
effectively, 
so more places to schedule into. 

91
00:06:27,566 --> 00:06:30,420
So there were some added hardware costs 
to this. 

92
00:06:31,540 --> 00:06:36,109
And to give you an idea of sort of area, 
24% area improvement increase, 

93
00:06:36,109 --> 00:06:40,870
a lot of this is just due to the cache, 
because they made the cache so much 

94
00:06:40,870 --> 00:06:43,599
larger. 
So, you have to ask yourself, does it 

95
00:06:43,599 --> 00:06:48,613
make sense to use a 25, 24% more area to 
run two threads or at one point does it 

96
00:06:48,613 --> 00:06:52,520
make sense just to plop down a second 
processor? 

97
00:06:52,520 --> 00:06:56,373
It's tough trade there. 
Well, if it's 24% bigger and you get a 

98
00:06:56,373 --> 00:06:58,696
two times performance boost, it was a 
good trade. 

99
00:06:58,696 --> 00:07:03,966
If it was 24% bigger and you get less 
than a 24% performance boost, it's pretty 

100
00:07:03,966 --> 00:07:07,253
questionable. 
So let's look at another example and then 

101
00:07:07,253 --> 00:07:11,900
we'll wrap up on the performance of this 
and see if they did get their 24% 

102
00:07:11,900 --> 00:07:16,300
performance boost. 
The first implementation of symmetric 

103
00:07:16,300 --> 00:07:20,832
multithreading, there's been many 
implementations of multi threading. 

104
00:07:20,832 --> 00:07:25,940
But excuse me, not symmetric, 
simultaneous multithreading, was the 

105
00:07:25,940 --> 00:07:30,112
Pentium 4 processor. 
Now they called it, hyperthreading. 

106
00:07:30,112 --> 00:07:35,220
that's a Intel marketing term. 
it is simultaneous multithreading, 

107
00:07:35,220 --> 00:07:40,268
but the Pentium 4, they didn't actually 
replicate very many structures. 

108
00:07:40,268 --> 00:07:43,725
They didn't really change the processor 
very much. 

109
00:07:43,725 --> 00:07:48,912
And in fact there's basically modes to 
turn off simultaneous multithreading 

110
00:07:48,912 --> 00:07:54,168
they're paying for, and it was not a 
flagship feature when the processor was 

111
00:07:54,168 --> 00:07:57,141
first shipped. 
So the, the, the basic idea in the 

112
00:07:57,141 --> 00:08:02,447
Pentium 4, the, the, the little bit of 
extra hardware they added was, they did 

113
00:08:02,447 --> 00:08:08,130
duplicate some, some resources here. So, 
overall they, they increased their dye 

114
00:08:08,130 --> 00:08:13,670
area by about 5%., So very, very small 
amount. Now the question is what is the 

115
00:08:13,670 --> 00:08:21,414
performance boost they can get from this? 
[COUGH], and one of the interesting 

116
00:08:21,414 --> 00:08:28,132
things here is that there was a big 
problem that started to come up with the 

117
00:08:28,132 --> 00:08:32,473
Pentium 4 that when you were trying to 
run one thread on it, well, you weren't 

118
00:08:32,473 --> 00:08:36,597
guaranteed that the performance of that 
one thread would be equal to if you 

119
00:08:36,597 --> 00:08:40,775
turned off hyperthreading, or sorry, 
symmetric simultaneous multithreading. 

120
00:08:40,775 --> 00:08:44,194
So they had a mode switch, 
they could turn off, where they would 

121
00:08:44,194 --> 00:08:48,046
take some structures that they would 
partition when they were in 

122
00:08:48,046 --> 00:08:50,217
multithreading mode and partition them 
50/50. 

123
00:08:50,217 --> 00:08:54,395
So you didn't have deadlock problems and 
you didn't have resource contention 

124
00:08:54,395 --> 00:08:59,303
problems on the shared structures. 
And when you turn this mode switch off, 

125
00:08:59,303 --> 00:09:04,220
[LAUGH], some programs got faster, to 
some extent what that meant is, they 

126
00:09:04,220 --> 00:09:10,000
didn't size those structures correctly to 
be running in multitraining mode at all 

127
00:09:10,000 --> 00:09:13,554
times. And the switch was not a little 
switch. It was like a big switch. 

128
00:09:13,554 --> 00:09:16,398
It's like reboot the computer and change 
the switch. 

129
00:09:16,398 --> 00:09:19,240
It was it was a boot time parameter, 
parameter to the chip. 

130
00:09:19,240 --> 00:09:25,678
And this left such a bad taste in, in 
Intel's mouth that simultaneous 

131
00:09:25,678 --> 00:09:32,397
multithreading got kicked out of Intel 
land for a few generations. 

132
00:09:32,397 --> 00:09:39,395
The Pentium M, Coreduo, Core2duo all 
kicked out simultaneous multithreading. 

133
00:09:39,395 --> 00:09:47,700
And it didn't work it's way back in until 
inhalem core i5, i7 sort of processors 

134
00:09:47,700 --> 00:09:53,194
there recently. 
So, it's interesting to see that, you can 

135
00:09:53,194 --> 00:09:57,865
use a little bit of hardware, but you 
have to be careful, now, what was, what 

136
00:09:57,865 --> 00:10:00,580
was the biggest problem that they had 
here? 

137
00:10:01,800 --> 00:10:10,030
The, the biggest problem was that the 
load store queue on the Pentium four was. 

138
00:10:10,030 --> 00:10:13,679
Split in half when they were running in 
two thread mode. 

139
00:10:13,679 --> 00:10:17,328
So the first Pentium 4 architecture that 
was shipped. 

140
00:10:17,328 --> 00:10:21,564
It had symmetric or simultaneous multi 
threading turned on. 

141
00:10:21,564 --> 00:10:26,647
They just split the structure in half. 
And it didn't have enough bandwidth for 

142
00:10:26,647 --> 00:10:30,296
lots of programs. 
Or didn't have enough entries in there 

143
00:10:30,296 --> 00:10:33,229
for lots of programs when it was cut in 
half. 

144
00:10:33,229 --> 00:10:38,051
So it was sized to run one thread. 
They cut it in half statically and when 

145
00:10:38,051 --> 00:10:41,820
they tried to run two threads it was, 
effectively. 

146
00:10:41,820 --> 00:10:45,443
Not providing enough performance. 
So they couldn't get enough loads out to 

147
00:10:45,443 --> 00:10:48,920
the memory system from one thread when 
they cut that structure in half. 

148
00:10:50,540 --> 00:10:54,530
So it really starts to bring up the 
question of what is the right allocation 

149
00:10:54,530 --> 00:10:57,172
of resources. 
And, should you use some sort of round 

150
00:10:57,172 --> 00:10:59,608
robin scheme? 
Should you use dynamic allocation? 

151
00:10:59,608 --> 00:11:02,406
Should you use static allocation? 
Should you replicate? 

152
00:11:02,406 --> 00:11:06,240
If you replicate it costs more area and 
this is, these are some of the big 

153
00:11:06,240 --> 00:11:10,040
challenges here. 
Okay so let's look at the, the, some of 

154
00:11:10,040 --> 00:11:15,348
the performance here. 
So the, The numbers were pretty abysmal 

155
00:11:15,348 --> 00:11:18,370
for the Pentium 4. 
Even when. 

156
00:11:18,370 --> 00:11:20,391
Everything was going well. 
So. 

157
00:11:20,391 --> 00:11:22,204
Example here. 
Pentium 4. 

158
00:11:22,204 --> 00:11:27,154
Extreme edition with simultaneous 
multithreading give you one% speed up 

159
00:11:27,154 --> 00:11:30,500
when you're running two threads, for spec 
int. 

160
00:11:30,500 --> 00:11:33,917
Rate. 
So you're running multiple, copies of the 

161
00:11:33,917 --> 00:11:37,193
same program, that's how spec, spec rate 
is done. 

162
00:11:37,193 --> 00:11:42,143
And that was the integer of one. 
For the floating.1, it did a little bit 

163
00:11:42,143 --> 00:11:43,886
better. 
Seven% improvement. 

164
00:11:43,886 --> 00:11:48,070
Now, that's not horrible, considering it 
was only a five% area. 

165
00:11:48,070 --> 00:11:53,242
Improvement or area increase, but still 
that's, that's not saying a whole lot. 

166
00:11:53,242 --> 00:11:58,415
I mean you have to question if that's 
really, this whole level of complexity 

167
00:11:58,415 --> 00:12:02,049
was worth it. 
And was at five% area, could have been 

168
00:12:02,049 --> 00:12:05,919
used for something else to get a few 
percent performance increase. 

169
00:12:05,919 --> 00:12:10,257
And this was to some extent pretty 
typical across applications for the 

170
00:12:10,257 --> 00:12:15,292
Pentium 4. 
One of the, the interesting things about 

171
00:12:15,292 --> 00:12:20,641
the, the, the Pentium 4 was. 
Because the cash wasn't huge, especially 

172
00:12:20,641 --> 00:12:24,973
the L1 cash was relatively small because 
they wanted it to operate so fast, at the 

173
00:12:24,973 --> 00:12:28,805
very high clock frequencies. 
You get a lot of data pollution in there 

174
00:12:28,805 --> 00:12:33,359
from the two threats that would actually 
destructively interfere in the level one 

175
00:12:33,359 --> 00:12:35,859
cache. 
So that was one of the big things they 

176
00:12:35,859 --> 00:12:39,802
had to try and make up was this is a 
problem of any multi-threading. 

177
00:12:39,802 --> 00:12:44,189
Not just simultaneous multi-threading, 
that you can actually fight for space in 

178
00:12:44,189 --> 00:12:46,633
the cache. 
And end up with both capacity and 

179
00:12:46,633 --> 00:12:50,996
[INAUDIBLE]. 
The Power five did a lot better here. 

180
00:12:50,996 --> 00:12:55,840
You know it was, they, they thought a 
little bit harder about sharing 

181
00:12:55,840 --> 00:13:01,398
structures, duplicating structures, and 
allocation of different resources in 

182
00:13:01,398 --> 00:13:04,390
there. 
So on, on Specint, they, Specint rates, 

183
00:13:04,390 --> 00:13:08,024
they were getting 23% improvement for 
their 25% area. 

184
00:13:08,024 --> 00:13:11,230
Okay. 
Well, this, this might actually be a good 

185
00:13:11,230 --> 00:13:13,367
idea. 
It might have some value. 

186
00:13:13,367 --> 00:13:18,497
[COUGH] And for the floating point 
version, they are not doing great, but 

187
00:13:18,497 --> 00:13:22,420
doing better here. 
And. 

188
00:13:22,420 --> 00:13:27,286
The, the floating point apps had a lot 
of, cache conflicts, so the floating 

189
00:13:27,286 --> 00:13:32,152
point spec FP apps, if you go look at 
them in the inside, they typically have 

190
00:13:32,152 --> 00:13:36,698
very large data sets, so they were 
basically conflicting in their large, 

191
00:13:36,698 --> 00:13:39,900
larger, last little caches and things 
like that. 

192
00:13:42,080 --> 00:13:51,315
So finally I wanted to just give a little 
bit of color to this idea of picking 

193
00:13:51,315 --> 00:13:59,649
fairly and not having starvation in 
different resources in a simultaneous 

194
00:13:59,649 --> 00:14:06,962
multi trading pipeline. 
So one processor that's sort of the 

195
00:14:06,962 --> 00:14:14,759
famous, simultaneous multiframe processor 
that was never built, was the EV eight, 

196
00:14:14,759 --> 00:14:21,319
or the Digital Equipment Corporation last 
alpha, that was never built. 

197
00:14:21,319 --> 00:14:27,940
So this last alpha, the, EV eight or also 
known as the 21464, 

198
00:14:27,940 --> 00:14:33,504
Have introduced they had eight way 
threaded processors note this was never, 

199
00:14:33,504 --> 00:14:39,140
never built but, it was [INAUDIBLE] super 
scale that it had eight way, eight way. 

200
00:14:39,140 --> 00:14:43,941
Simultaneous multi-threading. 
It also was a very aggressive processor. 

201
00:14:43,941 --> 00:14:48,673
And when I say it was never built, it was 
never shipped commercially. 

202
00:14:48,673 --> 00:14:53,891
They went pretty far in the design of it. 
So one of the questions that they 

203
00:14:53,891 --> 00:14:59,458
realized upon here is, what is the right 
way in an out of order pipeline to make 

204
00:14:59,458 --> 00:15:04,538
sure that you're getting correct 
utilization of the pipeline when you're 

205
00:15:04,538 --> 00:15:07,600
issuing instructions from different 
threads? 

206
00:15:07,600 --> 00:15:12,240
So in this example here, we have four 
different threads we'll say. 

207
00:15:12,240 --> 00:15:17,809
And, you need to choose to go into your 
multi issue out of order pipeline here, 

208
00:15:17,809 --> 00:15:23,493
which thread to go pick from. 
And, by definition you want to try to 

209
00:15:23,493 --> 00:15:28,481
fill in the holes of one of the threads 
with work from another thread. 

210
00:15:28,481 --> 00:15:34,537
So, complete fairness here or lockstep 
round robin choosing for instance is not 

211
00:15:34,537 --> 00:15:37,750
what you want to do. 
Completely wrong thing to do. 

212
00:15:37,750 --> 00:15:41,725
Because you want to fill in when one 
processor is stalled, with other 

213
00:15:41,725 --> 00:15:44,664
processors work. 
But, you want to make sure that the 

214
00:15:44,664 --> 00:15:47,948
thread which let's say, has no 
dependencies on each other. 

215
00:15:47,948 --> 00:15:52,153
Or could run very fast, doesn't go to 
memory very often, we'll say, it just 

216
00:15:52,153 --> 00:15:55,760
doesn't stall very often. 
Doesn't just hog the processor. 

217
00:15:55,760 --> 00:16:00,320
Because it's never going to introduce 
stall cycles on itself by itself. 

218
00:16:00,320 --> 00:16:05,785
So they came up with this idea called 
icount, and it was a choosing policy 

219
00:16:05,785 --> 00:16:11,541
which basically looked at the number of 
instructions that were retiring out of 

220
00:16:11,541 --> 00:16:17,444
the back of the pipeline, had like a 
moving window to try to estimate how many 

221
00:16:17,444 --> 00:16:22,836
instructions from each thread were 
completing and then, feedback around to 

222
00:16:22,836 --> 00:16:28,440
the front of the pipeline here to 
determine where to issue from. 

223
00:16:28,440 --> 00:16:31,890
So I, I just wanted to get across the 
idea that, if you try to intermix 

224
00:16:31,890 --> 00:16:35,000
different threads in a data structure, 
it's relatively simple. 

225
00:16:35,000 --> 00:16:38,985
You can either cut the data structure in 
half, or have some round-robin allocation 

226
00:16:38,985 --> 00:16:41,172
there. 
It gets much more complicated when you 

227
00:16:41,172 --> 00:16:45,157
have a full processor pipeline that you 
need to figure out how to issue into that 

228
00:16:45,157 --> 00:16:48,219
processor pipeline. 
And, then you know 20 stages later, if 

229
00:16:48,219 --> 00:16:51,086
something happens, 
there's so much stuff in flight that you 

230
00:16:51,086 --> 00:16:53,661
need to, need to worry about this a 
little bit harder. 

231
00:16:53,661 --> 00:16:57,258
So here, they, they basically were 
looking at how many instructions were in 

232
00:16:57,258 --> 00:17:00,852
flight and how many instructions had 
recently committed from a thread to 

233
00:17:00,852 --> 00:17:05,477
determine what to do. And, on average 
they were hoping that they had some 

234
00:17:05,477 --> 00:17:07,340
fairness between the threads. 

