1
00:00:00,700 --> 00:00:05,370
Hi, and welcome to our special
module on the MP3 audio encoder.

2
00:00:05,370 --> 00:00:08,890
MP3 is a shorthand for MPEG layer 3, and

3
00:00:08,890 --> 00:00:14,380
MPEG is a shorthand for
the Motion Picture Expert Group.

4
00:00:14,380 --> 00:00:18,660
And what this all means is
that at one point in the 90s,

5
00:00:18,660 --> 00:00:23,488
a lot of experts got together and
agreed on a set of standards for

6
00:00:23,488 --> 00:00:26,881
video and audio compression and encoding.

7
00:00:26,881 --> 00:00:30,963
MP3 turned out to be the most
used audio digital format for

8
00:00:30,963 --> 00:00:35,000
digital audio storage,
streaming and playback.

9
00:00:35,000 --> 00:00:37,130
And today a portable music device,

10
00:00:37,130 --> 00:00:41,820
thanks to the MP3 encoding,
can store up to 30,000 songs, which

11
00:00:41,820 --> 00:00:45,610
really means you can carry your entire
music collection with you everywhere.

12
00:00:45,610 --> 00:00:50,420
So in this video we will look at the
technology behind the success of MP3, and

13
00:00:50,420 --> 00:00:53,834
we will describe in detail
how the MP3 encoder works.

14
00:00:53,834 --> 00:00:58,345
You will see how all the tools that
you have learned in our DSB class from

15
00:00:58,345 --> 00:01:02,862
Fourier transform to filtering,
from sampling into quantization,

16
00:01:02,862 --> 00:01:05,793
they all come together
in this application.

17
00:01:05,793 --> 00:01:10,000
So how does the encoding and
decoding process take place?

18
00:01:10,000 --> 00:01:13,802
Supposed you start with a discrete
time sound of signal x[n].

19
00:01:13,802 --> 00:01:18,568
This is processed by the encoder and
converted into a binary string.

20
00:01:18,568 --> 00:01:21,494
The decoder will take
that binary string and

21
00:01:21,494 --> 00:01:25,070
convert it back in to a sound signal y[n].

22
00:01:25,070 --> 00:01:26,700
The goal of this encoding and

23
00:01:26,700 --> 00:01:33,130
decoding chain is to reduce the memory
requirements to store the sound waveform.

24
00:01:33,130 --> 00:01:37,580
And the real achievement of MP3
is its ability to greatly reduce

25
00:01:37,580 --> 00:01:39,970
the amount of data needed
to encode a file, and

26
00:01:39,970 --> 00:01:45,600
this at a very reasonable trade-off with
respect to sound quality degradation.

27
00:01:45,600 --> 00:01:49,940
The data reduction is determined
by looking at the amount of memory

28
00:01:49,940 --> 00:01:53,920
that is necessary to store
the output of the encoder,

29
00:01:53,920 --> 00:01:58,230
by comparing this quantity to the amount
of memory that we should have used

30
00:01:58,230 --> 00:02:03,400
if we had wanted to store the original
signal in uncompressed format.

31
00:02:03,400 --> 00:02:09,070
And remember that an uncoded raw audio
file will require quite a bit of storage.

32
00:02:09,070 --> 00:02:14,420
For instance, if we sample at 48
kilohertz, which is the DVD standard,

33
00:02:14,420 --> 00:02:16,980
and we use 16 bits per sample,

34
00:02:16,980 --> 00:02:22,400
we will need 12 megabytes to store
a single minute of audio in stereo.

35
00:02:22,400 --> 00:02:28,020
On the other hand, a high-quality
MP3 will require just 1.5 megabytes,

36
00:02:28,020 --> 00:02:31,180
which represents almost an order
of magnitude in data reduction.

37
00:02:32,940 --> 00:02:37,850
To achieve this performance, the coding
has to be done in a very clever way, and

38
00:02:37,850 --> 00:02:43,590
one of the key ingredients in MP3 is
a model of the human auditory system.

39
00:02:43,590 --> 00:02:47,310
So MP3 does not attempt to
preserve the original framework.

40
00:02:47,310 --> 00:02:51,870
But rather it focuses on coding
the elements of the waveform

41
00:02:51,870 --> 00:02:56,580
that are most important to our way of
listening to music and hearing sounds.

42
00:02:57,720 --> 00:03:00,470
In particular,
the distortion introduced by the encoder,

43
00:03:00,470 --> 00:03:04,820
the loss of information introduced
by the encoding mechanism,

44
00:03:04,820 --> 00:03:10,902
is placed in parts of the spectrum of
the original signal that we cannot hear.

45
00:03:10,902 --> 00:03:12,679
We will see that in more
detail in just a minute.

46
00:03:13,740 --> 00:03:18,310
As we said, the origins of
MP3 date back to the 90s when

47
00:03:18,310 --> 00:03:22,420
the Moving Picture Expert Group,
in short MPEG,

48
00:03:22,420 --> 00:03:26,540
was set up by the International Standard
Organization to develop algorithms and

49
00:03:26,540 --> 00:03:30,240
standards for audio and video compression.

50
00:03:30,240 --> 00:03:34,160
And the audio compression part of
the standard, the MP3 protocol,

51
00:03:35,260 --> 00:03:40,140
had its origins in a set of
compression algorithms that had been

52
00:03:40,140 --> 00:03:45,270
developed in the 80s by
the Fraunhofer Institute in Germany.

53
00:03:45,270 --> 00:03:48,210
We see a photo of the team
here in this picture.

54
00:03:48,210 --> 00:03:51,990
The MP3 standard was quickly
embraced by the industry, and

55
00:03:51,990 --> 00:03:57,170
this widespread acceptance is what
decreed its success ultimately.

56
00:03:58,900 --> 00:04:04,170
Now let's try to understand how MP3
works using the simple block diagram.

57
00:04:04,170 --> 00:04:09,560
Your input signal x[n] enters
a bank of subband filters.

58
00:04:09,560 --> 00:04:15,400
There are 32 parallel filters
that subdivide the input signal

59
00:04:15,400 --> 00:04:22,410
into 32 independent channels that span
the full spectral range of the input.

60
00:04:22,410 --> 00:04:28,700
Each channel is then quantized
independently using a very clever method,

61
00:04:28,700 --> 00:04:32,920
and the quantized sample
are then formatted and

62
00:04:32,920 --> 00:04:34,800
encoded in a continuous bitstream.

63
00:04:35,860 --> 00:04:39,260
The quantization scheme is clever
because the number of bits

64
00:04:39,260 --> 00:04:44,500
allocated to each subband is dependent
on the perceptual importance

65
00:04:44,500 --> 00:04:49,910
of each subband with respect to
the overall quality of the audio waveform.

66
00:04:49,910 --> 00:04:51,380
In other words,

67
00:04:51,380 --> 00:04:56,600
subbands that are deemed by the
psycho-acoustic model not to be important,

68
00:04:56,600 --> 00:05:02,000
or are difficult to be perceived,
are allocated very few or no bits at all.

69
00:05:02,000 --> 00:05:06,817
Whereas the most perceptually relevant
subbands are allocated the bulk of

70
00:05:06,817 --> 00:05:08,670
the entire bit budget.

71
00:05:08,670 --> 00:05:11,960
The reason why we can safely
allocate different amounts

72
00:05:11,960 --> 00:05:17,320
of bits to the different subbands is to
be found in the so-called masking effect

73
00:05:17,320 --> 00:05:19,310
of the human auditory system.

74
00:05:19,310 --> 00:05:22,270
Suppose you have a sound with
a strong sinusoidal component,

75
00:05:22,270 --> 00:05:23,160
as in this picture here.

76
00:05:23,160 --> 00:05:25,610
The blue line represent
the spectrum of the sound, and

77
00:05:25,610 --> 00:05:29,370
here with the red dot we indicate
the strong sinusoidal component.

78
00:05:29,370 --> 00:05:34,700
When your ear listens to a sound
like this, a masking effect takes

79
00:05:34,700 --> 00:05:40,340
place whereby frequency components
in the vicinity of the dominant peak

80
00:05:40,340 --> 00:05:45,380
are not heard unless they are louder
than a given masking threshold.

81
00:05:45,380 --> 00:05:46,830
In this figure, for example,

82
00:05:46,830 --> 00:05:50,820
the masking threshold is
indicated by the red dotted line.

83
00:05:50,820 --> 00:05:55,990
And what it indicates is that anything in
the spectrum that falls below the red line

84
00:05:55,990 --> 00:05:57,100
will not be heard and

85
00:05:57,100 --> 00:06:01,530
therefore can be removed without
any loss of perceptual quality.

86
00:06:01,530 --> 00:06:04,700
Masking effects are something
that we experience every day.

87
00:06:04,700 --> 00:06:08,780
Imagine being in a perfectly quiet room,
like in your home at night.

88
00:06:08,780 --> 00:06:10,680
You can even hear your wristwatch ticking.

89
00:06:11,690 --> 00:06:16,090
But of course, you wouldn't be able to
hear that noise in normal conditions

90
00:06:16,090 --> 00:06:21,860
during the day, when a lot of other
auditory stimuli are reaching your ears.

91
00:06:21,860 --> 00:06:26,110
Although if you were to record the audio
environment and analyze its spectrum,

92
00:06:26,110 --> 00:06:30,970
you will see that it still contains the
information about your wristwatch ticking.

93
00:06:30,970 --> 00:06:35,070
The shape of the masking threshold
is a function of the loudness and

94
00:06:35,070 --> 00:06:37,310
the frequency of the dominant tone.

95
00:06:37,310 --> 00:06:40,430
And it has been determined experimentally

96
00:06:40,430 --> 00:06:44,590
by running a lot of listening
tests with human subjects.

97
00:06:44,590 --> 00:06:48,880
Masking in the human ear takes
place within the critical bands.

98
00:06:48,880 --> 00:06:52,460
And critical bands
are portions of the spectrum

99
00:06:52,460 --> 00:06:55,460
that are treated by
the ear as a single unit.

100
00:06:55,460 --> 00:06:59,610
Everything that happens within a critical
band cannot be further resolved

101
00:06:59,610 --> 00:07:03,610
by the ear, so two different frequencies
taking place in the same critical band

102
00:07:03,610 --> 00:07:05,790
are perceived as a single tone.

103
00:07:05,790 --> 00:07:09,940
There are approximately 24 critical
bands in the human ear, and

104
00:07:09,940 --> 00:07:12,680
here is a picture of their
distribution in frequency.

105
00:07:12,680 --> 00:07:16,431
As you can see,
they get wider as we go up in frequency,

106
00:07:16,431 --> 00:07:18,770
they follow a logarithmic scale.

107
00:07:18,770 --> 00:07:24,010
Which means that the resolution power of
the ear is stronger at low frequencies,

108
00:07:24,010 --> 00:07:27,650
whereas at high frequencies
we're less discriminant.

109
00:07:27,650 --> 00:07:31,280
And therefore, when we quantize
things across critical bands,

110
00:07:31,280 --> 00:07:35,990
we can probably fit more noise in the high
frequencies than in the low frequencies.

111
00:07:35,990 --> 00:07:39,125
In the end, the purpose of
the psycho-acoustic model is to compute

112
00:07:39,125 --> 00:07:46,155
the minimum number of bits that we need
to use to quantize each of the 32 subband

113
00:07:46,155 --> 00:07:51,395
filter outputs so that the perceptual
distortion is as little as possible.

114
00:07:51,395 --> 00:07:54,845
In the end we're given
a non-uniform bit allocation

115
00:07:54,845 --> 00:07:59,305
which will allocate fewer bits to
the bands where the masking is strongest.

116
00:08:00,620 --> 00:08:05,920
Interestingly enough, the specifications
of the psycho-acoustic model are not part

117
00:08:05,920 --> 00:08:10,920
of the MP3 standard, which means
that manufacturers of MP3 encoders

118
00:08:10,920 --> 00:08:16,400
can compete with better and better
versions of their psycho-acoustic model.

119
00:08:16,400 --> 00:08:18,450
In the end, the number of bits used for

120
00:08:18,450 --> 00:08:23,110
each subband is sent along with
the quantized data to the decoder, so

121
00:08:23,110 --> 00:08:27,200
it doesn't really matter how this
bit distribution has been generated.

122
00:08:27,200 --> 00:08:31,610
From a technical point of view, as you can
imagine, there are a lot of fine details

123
00:08:31,610 --> 00:08:34,660
in the inner workings of
the psycho-acoustic model and

124
00:08:34,660 --> 00:08:36,530
the bit allocation procedure.

125
00:08:36,530 --> 00:08:41,270
And we will not have time to examine
all of this in this presentation, but

126
00:08:41,270 --> 00:08:46,710
we can roughly sum up what happens inside
the psycho-acoustic model like so.

127
00:08:46,710 --> 00:08:47,682
First of all,

128
00:08:47,682 --> 00:08:53,435
remember that all processing is performed
on subsequent windows of a given length.

129
00:08:53,435 --> 00:08:57,536
So that input signal comes in and
the stream of input samples

130
00:08:57,536 --> 00:09:02,060
is cut into chunks of a given length,
say, 1,024 samples.

131
00:09:03,090 --> 00:09:07,450
An FFT is then used to estimate
the energy of the signal in each

132
00:09:07,450 --> 00:09:09,805
of the subbands computed
by the filter bank.

133
00:09:11,290 --> 00:09:15,857
For each subband we try to distinguish
between tonal and non-tonal components,

134
00:09:15,857 --> 00:09:20,650
components that have a strong sinusoidal
shape and noise-like components.

135
00:09:20,650 --> 00:09:23,100
We have looked at masking for
tonal components, but

136
00:09:23,100 --> 00:09:25,525
a similar type of masking takes place for

137
00:09:25,525 --> 00:09:29,630
non-tonal components, and we will have
to take that into account as well.

138
00:09:30,720 --> 00:09:34,770
The individual masking effect for tonal
and non-tonal components is computed for

139
00:09:34,770 --> 00:09:36,300
each critical band.

140
00:09:36,300 --> 00:09:39,740
And then these results are summed together

141
00:09:39,740 --> 00:09:45,040
to obtain a global masking curve for
the audio frame that we are analyzing.

142
00:09:45,040 --> 00:09:48,920
This masking curve is mapped
onto the 32 subbands.

143
00:09:48,920 --> 00:09:51,790
And the number of bits
that we will use for

144
00:09:51,790 --> 00:09:56,830
each subband is computed as a function
of the signal-to-mask ratio,

145
00:09:56,830 --> 00:10:00,550
the power of the signal versus
the masking power for each critical band.

146
00:10:01,700 --> 00:10:05,900
Let's now talk about the implementation
of this subband filtering in MP3.

147
00:10:07,070 --> 00:10:11,998
As we said, the input is split
across a filter bank that

148
00:10:11,998 --> 00:10:17,650
contains 32 filters isolating
different parts of the spectrum.

149
00:10:17,650 --> 00:10:22,770
These filters are implemented as
512-tap FIRs, and they're followed

150
00:10:22,770 --> 00:10:28,210
by a 32-times down-sampler to provide
the independent subband samples.

151
00:10:28,210 --> 00:10:33,232
The filter prototype is a simple
low pass with a cutoff frequency

152
00:10:33,232 --> 00:10:37,510
of pi over 64, and
a total bandwidth of pi over 32.

153
00:10:37,510 --> 00:10:43,339
The different subbands are obtained by
modulating the base filter with a cosine

154
00:10:43,339 --> 00:10:49,350
at multiples of pi over 64, and
the resulting filter bank looks like this.

155
00:10:49,350 --> 00:10:52,380
We're showing the positive
half of the frequency axis.

156
00:10:52,380 --> 00:10:57,040
This would be the first low
pass filter with an x zero.

157
00:10:57,040 --> 00:11:00,970
This is the second one,
the third, the fourth, and so

158
00:11:00,970 --> 00:11:03,420
on, covering the entire spectrum.

159
00:11:04,770 --> 00:11:08,120
Now let's go back to
the implementation of the filter bank.

160
00:11:08,120 --> 00:11:12,959
As you can see here from this block
diagram, each branch in the filter bank

161
00:11:12,959 --> 00:11:18,880
comprises an FIR filter of length 512 and
a 32-time downsampler here.

162
00:11:18,880 --> 00:11:23,970
What this means, of course, is that 31
out of 32 output samples of this filter

163
00:11:23,970 --> 00:11:28,550
are discarded, and so this is, of course,
a very wasteful implementation.

164
00:11:28,550 --> 00:11:30,850
Let's try and
make this a little bit more efficient.

165
00:11:30,850 --> 00:11:34,580
This is actually explained
in the MP3 standard.

166
00:11:34,580 --> 00:11:41,420
We start with the equation that expresses
the output of the subband number i as

167
00:11:41,420 --> 00:11:47,110
a convolution of the impulse response of
the filter for that branch with the input.

168
00:11:47,110 --> 00:11:49,800
And here you see that
the downsampling factor

169
00:11:49,800 --> 00:11:54,660
translates to a factor of 32
in front of the input index.

170
00:11:54,660 --> 00:11:59,807
We can now replace the expression for
the impulse response of the filter

171
00:11:59,807 --> 00:12:04,699
as the prototype impulse response
times the modulating factor that

172
00:12:04,699 --> 00:12:09,354
brings the filter to the proper
position in the frequency band.

173
00:12:09,354 --> 00:12:11,850
And then here we're going
to apply a little trick.

174
00:12:11,850 --> 00:12:15,646
We're going to express the index
k as the sum of two indices.

175
00:12:15,646 --> 00:12:20,336
Namely, we're going to say that k is

176
00:12:20,336 --> 00:12:25,510
equal to 64 times and index p plus q,

177
00:12:25,510 --> 00:12:32,966
where q ranges from 0 to 63 and
p ranges from 0 to 7.

178
00:12:32,966 --> 00:12:38,296
So with this split of the summation,
we can write the previous line as a double

179
00:12:38,296 --> 00:12:43,380
summation, for p that goes from 0 to 7 and
for q that goes from 0 to 63,

180
00:12:43,380 --> 00:12:48,470
as the same term, the modulation
term that we've seen before.

181
00:12:48,470 --> 00:12:51,680
The prototype impulse response and
the input.

182
00:12:51,680 --> 00:12:58,136
Where, again, we have made
the substitution k = 64 p + q.

183
00:12:58,136 --> 00:12:58,832
Okay, so

184
00:12:58,832 --> 00:13:05,170
with this trick we can actually simplify
the first term of this double summation.

185
00:13:05,170 --> 00:13:06,780
Consider the cosine term.

186
00:13:06,780 --> 00:13:12,049
We can write cosine of pi over 64 times 2i

187
00:13:12,049 --> 00:13:18,180
plus 1 times 64p plus some other term,
let's call it f,

188
00:13:18,180 --> 00:13:23,390
that depends only on i and q, and
we don't really care about that.

189
00:13:23,390 --> 00:13:27,979
Now here, 64 is canceled out and

190
00:13:27,979 --> 00:13:35,740
we're left with cosine of 2ip
pi plus p pi plus this term.

191
00:13:37,090 --> 00:13:42,380
And now, well, this is a multiple of 2 pi,
so it doesn't influence the angle.

192
00:13:42,380 --> 00:13:45,410
And here we have a multiple of pi.

193
00:13:45,410 --> 00:13:52,580
And we know that cosine of pi plus alpha
is equal to minus cosine of alpha.

194
00:13:52,580 --> 00:13:58,157
And so in the end,
what we can do is simplify this cosine

195
00:13:58,157 --> 00:14:04,599
as cosine of pi over 64,
times 2i plus 1 times q minus 16.

196
00:14:04,599 --> 00:14:09,099
And add a term minus 1 to the power
of p that we can move over to

197
00:14:09,099 --> 00:14:11,350
the second summation.

198
00:14:11,350 --> 00:14:15,445
And we have a simplified, quote unquote,
expression that looks like so.

199
00:14:15,445 --> 00:14:21,150
An outside sum here that only involves the
cosine modulation, and an inner sum here,

200
00:14:21,150 --> 00:14:24,970
which is a pre-subsampled implementation
of the filtering operation.

201
00:14:26,110 --> 00:14:28,570
If we work out the indices and

202
00:14:28,570 --> 00:14:32,960
convert this to an algorithmic procedure,
this is what we need to do.

203
00:14:32,960 --> 00:14:37,410
We will use a 512-tap
input circular buffer, and

204
00:14:37,410 --> 00:14:44,510
we will shift at each step 32 new input
audio samples, starting from the newest.

205
00:14:44,510 --> 00:14:45,577
So at any time,

206
00:14:45,577 --> 00:14:51,510
the circular buffer is holding 512
input samples in time-reversed order.

207
00:14:51,510 --> 00:14:56,170
Then we take a new 512-point buffer and
we fill it,

208
00:14:56,170 --> 00:15:01,290
sample by sample, with the product between
the prototype impulse response and

209
00:15:01,290 --> 00:15:02,950
the content of the circular buffer.

210
00:15:04,310 --> 00:15:08,740
Next we compute this
intermediate quantity here,

211
00:15:08,740 --> 00:15:14,980
which is the sum of the contents of
this new buffer, 64 points apart.

212
00:15:14,980 --> 00:15:17,810
We can do that for 63 different points.

213
00:15:19,870 --> 00:15:23,460
And if you do the math, there are seven
points that we have summed together for

214
00:15:23,460 --> 00:15:25,260
each q index.

215
00:15:25,260 --> 00:15:29,840
Finally, each subband output is given
by this sum here, where we're taking

216
00:15:29,840 --> 00:15:34,870
the intermediate quantity c[q] that
we computed before and we modulate it

217
00:15:34,870 --> 00:15:38,950
with the cosines at the frequencies
that we had defined in the beginning.

218
00:15:39,990 --> 00:15:41,600
And finally, quantization.

219
00:15:41,600 --> 00:15:44,947
This is where the great bit rate
savings are going to be achieved.

220
00:15:44,947 --> 00:15:48,920
MP3 uses uniform quantization
of subband samples.

221
00:15:48,920 --> 00:15:52,150
And the number of bits per
sample in each subband

222
00:15:52,150 --> 00:15:56,230
is determined by the psycho-acoustic
model, as we explained before.

223
00:15:56,230 --> 00:16:01,040
We also said before that MP3
works on subsequent audio frames,

224
00:16:01,040 --> 00:16:05,830
a frame being a window of input samples
that is processed independently.

225
00:16:05,830 --> 00:16:10,550
There are 36 samples per band and
per frame in the MP3 standard.

226
00:16:10,550 --> 00:16:11,755
And so

227
00:16:11,755 --> 00:16:17,235
since all of these 36 samples are going
to be quantized by the same quantizer,

228
00:16:17,235 --> 00:16:22,215
a rescaling is needed so that we are using
the full range of the quantizer.

229
00:16:22,215 --> 00:16:24,950
Remember how uniform quantization works.

230
00:16:24,950 --> 00:16:31,160
A quantizer maps an input interval
to a set of quantization levels.

231
00:16:32,400 --> 00:16:36,550
Of course, you have to make sure that
the range of your input signal matches

232
00:16:36,550 --> 00:16:38,210
the range of the quantizer.

233
00:16:38,210 --> 00:16:43,170
For instance, this quantizer expects
input to range from minus 1 to 1,

234
00:16:43,170 --> 00:16:47,940
but if your actual input only
lives in this small sub-interval,

235
00:16:47,940 --> 00:16:51,280
you will not be able to make use
of the full quantization range.

236
00:16:51,280 --> 00:16:56,760
So rescaling normally would imply
perfect renormalization of the 36

237
00:16:56,760 --> 00:17:02,045
samples by dividing the samples by
the largest sampling magnitude.

238
00:17:02,045 --> 00:17:03,125
Of course, in order for

239
00:17:03,125 --> 00:17:06,925
the decoder to then reconstruct
the actual levels of the input,

240
00:17:06,925 --> 00:17:11,480
we would have to send this normalization
factor alongside with quantized data.

241
00:17:11,480 --> 00:17:14,630
But this would require
a lot of side information.

242
00:17:14,630 --> 00:17:19,930
We would use 16 or 32 bits to encode
normalization factor and send it along.

243
00:17:19,930 --> 00:17:25,580
Instead, the MPEG standard defines
16 predefined scale factors.

244
00:17:25,580 --> 00:17:30,100
We will choose the one that best matches
the actual range of the input and

245
00:17:30,100 --> 00:17:34,750
only use 4 bits to communicate
this range to the decoder,

246
00:17:34,750 --> 00:17:38,390
thanks to the fact that these
predefined levels are set in stone.

247
00:17:39,650 --> 00:17:43,690
Finally, the actual quantization is
performed according to this formula,

248
00:17:43,690 --> 00:17:48,430
where b is the number of bits as
provided by the psycho-acoustic model,

249
00:17:48,430 --> 00:17:52,820
and Qa and Qb,
functions of the number of bits,

250
00:17:52,820 --> 00:17:56,840
are parameters that are encoded
inside the MP3 standard.

251
00:17:56,840 --> 00:17:59,160
Finally, let's listen to some examples.

252
00:17:59,160 --> 00:18:01,970
We all know that MP3 works very well, so

253
00:18:01,970 --> 00:18:07,610
what we want to concentrate on here
is the importance of the variable

254
00:18:07,610 --> 00:18:12,950
bit allocation across the subbands as
performed by the psycho-acoustic model.

255
00:18:12,950 --> 00:18:17,360
So for a fixed bit budget we
could choose to allocate the same

256
00:18:17,360 --> 00:18:19,700
number of bits to all subbands.

257
00:18:19,700 --> 00:18:22,540
This would be uniform bit allocation.

258
00:18:22,540 --> 00:18:24,910
Or we could use
a psycho-acoustic model and

259
00:18:24,910 --> 00:18:28,050
allocate the bits smartly across subbands.

260
00:18:28,050 --> 00:18:33,844
And so here are the examples,
starting with the original signal [NOISE].

261
00:18:33,844 --> 00:18:38,749
Now let's listen to the same
signal encoded with uniform bit

262
00:18:38,749 --> 00:18:40,392
allocation [NOISE].

263
00:18:40,392 --> 00:18:44,907
And finally, this is the result of
a full-fledged MP3 implementation with

264
00:18:44,907 --> 00:18:48,037
psycho-acoustically based
bit allocation [NOISE].

265
00:18:48,037 --> 00:18:50,336
Of course, in both the uniform and

266
00:18:50,336 --> 00:18:55,652
non-uniform bit allocation encoding
schemes, the target bit rate was very,

267
00:18:55,652 --> 00:18:59,798
very low in order to exacerbate
the effects of quantization.

268
00:18:59,798 --> 00:19:02,120
But the principle holds for all bit rates.

