1
00:00:01,990 --> 00:00:05,900
Hi and
welcome to Module 8.1 on Image Processing.

2
00:00:05,900 --> 00:00:06,670
In this module,

3
00:00:06,670 --> 00:00:12,790
we will look at images as a particular
class of multidimensional digital signals.

4
00:00:12,790 --> 00:00:17,070
We will look at different ways to
represent these two-dimensional signals.

5
00:00:17,070 --> 00:00:21,230
And we will explore some
basic instances of images and

6
00:00:21,230 --> 00:00:22,820
operators that we can apply to them.

7
00:00:23,890 --> 00:00:26,790
Before we start let's make
the acquaintance of this cute little dog.

8
00:00:26,790 --> 00:00:31,590
We decided to boost the ratings of our
class by using the picture of a puppy for

9
00:00:31,590 --> 00:00:34,850
all the image examples that
we will use in the following.

10
00:00:34,850 --> 00:00:39,170
But even behind a puppy,
there's a hard mathematical reality.

11
00:00:39,170 --> 00:00:45,040
And so digital images can be
expressed as a two-dimensional

12
00:00:45,040 --> 00:00:51,620
signal x[n1,n2], where n1 and
n2 are integers numbers, integer indices.

13
00:00:51,620 --> 00:00:56,000
And each combination of n1 and
n2 indicates a point on a grid.

14
00:00:56,000 --> 00:01:00,980
We know already from every day life
that we call these points pixels,

15
00:01:00,980 --> 00:01:02,440
as picture elements.

16
00:01:03,860 --> 00:01:06,350
Usually the grid is regularly spaced, so

17
00:01:06,350 --> 00:01:11,840
we have a regular arrangement
of points on the grid.

18
00:01:11,840 --> 00:01:17,550
And the value of the signal at coordinates
(n1, n2) refers to the pixel's appearance.

19
00:01:17,550 --> 00:01:18,630
Now what do we mean by that?

20
00:01:19,920 --> 00:01:21,680
Take for instance a grayscale image.

21
00:01:21,680 --> 00:01:23,510
A black and white picture.

22
00:01:23,510 --> 00:01:27,870
In digital format, we can encode
the appearance of each pixel in the image

23
00:01:27,870 --> 00:01:31,730
by associating the grayscale
level to a scalar value.

24
00:01:31,730 --> 00:01:34,750
The appearance will vary
between a lowest value,

25
00:01:34,750 --> 00:01:38,900
which we associate to black,
to a highest possible value.

26
00:01:38,900 --> 00:01:42,570
Which we associate to white and
all gray levels in between.

27
00:01:42,570 --> 00:01:45,620
For color images we need something
a little bit more complicated.

28
00:01:45,620 --> 00:01:48,890
We need to use the concept
of a color space.

29
00:01:48,890 --> 00:01:51,982
Now the theory of color space is
a very fascinating subject but

30
00:01:51,982 --> 00:01:55,875
it's way too complicated for
even a cursory introduction here.

31
00:01:55,875 --> 00:01:59,505
So we will just rely on your everyday
experience, you're all familiar with

32
00:01:59,505 --> 00:02:03,995
the RGB color model for instance, which
is used in monitors for your computer.

33
00:02:03,995 --> 00:02:09,635
The RGB coding associates to each
color pixel three scholar values,

34
00:02:09,635 --> 00:02:13,185
R, G, and B, which represent
the amount of red, green, and

35
00:02:13,185 --> 00:02:17,040
blue that we need to mix up
to obtain the desired color.

36
00:02:17,040 --> 00:02:20,470
It's a fascinating fact about the human
visual system that we can actually

37
00:02:20,470 --> 00:02:24,900
represent such a wide array of
colors using just three components.

38
00:02:24,900 --> 00:02:26,360
Now when we use a color model,

39
00:02:26,360 --> 00:02:30,580
it means that each pixel has
a multidimensional value.

40
00:02:30,580 --> 00:02:32,450
But since these values are independent,

41
00:02:32,450 --> 00:02:37,305
we can split the original color images
into three fundamental components

42
00:02:37,305 --> 00:02:41,260
that are associated to
the components of the vector space.

43
00:02:41,260 --> 00:02:45,090
So in the case of RGB, for instance,
we will have three independent components

44
00:02:45,090 --> 00:02:49,720
that encode the red part,
the green part, and the blue part.

45
00:02:49,720 --> 00:02:53,520
Each of these images is now a scalar
image and so in the following,

46
00:02:53,520 --> 00:02:56,440
we will be simply concentrating
these color images for processing.

47
00:02:57,920 --> 00:03:01,710
So in image processing,
we're moving from one to two dimensions.

48
00:03:01,710 --> 00:03:06,140
And we know quite a bit about one
dimensional signal processing already.

49
00:03:07,150 --> 00:03:10,490
When we move to two dimensions,
something still works.

50
00:03:10,490 --> 00:03:15,180
We can still use some concepts
that we used in one dimension.

51
00:03:15,180 --> 00:03:17,440
Something unfortunately breaks down.

52
00:03:17,440 --> 00:03:19,200
And new things appear.

53
00:03:19,200 --> 00:03:22,340
So let's look and
turn at these three scenarios.

54
00:03:23,650 --> 00:03:27,980
The things that work and work rather well,
are the concepts of linearity and

55
00:03:27,980 --> 00:03:29,640
convolution.

56
00:03:29,640 --> 00:03:34,605
The Fourier transform into dimensions is
a simple extension of the one-dimensional

57
00:03:34,605 --> 00:03:35,930
Fourier transform.

58
00:03:35,930 --> 00:03:40,430
And interpolation and sampling work
exactly the same way in two dimensions.

59
00:03:40,430 --> 00:03:44,850
What works less well in image processing
is that, for instance, fourier analysis,

60
00:03:44,850 --> 00:03:48,250
which algorithmically is just
an extension of the one d case,

61
00:03:48,250 --> 00:03:51,660
becomes much less relevant
in the case of images.

62
00:03:51,660 --> 00:03:56,200
Filter design is much harder as well,
and IIR filters are rare.

63
00:03:56,200 --> 00:03:58,920
And linear operators
are only mildly useful and

64
00:03:58,920 --> 00:04:03,300
the reason is images
are very diverse signals.

65
00:04:03,300 --> 00:04:07,810
Imagine a photograph of a landscape
where you have all sorts of objects and

66
00:04:07,810 --> 00:04:08,830
textures.

67
00:04:08,830 --> 00:04:10,400
Now, a linear operator, and

68
00:04:10,400 --> 00:04:14,260
in 2D case it would be a linear
space invariant operator, would

69
00:04:14,260 --> 00:04:18,620
apply the same kind of transformation
to all different parts of an image.

70
00:04:18,620 --> 00:04:20,810
Regardless of what they represent.

71
00:04:20,810 --> 00:04:24,390
And of course, it's kind of difficult
to imagine that the same filter,

72
00:04:24,390 --> 00:04:26,590
unless it's a very, very simple operation.

73
00:04:26,590 --> 00:04:32,000
Will yield the desired result when applied
to very heterogeneous parts of an image.

74
00:04:32,000 --> 00:04:35,520
New concepts that appear in image
processing is the fact that we can

75
00:04:35,520 --> 00:04:38,580
introduce a new class of manipulations
called affine transforms,

76
00:04:38,580 --> 00:04:42,490
this include rotation,
scaling, skewing of images.

77
00:04:42,490 --> 00:04:44,810
Things you do in Photoshop all the time.

78
00:04:44,810 --> 00:04:47,860
The fact that images
are finite support signals.

79
00:04:47,860 --> 00:04:50,840
By definition, you take an image
with a camera and the CCD,

80
00:04:50,840 --> 00:04:53,720
the sensor of the camera
has a finite surface.

81
00:04:53,720 --> 00:04:58,200
So they're intrinsically finite
support and because of that

82
00:04:58,200 --> 00:05:02,130
an image signal is available in its
entirety from the beginning of processing.

83
00:05:02,130 --> 00:05:05,890
So while in 1D, you can imagine a system
that works online with the samples coming,

84
00:05:05,890 --> 00:05:08,280
and you never know when
the samples are going to stop.

85
00:05:08,280 --> 00:05:11,640
When it comes to images, you sort of
assume that you have the whole image

86
00:05:11,640 --> 00:05:14,020
already in memory before
you start processing.

87
00:05:14,020 --> 00:05:17,330
So causality is less of
an issue in image processing.

88
00:05:17,330 --> 00:05:21,650
However, all this applies to images,
not to general 2D singles.

89
00:05:21,650 --> 00:05:25,047
Images are very specialized signal,
and they are designed for

90
00:05:25,047 --> 00:05:29,000
a very specific type of receiver,
the human visual system.

91
00:05:29,000 --> 00:05:32,010
So images are a very small
subset of 2D singles, and

92
00:05:32,010 --> 00:05:34,630
a subset that is imbued with semantics.

93
00:05:34,630 --> 00:05:39,550
Now semantics are very, very hard to
deal with in our linear space and

94
00:05:39,550 --> 00:05:40,200
variant paradigm.

95
00:05:41,900 --> 00:05:46,050
Let's now look in more detail at the
extension of our digital signal processing

96
00:05:46,050 --> 00:05:47,446
paradigm to two dimensions.

97
00:05:47,446 --> 00:05:52,401
So a discrete-space signal is
the signal that we indicate with

98
00:05:52,401 --> 00:05:55,117
this notation, x[n1, n2].

99
00:05:55,117 --> 00:05:59,550
n1 and n2 are two discreet valued indices.

100
00:05:59,550 --> 00:06:01,460
The signal could be complex valued, but

101
00:06:01,460 --> 00:06:05,760
of course in the case of images, skylar
images, the values are going to be real.

102
00:06:06,760 --> 00:06:10,060
So, how do we represent this 2D signal?

103
00:06:10,060 --> 00:06:14,250
Well, from a standard mathematical point
of view we could represent it with

104
00:06:14,250 --> 00:06:19,630
a Cartesian plot where we have one
axis that indicates the first index.

105
00:06:19,630 --> 00:06:23,168
The second axis indicates
the second index, and

106
00:06:23,168 --> 00:06:27,663
the value of the signal is
represented as a third coordinate.

107
00:06:27,663 --> 00:06:32,480
So we have a 3D plot where
the scalar values form a 3D surface.

108
00:06:32,480 --> 00:06:36,380
Sometimes, especially in conjunction
with the description of filters,

109
00:06:36,380 --> 00:06:40,700
we are interested in what we call
the support representation of a 2D signal.

110
00:06:40,700 --> 00:06:44,340
In this representation we take a bird's
eye view of the signal and we only

111
00:06:44,340 --> 00:06:50,200
represent the known 0 values of the signal
as dots in the two dimensional plane.

112
00:06:50,200 --> 00:06:52,130
Since the height of the pixels,

113
00:06:52,130 --> 00:06:56,800
namely their scalar value, cannot be
inferred just by the dot representation,

114
00:06:56,800 --> 00:07:01,520
we often write the value of the signal of
the particular location next to the dot.

115
00:07:01,520 --> 00:07:05,780
So this plot for instance, represents
the two dimensional delta signal,

116
00:07:05,780 --> 00:07:09,480
which is a signal which is 0 everywhere,
except in the origin.

117
00:07:09,480 --> 00:07:10,310
Where it is 1.

118
00:07:10,310 --> 00:07:15,680
And so you have just 1 red dot
here at the origin with value 1.

119
00:07:15,680 --> 00:07:18,060
Of course the most common
representation for

120
00:07:18,060 --> 00:07:22,010
a 2D signal which is also an image
is an image representation.

121
00:07:22,010 --> 00:07:25,850
In this case we exploit
the dynamic range of the medium.

122
00:07:25,850 --> 00:07:28,230
In this case we have the computer monitor.

123
00:07:28,230 --> 00:07:33,570
And we know that each pixel can be driven
to represent a different shade of gray.

124
00:07:33,570 --> 00:07:38,900
And since the pixel values are packed
very closely together in space, here for

125
00:07:38,900 --> 00:07:45,250
instance we have 512 by 512 pixel values,
the density will be high,

126
00:07:45,250 --> 00:07:50,530
and the eye will create
the illusion of a continuous image.

127
00:07:51,730 --> 00:07:54,480
So one question that could come
up naturally at this point is,

128
00:07:54,480 --> 00:07:58,220
why do we go through the trouble
of defining a whole new

129
00:07:58,220 --> 00:08:00,970
two dimensional single
processing paradigm.

130
00:08:00,970 --> 00:08:04,770
Can we just convert images
into one of these signals and

131
00:08:04,770 --> 00:08:08,130
use the standard things that we've used so
far?

132
00:08:08,130 --> 00:08:10,480
And of course sometimes
that's exactly what we do.

133
00:08:10,480 --> 00:08:15,230
If you think of a printer that prints one
line at a time as the paper rolls out or

134
00:08:15,230 --> 00:08:18,210
a fax machine that's exactly what happens.

135
00:08:18,210 --> 00:08:23,100
However, if we do that we miss out on the
spatial correlation between pixels, and

136
00:08:23,100 --> 00:08:27,220
therefore the properties of an image
will be more difficult to understand.

137
00:08:27,220 --> 00:08:28,510
Let's look at an example.

138
00:08:28,510 --> 00:08:32,280
Here we have a 41 x 41 pixel image.

139
00:08:32,280 --> 00:08:35,740
And the content of this image
is simply a straight line.

140
00:08:35,740 --> 00:08:38,400
We will see that the angle of
the straight line will change later.

141
00:08:38,400 --> 00:08:42,010
What you have in the bottom panel
is what we call a raster scan

142
00:08:42,010 --> 00:08:43,330
representation of the image.

143
00:08:43,330 --> 00:08:47,880
In other words, we go through
the lines of the image one by one and

144
00:08:48,890 --> 00:08:52,490
we plot the correspondence
pixels on this axis.

145
00:08:52,490 --> 00:08:58,260
Now, for a horizontal line that coincides
with the N1 axis the resulting unrolled

146
00:08:58,260 --> 00:09:03,170
representation is just a series of 0
pixels except one with scanned this line

147
00:09:03,170 --> 00:09:09,420
which point we will have 41 pixels equal
to 1 and then we will go back to 0.

148
00:09:09,420 --> 00:09:11,630
So this is rather simple to understand,
but

149
00:09:11,630 --> 00:09:15,740
if we change the angle of the line
we see that the representation

150
00:09:15,740 --> 00:09:20,930
in a row fashion changes in ways
that are not very intuitive.

151
00:09:20,930 --> 00:09:27,410
When the angle is small we have clusters
of pixels interspersed with zeros.

152
00:09:27,410 --> 00:09:31,360
As the angle increases,
the spacing of the clusters changes and

153
00:09:31,360 --> 00:09:33,540
also number of pixels per cluster.

154
00:09:33,540 --> 00:09:38,280
It's very hard to understand to
visual characteristic of the line

155
00:09:38,280 --> 00:09:41,920
from the position of the clusters and
the number of pixels.

156
00:09:41,920 --> 00:09:47,770
After we pass the 45 degree angle we will
have collections of single pixel clusters,

157
00:09:47,770 --> 00:09:52,850
and the spacing of these clusters
will change in even more subtle ways,

158
00:09:52,850 --> 00:09:55,208
according to the angle of the line.

159
00:09:55,208 --> 00:09:59,950
Finally when each line that is
coinciding with the N2 axis we will have

160
00:09:59,950 --> 00:10:04,550
single pixels,
they are separated by 40 zeros.

161
00:10:04,550 --> 00:10:07,620
Because as we scan the image we
will hit a no zero pixel and

162
00:10:07,620 --> 00:10:10,740
then we will have to go 40 pixel
before we hit another one.

163
00:10:11,780 --> 00:10:16,280
This simple example should convince
you that a full 2D representation is

164
00:10:16,280 --> 00:10:20,720
necessary to best describe and
interpret an image signal.

165
00:10:22,060 --> 00:10:26,640
Just like we did for the one d k's,
here are some basic signals.

166
00:10:26,640 --> 00:10:29,900
The first one we have already
seen in passing is the delta,

167
00:10:29,900 --> 00:10:35,110
the impulse which is 0
everywhere except in the origin

168
00:10:35,110 --> 00:10:38,870
where it is equal to 1 and
the support representation is LAC cell.

169
00:10:40,030 --> 00:10:43,360
The two-dimensional erect signal
is defined by two parameters.

170
00:10:43,360 --> 00:10:49,040
Which we may call the width and the height
of the rect, and it is zero everywhere

171
00:10:49,040 --> 00:10:54,575
except in a rectangular region which is
defined by those values of the n1 index,

172
00:10:54,575 --> 00:10:59,430
that are smaller than
capital N1 in magnitude, and

173
00:10:59,430 --> 00:11:04,990
those values of the n2 index that
are smaller in magnitude than the N2.

174
00:11:04,990 --> 00:11:06,530
And of course it looks like this.

175
00:11:08,010 --> 00:11:11,740
One fundamental property for two
dimensional signals that has no equivalent

176
00:11:11,740 --> 00:11:15,050
in one dimension is that of separability.

177
00:11:15,050 --> 00:11:20,574
Now separability simply means that we
can write a two-dimensional signal

178
00:11:20,574 --> 00:11:26,375
as the product of two independent 1d
signals, defined in this as n1 and n2.

179
00:11:27,907 --> 00:11:31,858
So the delta signal, for instance,
is fully separable because

180
00:11:31,858 --> 00:11:36,190
this just the product of two delta
functions applied to both indexes.

181
00:11:37,270 --> 00:11:41,920
And similarly, the rectangular
function is again the product of two

182
00:11:41,920 --> 00:11:45,811
one dimensional rect functions,
defined over n1 and n2.

183
00:11:47,205 --> 00:11:50,330
Separability is a fragile property
in the sense that we can have simple

184
00:11:50,330 --> 00:11:51,250
transformations, or

185
00:11:51,250 --> 00:11:55,800
linear combinations of separable objects,
which are no longer separable.

186
00:11:55,800 --> 00:12:00,210
In this picture here you have a square for
instance which is simply a rect function

187
00:12:00,210 --> 00:12:05,370
with equal width and height but
which is rotated by 45 degrees.

188
00:12:05,370 --> 00:12:07,470
This signal is now separable.

189
00:12:07,470 --> 00:12:11,400
It is expressed as one when the sum

190
00:12:11,400 --> 00:12:15,780
of the indices in magnitude is smaller
than capital N and zero otherwise.

191
00:12:15,780 --> 00:12:18,800
But there is no way that
we can express this signal

192
00:12:18,800 --> 00:12:21,147
as the product of two
elementary 1d signals.

193
00:12:22,280 --> 00:12:29,144
Similarly the difference

194
00:12:29,144 --> 00:12:38,060
of two rect signals [CROSSTALK].

195
00:12:38,060 --> 00:12:42,388
[CROSSTALK] which is really a simple,

196
00:12:42,388 --> 00:12:49,514
straight forward extension
of the one-dimensional case.

197
00:12:49,514 --> 00:12:55,303
The convolution

198
00:12:55,303 --> 00:13:01,091
of two sequences

199
00:13:01,091 --> 00:13:07,706
in this great space

200
00:13:07,706 --> 00:13:14,734
is simply the sum for

201
00:13:14,734 --> 00:13:20,108
the first index

202
00:13:20,108 --> 00:13:27,550
to go using [CROSSTALK]

203
00:13:27,550 --> 00:13:31,271
Whereas as

204
00:13:31,271 --> 00:13:35,405
a separable

205
00:13:35,405 --> 00:13:41,608
convolution will

206
00:13:41,608 --> 00:13:47,396
require M1 plus

207
00:13:47,396 --> 00:13:52,771
M2 operations

208
00:13:52,771 --> 00:13:58,973
per output sample

209
00:13:58,973 --> 00:14:05,588
which is generally

210
00:14:05,588 --> 00:14:10,136
much smaller

211
00:14:10,136 --> 00:14:15,511
than the number

212
00:14:15,511 --> 00:14:20,473
of operations

213
00:14:20,473 --> 00:14:24,608
required in

214
00:14:24,608 --> 00:14:31,660
the previous case.

215
00:14:31,660 --> 00:14:35,143
Which is at least one order of
magnitude less than the previous case.

216
00:14:43,086 --> 00:14:44,351
Which is generally much,

217
00:14:44,351 --> 00:14:47,921
much smaller than the number of
operation required in the previous case.

