1
00:00:00,260 --> 00:00:05,600
So this chapter deals with probably
the core problem that most

2
00:00:05,600 --> 00:00:10,490
people face when they're wrestling with
how to represent high dimensional data.

3
00:00:10,490 --> 00:00:16,160
And they have to decide what
mapping they would like to make.

4
00:00:16,160 --> 00:00:21,300
And finding the appropriate mapping
is probably the hardest part of

5
00:00:21,300 --> 00:00:24,540
doing any visualization well.

6
00:00:24,540 --> 00:00:29,270
So one of the things that I think
helps people make good decisions is to

7
00:00:29,270 --> 00:00:36,190
increase your literacy in the full
scale of possible dimensional mappings.

8
00:00:36,190 --> 00:00:39,190
So this chapter is really going to be

9
00:00:39,190 --> 00:00:43,090
high speed walk through
a number of visual mappings.

10
00:00:43,090 --> 00:00:46,180
Focusing on a few that I think
are particularly effective.

11
00:00:47,690 --> 00:00:52,870
So the first is this
idea of small multiples.

12
00:00:52,870 --> 00:00:57,400
Small multiples is a name that Tufte gives
to a technique that's been around for

13
00:00:57,400 --> 00:00:59,040
hundreds of years.

14
00:00:59,040 --> 00:01:05,600
And he describes them as postage
size postage-stamp size graphics,

15
00:01:05,600 --> 00:01:08,840
indexed by some kind of category or label.

16
00:01:08,840 --> 00:01:13,220
And they're sequenced over time,
like the frames of a movie.

17
00:01:13,220 --> 00:01:18,230
And they're ordered by some sort of
quantitative variable that's not part of

18
00:01:18,230 --> 00:01:19,850
the image itself.

19
00:01:19,850 --> 00:01:23,900
And what's key to small
multiples is this idea that we

20
00:01:23,900 --> 00:01:30,170
have a repetition of form
with a uniform presentation.

21
00:01:30,170 --> 00:01:34,760
And this allows us to focus on
the differences in the images.

22
00:01:34,760 --> 00:01:38,620
So if you look here at this very,
very wonderful early

23
00:01:38,620 --> 00:01:43,064
example of small multiple which
comes from Galileo's notebooks and

24
00:01:43,064 --> 00:01:48,330
I'm taking this particular example
from Tufte visually and he describes.

25
00:01:49,670 --> 00:01:52,310
How Galileo shows Jupiter and

26
00:01:52,310 --> 00:01:57,290
the moons of Jupiter that
are visible on consecutive nights.

27
00:01:57,290 --> 00:02:01,640
And what you see here is the organizing
principle is a vertical axis,

28
00:02:01,640 --> 00:02:03,630
that here represents time.

29
00:02:03,630 --> 00:02:08,190
But what you're able to see is that
consecutively, on different nights,

30
00:02:08,190 --> 00:02:11,980
moons appear in different
locations relative to the planet.

31
00:02:11,980 --> 00:02:15,950
And this presentation allows
you to focus immediately on

32
00:02:15,950 --> 00:02:20,010
the differences between each
of the individual knights.

33
00:02:21,400 --> 00:02:24,370
And you'll see this is used again and
again in Science.

34
00:02:25,940 --> 00:02:32,320
Here is an example showing measures
of atmospheric contaminants.

35
00:02:32,320 --> 00:02:34,310
Within the Los Angeles area.

36
00:02:34,310 --> 00:02:39,200
And what you see here is that actually
the small multiple is organized into

37
00:02:39,200 --> 00:02:40,760
three rows.

38
00:02:40,760 --> 00:02:45,650
So we've taken a nominal variable
which includes the pollutant type and

39
00:02:45,650 --> 00:02:51,550
allowed that to represent each of
the different that's shown on each row.

40
00:02:51,550 --> 00:02:54,780
And then time on the x axis, and so

41
00:02:54,780 --> 00:03:00,420
what you can see is something like
nitrogen oxide peaks very late,

42
00:03:00,420 --> 00:03:05,480
while something like hydro
carbons are peaking, are not

43
00:03:05,480 --> 00:03:10,450
active between midnight to three, but very
active throughout the rest of the day.

44
00:03:10,450 --> 00:03:16,370
So that the beauty of small multiples is
that this unique form of presentation

45
00:03:16,370 --> 00:03:22,290
allows you to quickly jump with your
eye between different iterations and

46
00:03:22,290 --> 00:03:26,300
you focus on the differences
between each of the different,

47
00:03:27,830 --> 00:03:29,720
each of the different measurements.

48
00:03:29,720 --> 00:03:34,876
So you can look at this example of
trilogies, movie trilogies represented as,

49
00:03:34,876 --> 00:03:37,180
excuse me, small multiples.

50
00:03:37,180 --> 00:03:40,410
And what you can see
here is that each set of

51
00:03:40,410 --> 00:03:45,040
trilogies is shown next to one
another with a little bar graph.

52
00:03:45,040 --> 00:03:50,180
That represents the person's
opinion of the different movies.

53
00:03:50,180 --> 00:03:55,050
And so you can see that the person
found Mad Max two to be

54
00:03:55,050 --> 00:03:57,430
even better than Mad Max one.

55
00:03:57,430 --> 00:04:02,020
Whereas if you look across Star Wars
they all seem to be fantastic.

56
00:04:02,020 --> 00:04:05,090
Whereas Jaws one was incredible and
two and three were eh,.

57
00:04:06,900 --> 00:04:11,910
Small multiples gives you a wonderful
way to compare to present this data.

58
00:04:11,910 --> 00:04:17,240
Now this data could be easily shown
in a histogram, or a line graph, but

59
00:04:17,240 --> 00:04:21,380
the small multiple representation
allows you to cluster each of

60
00:04:21,380 --> 00:04:26,100
the different variants together, and
allows you to focus on their differences.

61
00:04:28,020 --> 00:04:30,960
Here's another wonderful
example of a small multiple.

62
00:04:30,960 --> 00:04:34,900
It's lipsticks shown by an artist,
Stacy Green.

63
00:04:34,900 --> 00:04:39,350
And what's wonderful about this set
of small multiples, is saying well,

64
00:04:39,350 --> 00:04:40,880
where is the data here?

65
00:04:40,880 --> 00:04:45,780
Well, each part of these
lipsticks is a lipstick taken from

66
00:04:45,780 --> 00:04:49,260
a different woman that,
that the artist knows.

67
00:04:49,260 --> 00:04:55,630
And what you actually see is a portrait
of the woman applying the lipstick.

68
00:04:55,630 --> 00:05:00,560
And so what you get to see is the way that
they move their hand through space as

69
00:05:00,560 --> 00:05:02,230
they apply the lipstick.

70
00:05:02,230 --> 00:05:05,440
And you see that the woman in the middle.

71
00:05:05,440 --> 00:05:09,460
Really applies this twisting motion
as she applies the lipstick.

72
00:05:09,460 --> 00:05:17,360
Whereas the woman on the top middle pushes
much harder on the bottom than top as she

73
00:05:17,360 --> 00:05:23,610
moves through her So what's beautiful
about this small multiple is that you see.

74
00:05:23,610 --> 00:05:28,170
Real comparisons between each of these
unique portraits of a person, but

75
00:05:28,170 --> 00:05:32,780
the actual data is hiding
behind the image itself.

76
00:05:34,800 --> 00:05:37,299
Here's another wonderful
example of a small multiple.

77
00:05:38,700 --> 00:05:44,690
Artists has taken the covers
of Playboy magazine

78
00:05:44,690 --> 00:05:49,360
from the 60s, 70s, 80s and
90s from left to right.

79
00:05:49,360 --> 00:05:55,460
And through a technique called image
averaging has averaged all the different

80
00:05:55,460 --> 00:06:01,080
images through the centerfold and
what you can see is a true trend.

81
00:06:01,080 --> 00:06:05,140
In the male historical gaze,

82
00:06:05,140 --> 00:06:10,460
as a preference towards a going from more

83
00:06:10,460 --> 00:06:16,978
of a dark haired woman,
to a blonde haired woman.

84
00:06:16,978 --> 00:06:19,590
From more of a robust woman
to a more skinny woman.

85
00:06:22,200 --> 00:06:28,140
If we turn to a slightly more empirical
data set, let's take a look at various

86
00:06:28,140 --> 00:06:36,520
ways one might compare Fisher's Historical
data on different types of Irises.

87
00:06:36,520 --> 00:06:41,070
So perhaps I think easiest to see is this.

88
00:06:41,070 --> 00:06:43,310
Parallel plot.

89
00:06:43,310 --> 00:06:49,830
And what you see here is that each
of the different dimensions a color

90
00:06:49,830 --> 00:06:55,810
is used to represent each of the different
nominal iris categories, or types.

91
00:06:55,810 --> 00:06:58,640
And then an axis, a vertical axis,

92
00:06:58,640 --> 00:07:03,880
is used to represent each of
the different features of each category.

93
00:07:03,880 --> 00:07:06,320
And so you can see simple lengths.

94
00:07:06,320 --> 00:07:10,850
Petal length are all
represented on this scale.

95
00:07:10,850 --> 00:07:15,220
And then,
each iris is represented as a single line

96
00:07:15,220 --> 00:07:19,010
drawn across from axis to axis.

97
00:07:19,010 --> 00:07:24,640
And you can see that iris setosa
clearly clusters in the bottom left and

98
00:07:24,640 --> 00:07:27,630
has, generally has a smaller petal length.

99
00:07:27,630 --> 00:07:30,880
Then then the other irises,
like the virginica,

100
00:07:30,880 --> 00:07:32,760
which has a much larger petal length.

101
00:07:34,010 --> 00:07:38,610
So one additional,
very powerful technique for

102
00:07:38,610 --> 00:07:42,980
representing data is something
that Tufte calls a spark line.

103
00:07:42,980 --> 00:07:48,290
And it's a very small,
simple word-sized graphic.

104
00:07:48,290 --> 00:07:52,630
And Tufte describes it as having
a typographic resolution.

105
00:07:52,630 --> 00:07:57,200
And what he means is these are graphics
that we can embed into the very text of

106
00:07:57,200 --> 00:07:59,080
our scientific documents.

107
00:07:59,080 --> 00:08:03,950
And so if you have a number here,
for example, glucose at a level 6.6.

108
00:08:03,950 --> 00:08:07,760
Well, in general for
many readers this number.

109
00:08:07,760 --> 00:08:11,390
Might have no context,
or scale, or history.

110
00:08:11,390 --> 00:08:18,060
So, we can change this inline in the
document by adding a very small line graph

111
00:08:18,060 --> 00:08:24,750
that would show the data values 6.6, and
comparing it to earlier measurements.

112
00:08:24,750 --> 00:08:27,810
So, you what can see now,
is that same number.

113
00:08:27,810 --> 00:08:33,072
Actually has a history and can show
you that for this particular person

114
00:08:33,072 --> 00:08:38,439
there was a peak at a, at some
point halfway through that history.

115
00:08:38,439 --> 00:08:41,167
And it's now more at normal levels.

116
00:08:41,167 --> 00:08:46,480
Another thing that one might be
able to do, is to tie that number.

117
00:08:46,480 --> 00:08:49,060
The number which we're
describing in the text

118
00:08:50,060 --> 00:08:54,460
to the actual line graph
using a color marker.

119
00:08:54,460 --> 00:09:00,130
So, you can see that semantically we've
created a relationship between 6.6 and

120
00:09:00,130 --> 00:09:02,689
the right-most number in the line graph.

121
00:09:04,690 --> 00:09:07,910
Another additional thing
that you can do for

122
00:09:07,910 --> 00:09:15,210
a spark line is to show something that
includes the normal distribution.

123
00:09:15,210 --> 00:09:18,590
And so by you using a slightly
different background color.

124
00:09:18,590 --> 00:09:22,615
What you're now additionally
able to see is not just that

125
00:09:22,615 --> 00:09:28,150
6.6 is the latest value, but
it's within normal range.

126
00:09:28,150 --> 00:09:32,940
And that, that peak value was only above
normal range for a short period of time.

127
00:09:34,220 --> 00:09:41,110
So, the graphic itself can be
moved on either side of the.

128
00:09:41,110 --> 00:09:44,270
The actual word it's describing,
or the number,

129
00:09:44,270 --> 00:09:48,610
and that they can be embedded right
into the text of a scientific document.

130
00:09:49,680 --> 00:09:55,570
Separately, very often we find that
these types of graphs are used

131
00:09:55,570 --> 00:10:01,130
in things like,
financial pages to describe a history

