1
00:00:02,540 --> 00:00:05,070
So now that we've talked
about data representation and

2
00:00:05,070 --> 00:00:10,070
how to create it,
your first attempt at showing data,

3
00:00:10,070 --> 00:00:14,690
and how to get better at it,
the, the more your,

4
00:00:14,690 --> 00:00:19,630
your science progresses, the more likely
you're bound to find bottlenecks.

5
00:00:19,630 --> 00:00:22,790
On how you create an application and
how that application handles your data.

6
00:00:23,810 --> 00:00:25,820
So let's talk a little bit
about those bottlenecks.

7
00:00:28,180 --> 00:00:30,310
It turns out in data visualization,

8
00:00:30,310 --> 00:00:34,230
a very usual bottleneck is
actually just a data injunction.

9
00:00:34,230 --> 00:00:37,060
If you're the one creating
the full simulation,

10
00:00:37,060 --> 00:00:40,540
you might have control over your
output and the format of your output.

11
00:00:41,650 --> 00:00:45,450
But if you don't, then not only might you
have problems with the data size, but

12
00:00:45,450 --> 00:00:47,550
even trying to understand and

13
00:00:47,550 --> 00:00:51,442
massaging your data format to actually
be read by your visualization tool.

14
00:00:51,442 --> 00:00:55,031
Or especially if it's
an off-the-shelf component,

15
00:00:55,031 --> 00:00:58,130
might actually require
quite a bit of time.

16
00:00:58,130 --> 00:01:01,400
And that's something that we encounter
even in our regular classes that we do

17
00:01:01,400 --> 00:01:06,030
in visualization, is how much
time people spend ingesting and

18
00:01:06,030 --> 00:01:07,610
manipulating and massaging the data.

19
00:01:10,850 --> 00:01:12,350
There has been work being done on that.

20
00:01:13,750 --> 00:01:16,578
There's this work that came out of,
out of,

21
00:01:16,578 --> 00:01:20,060
out of Stanford which is
called Data Wrangler,

22
00:01:20,060 --> 00:01:24,879
which is an attempt to actually do smarter
text manipulation to be able to convert

23
00:01:26,340 --> 00:01:30,530
mangled, confusing data
files into cleaner tables.

24
00:01:32,020 --> 00:01:32,680
So for instance,

25
00:01:32,680 --> 00:01:36,750
let's look at this table data from
the FBI which shows crime per state.

26
00:01:38,390 --> 00:01:41,180
It has a very nice table and
if you press Download,

27
00:01:41,180 --> 00:01:44,910
it actually gives you this data file
which actually has state and then, and

28
00:01:44,910 --> 00:01:47,400
then the crime rates for
the last five years.

29
00:01:48,590 --> 00:01:53,470
Now, while this seems clean, it's actually
a little more complicated to see okay,

30
00:01:53,470 --> 00:01:56,200
how am I going to,
if I want to ingest a table,

31
00:01:56,200 --> 00:01:59,960
how do I clean this up to see how much
stuff is here and how many, how many,

32
00:01:59,960 --> 00:02:04,260
many years, it has and how clean and
how to ingest all this information so

33
00:02:04,260 --> 00:02:07,403
I can read it easily into a system
like Mondrian or Excel even.

34
00:02:09,110 --> 00:02:13,750
So Data Wrangler, what it does is it loads
in the data the way it is, and the moment

35
00:02:13,750 --> 00:02:18,380
you start interacting with it, it starts
giving you suggestions on what you can do.

36
00:02:18,380 --> 00:02:20,520
So if for,
imagine you have this data file.

37
00:02:20,520 --> 00:02:25,120
It automatically notices that it has years
and a column for, for, for different data,

38
00:02:25,120 --> 00:02:28,280
but it also has this header that
says reported crime in each state.

39
00:02:28,280 --> 00:02:33,050
So now, if you click in,
right after the word in, it gives you many

40
00:02:33,050 --> 00:02:38,140
suggestions including this, the ability
to split anything that has the word crime

41
00:02:38,140 --> 00:02:42,850
in by the data that's, the word after in,
and everything that, that's before.

42
00:02:44,520 --> 00:02:49,855
So if you do that, and it gives you the,
the results right here in yellow as well,

43
00:02:49,855 --> 00:02:54,798
you actually get this table, that
actually has the state separated from,

44
00:02:54,798 --> 00:03:00,082
from the, from the words that were not
important, which is reported crime in.

45
00:03:00,082 --> 00:03:04,541
Now see what we can do is we can click on
the empty row and say okay, go ahead and

46
00:03:04,541 --> 00:03:08,545
delete all the empty rows, and
it gives us already a cleaner table.

47
00:03:08,545 --> 00:03:10,312
Then if we click on the word Alabama,

48
00:03:10,312 --> 00:03:13,980
then it notices that everything
underneath is actually empty.

49
00:03:13,980 --> 00:03:15,434
So one of the suggestions
it actually gives you,

50
00:03:15,434 --> 00:03:16,400
it says do you want to copy down?

51
00:03:16,400 --> 00:03:21,770
And you say sure, you copy that, and
it gives you a table that is now full.

52
00:03:21,770 --> 00:03:25,406
Now, one of the first problems that you
have is, okay, now you still have this one

53
00:03:25,406 --> 00:03:28,180
row that says crime reported in,
without actually any data.

54
00:03:28,180 --> 00:03:31,030
Well, if you click back on crime
reported in, you can say you know what?

55
00:03:31,030 --> 00:03:32,730
Let me go ahead and delete all those rows.

56
00:03:35,190 --> 00:03:40,090
So now you have a, a data table
that has year, state and crime.

57
00:03:40,090 --> 00:03:43,690
Now the problem is that you don't want
the year in a cl, in a straight sequence.

58
00:03:43,690 --> 00:03:47,530
So what you want to do is you want to
flip this, you want to pivot it.

59
00:03:47,530 --> 00:03:52,210
So what you do is you select a range of
years, and automatically it tells you,

60
00:03:52,210 --> 00:03:56,865
do you want to unfold this data
to actually put the years as, as,

61
00:03:56,865 --> 00:03:58,990
as your different variables?

62
00:03:58,990 --> 00:04:01,887
And in fact, that's exactly what you want
to do, so you just click on the years and

63
00:04:01,887 --> 00:04:05,690
then you tell it to unfold and
then finally you get a clean data set.

64
00:04:07,520 --> 00:04:11,730
Now this actually creates a set of
scripts that you can actually then

65
00:04:11,730 --> 00:04:15,010
go through even a much larger
data set and, for the most part,

66
00:04:15,010 --> 00:04:19,100
this will actually be quite efficient and
class it quite fast.

67
00:04:19,100 --> 00:04:22,190
Of course, you could actually just
learn description languages yourself

68
00:04:22,190 --> 00:04:25,130
like sed and AWK and
figure out how to do this.

69
00:04:25,130 --> 00:04:28,910
But, for the most part, a tool like
this is already quite an advantage for

70
00:04:28,910 --> 00:04:33,802
people who don't want to spend time
learning more complicated tools like sed

71
00:04:33,802 --> 00:04:37,090
and AWK.

72
00:04:37,090 --> 00:04:38,950
So that was the issue with data ingestion.

73
00:04:38,950 --> 00:04:40,670
What about the rest of it?

74
00:04:40,670 --> 00:04:41,830
What about computing power?

75
00:04:41,830 --> 00:04:44,510
When you have,
when you try to do processing algorithms,

76
00:04:44,510 --> 00:04:47,140
most of the times if you have a large data
set, you're just going to get clogged up.

77
00:04:49,070 --> 00:04:52,670
And then eventually, once you create
your geometry, if you have millions or

78
00:04:52,670 --> 00:04:55,750
billions of points, how are you going
to represent those on the screen?

79
00:04:55,750 --> 00:04:58,640
There's no way you can get your graphic's
card to be able to swallow that many

80
00:04:58,640 --> 00:04:59,180
data points.

81
00:05:00,280 --> 00:05:05,060
So clearly you have a problem here where
you have a throughput issue, and that has

82
00:05:05,060 --> 00:05:09,540
a lot to do both with your graphics card
as well as your computing performance.

83
00:05:10,710 --> 00:05:12,990
And there is an answer for that as well.

84
00:05:12,990 --> 00:05:16,430
And the, and the answer is to
actually either take it in parallel,

85
00:05:16,430 --> 00:05:21,620
which is more of a brute force approach,
basically bring many machines to

86
00:05:21,620 --> 00:05:26,010
do the job, although actually programming
it parallel is quite complicated.

87
00:05:26,010 --> 00:05:30,030
Or you can figure out a way to structure
your data so that you can read

88
00:05:30,030 --> 00:05:34,170
at different levels in the hierarchy and
create faster representations of your data

89
00:05:34,170 --> 00:05:40,280
and faster faster computations of your,
of your, of your data at a smaller level.

90
00:05:40,280 --> 00:05:43,760
Either a subset or a, or
a derived lower detail.

91
00:05:45,750 --> 00:05:48,500
So for instance, let's look at this mesh.

92
00:05:48,500 --> 00:05:51,920
And this is the same mesh that we
talked about for MCell before.

93
00:05:51,920 --> 00:05:56,050
It's a, it's, it's a,
it's a, a cell membrane.

94
00:05:56,050 --> 00:05:59,140
So here we have this same cell at
three different levels of detail.

95
00:06:00,310 --> 00:06:03,560
Now, it takes quite a while to compute
different levels of details, and

96
00:06:03,560 --> 00:06:06,327
of course, it also depends on
how good your algorithm is.

97
00:06:06,327 --> 00:06:10,970
And then when you, when you actually
create different levels of detail,

98
00:06:10,970 --> 00:06:12,530
you need to find its resolution but

99
00:06:12,530 --> 00:06:15,990
you're probably going to store it as well
as your lower levels of resolution, so

100
00:06:15,990 --> 00:06:20,080
you're going to need larger
storage space as well.

101
00:06:20,080 --> 00:06:24,000
But these are the results that you can
get out of, out of, out of creating that.

102
00:06:24,000 --> 00:06:27,450
And what you can do, for instance, if
you're just trying to visualize something,

103
00:06:27,450 --> 00:06:30,300
well, you can make it so that when
you're seeing something complicated from

104
00:06:30,300 --> 00:06:32,890
far away,
you use the lower-level resolution

105
00:06:32,890 --> 00:06:35,680
which you can hardly tell there's any
problems in there because it's so far.

106
00:06:35,680 --> 00:06:37,510
It's using so many pixels anyway.

107
00:06:37,510 --> 00:06:41,810
And then you use the higher resolution for
when you're looking at something closer.

108
00:06:41,810 --> 00:06:45,018
Notice that you can also do it to
actually create interactive applications.

109
00:06:45,018 --> 00:06:50,520
A classic paper that was
presented at SIGGRAPH in 2001,

110
00:06:50,520 --> 00:06:56,740
was this cues plot paper,
which has scans of Michelangelo statues.

111
00:06:58,480 --> 00:07:02,520
And what it actually did is actually
created these levels of hierarchy, and

112
00:07:02,520 --> 00:07:06,010
it created a much lower resolution
version of each of these plots.

113
00:07:06,010 --> 00:07:08,970
So if you were interacting with this plot,
this is what you would see.

114
00:07:08,970 --> 00:07:11,390
You would actually see
a lower resolution version,

115
00:07:11,390 --> 00:07:13,940
which has been drawn with
these big fat points.

116
00:07:13,940 --> 00:07:17,470
And in, in OpenGL, the graphics are,
points are actually square.

117
00:07:17,470 --> 00:07:21,450
And then you could spin it around, and
we're talking billions of, of data points.

118
00:07:21,450 --> 00:07:25,535
And render to visualize them
in over 60 frames per second.

119
00:07:25,535 --> 00:07:28,940
Then the moment you started
slowing down and the, and

120
00:07:28,940 --> 00:07:31,660
the computer was telling you that
you had more time to render it

121
00:07:31,660 --> 00:07:33,970
will just go one lever
deeper in your resolution.

122
00:07:33,970 --> 00:07:38,640
If you still didn't move your mouse, then
eventually it just started going down and

123
00:07:38,640 --> 00:07:41,709
down until it finally reached
a higher level of resolution.

124
00:07:42,780 --> 00:07:44,510
The same thing can be done for data.

125
00:07:44,510 --> 00:07:47,420
Not only can it be done for data but
it can actually be done for analyzing and

126
00:07:47,420 --> 00:07:48,060
exploring data.

127
00:07:49,578 --> 00:07:52,310
Here at CalTech we did a system that
actually did the same thing but

128
00:07:52,310 --> 00:07:56,930
actually encoded the internal structure in
one of these hierarchical trees as well.

129
00:07:56,930 --> 00:08:01,570
So you can actually do a CSG probe where
we actually moved interactively a probe

130
00:08:01,570 --> 00:08:05,280
and are able to cut or
even take a sphere through an object and

131
00:08:05,280 --> 00:08:08,970
see the inside of an object
at any frame rate.

132
00:08:08,970 --> 00:08:13,236
It would actually be able to, it was
actually able to handle over a million,

133
00:08:13,236 --> 00:08:17,436
a million vertices with about 6 million
tetrahedra at any frame rate, and

134
00:08:17,436 --> 00:08:21,788
eventually when you let the mouse go,
it would go to the highest resolution.

135
00:08:24,889 --> 00:08:28,939
Of course, you have to program the ability
to actually encode these hierarchical

136
00:08:28,939 --> 00:08:32,310
models as well as store the data
at different levels of resolution.

137
00:08:32,310 --> 00:08:33,340
Of course, if you're smart enough,

138
00:08:34,970 --> 00:08:39,920
you can use higher algorithms to actually
do some data encoding that actually

139
00:08:39,920 --> 00:08:43,929
have some lossy encoding to actually
make the data sets maybe smaller.

140
00:08:46,110 --> 00:08:50,480
However, the alternative which
seems simpler, in many cases,

141
00:08:50,480 --> 00:08:52,400
is just to do the same
thing just in parallel.

142
00:08:54,850 --> 00:09:00,359
Here are some results that we did from
a parallel volume rendering cluster,

143
00:09:00,359 --> 00:09:06,044
in which we visualized a large volume of
a really tailored simulation done here

144
00:09:06,044 --> 00:09:11,833
at CalTech as well as in
Lawrence Livermore National Lab between.

145
00:09:11,833 --> 00:09:14,508
And we actually were able not only to
visualize this large data set in, in,

146
00:09:14,508 --> 00:09:15,953
in real time at interactive speeds, but

147
00:09:15,953 --> 00:09:17,960
actually visualize it in
a high-resolution screens.

148
00:09:20,290 --> 00:09:23,099
Now, obviously the problem is that
first of all you need to have

149
00:09:23,099 --> 00:09:26,067
access to a parallel system, and
then you have to have the software

150
00:09:26,067 --> 00:09:28,720
that is actually able to handle
this complex environment.

151
00:09:30,140 --> 00:09:32,720
And of course, once you develop this
thing, it's less likely to be portable.

152
00:09:35,090 --> 00:09:38,450
Now let me put a side note right
here to say that while parallel

153
00:09:38,450 --> 00:09:42,770
can be complicated,
there is people working on it.

154
00:09:42,770 --> 00:09:47,150
For instance, if you notice, ParaView
has the word para at the beginning and

155
00:09:47,150 --> 00:09:49,975
the reason is because it was
meant to be a parallel system.

156
00:09:49,975 --> 00:09:53,470
When you double-click in ParaView,

157
00:09:53,470 --> 00:09:56,270
it actually does everything on your
workstation, but if you want to,

158
00:09:56,270 --> 00:10:00,800
you can use batch tools to actually
just deploy the interface in your,

159
00:10:00,800 --> 00:10:04,560
your workstation and
actually deploy both computational nodes,

160
00:10:04,560 --> 00:10:06,610
as well as rendering
nodes in the back end.

161
00:10:06,610 --> 00:10:09,480
And you can have multiple
computational logs, and

162
00:10:09,480 --> 00:10:13,150
multiple rendering nodes that will
do the best at splitting the data,

163
00:10:13,150 --> 00:10:17,020
doing the computation in pieces, then
rendering the pieces, and then finally,

164
00:10:17,020 --> 00:10:20,890
rendering the pieces, sending the,
rendering the cells to your front end,

165
00:10:20,890 --> 00:10:23,130
where it will get composed
into a single image.

166
00:10:24,700 --> 00:10:28,430
The other package we didn't talk
about much is called VisIt.

167
00:10:28,430 --> 00:10:32,670
That while its interface is not as
intuitive as ParaView, is actually even

168
00:10:32,670 --> 00:10:38,500
better situated and has been tweaked even
better to actually perform in parallel.

169
00:10:38,500 --> 00:10:44,038
In fact, ParaView, while it can work on
any system, VisIt actually comes with

170
00:10:44,038 --> 00:10:49,740
profiles to work on the large clusters and
large supercomputers at the national labs.

