So now that we've talked about data representation and how to create it, your first attempt at showing data, and how to get better at it, the, the more your, your science progresses, the more likely you're bound to find bottlenecks. On how you create an application and how that application handles your data. So let's talk a little bit about those bottlenecks. It turns out in data visualization, a very usual bottleneck is actually just a data injunction. If you're the one creating the full simulation, you might have control over your output and the format of your output. But if you don't, then not only might you have problems with the data size, but even trying to understand and massaging your data format to actually be read by your visualization tool. Or especially if it's an off-the-shelf component, might actually require quite a bit of time. And that's something that we encounter even in our regular classes that we do in visualization, is how much time people spend ingesting and manipulating and massaging the data. There has been work being done on that. There's this work that came out of, out of, out of Stanford which is called Data Wrangler, which is an attempt to actually do smarter text manipulation to be able to convert mangled, confusing data files into cleaner tables. So for instance, let's look at this table data from the FBI which shows crime per state. It has a very nice table and if you press Download, it actually gives you this data file which actually has state and then, and then the crime rates for the last five years. Now, while this seems clean, it's actually a little more complicated to see okay, how am I going to, if I want to ingest a table, how do I clean this up to see how much stuff is here and how many, how many, many years, it has and how clean and how to ingest all this information so I can read it easily into a system like Mondrian or Excel even. So Data Wrangler, what it does is it loads in the data the way it is, and the moment you start interacting with it, it starts giving you suggestions on what you can do. So if for, imagine you have this data file. It automatically notices that it has years and a column for, for, for different data, but it also has this header that says reported crime in each state. So now, if you click in, right after the word in, it gives you many suggestions including this, the ability to split anything that has the word crime in by the data that's, the word after in, and everything that, that's before. So if you do that, and it gives you the, the results right here in yellow as well, you actually get this table, that actually has the state separated from, from the, from the words that were not important, which is reported crime in. Now see what we can do is we can click on the empty row and say okay, go ahead and delete all the empty rows, and it gives us already a cleaner table. Then if we click on the word Alabama, then it notices that everything underneath is actually empty. So one of the suggestions it actually gives you, it says do you want to copy down? And you say sure, you copy that, and it gives you a table that is now full. Now, one of the first problems that you have is, okay, now you still have this one row that says crime reported in, without actually any data. Well, if you click back on crime reported in, you can say you know what? Let me go ahead and delete all those rows. So now you have a, a data table that has year, state and crime. Now the problem is that you don't want the year in a cl, in a straight sequence. So what you want to do is you want to flip this, you want to pivot it. So what you do is you select a range of years, and automatically it tells you, do you want to unfold this data to actually put the years as, as, as your different variables? And in fact, that's exactly what you want to do, so you just click on the years and then you tell it to unfold and then finally you get a clean data set. Now this actually creates a set of scripts that you can actually then go through even a much larger data set and, for the most part, this will actually be quite efficient and class it quite fast. Of course, you could actually just learn description languages yourself like sed and AWK and figure out how to do this. But, for the most part, a tool like this is already quite an advantage for people who don't want to spend time learning more complicated tools like sed and AWK. So that was the issue with data ingestion. What about the rest of it? What about computing power? When you have, when you try to do processing algorithms, most of the times if you have a large data set, you're just going to get clogged up. And then eventually, once you create your geometry, if you have millions or billions of points, how are you going to represent those on the screen? There's no way you can get your graphic's card to be able to swallow that many data points. So clearly you have a problem here where you have a throughput issue, and that has a lot to do both with your graphics card as well as your computing performance. And there is an answer for that as well. And the, and the answer is to actually either take it in parallel, which is more of a brute force approach, basically bring many machines to do the job, although actually programming it parallel is quite complicated. Or you can figure out a way to structure your data so that you can read at different levels in the hierarchy and create faster representations of your data and faster faster computations of your, of your, of your data at a smaller level. Either a subset or a, or a derived lower detail. So for instance, let's look at this mesh. And this is the same mesh that we talked about for MCell before. It's a, it's, it's a, it's a, a cell membrane. So here we have this same cell at three different levels of detail. Now, it takes quite a while to compute different levels of details, and of course, it also depends on how good your algorithm is. And then when you, when you actually create different levels of detail, you need to find its resolution but you're probably going to store it as well as your lower levels of resolution, so you're going to need larger storage space as well. But these are the results that you can get out of, out of, out of creating that. And what you can do, for instance, if you're just trying to visualize something, well, you can make it so that when you're seeing something complicated from far away, you use the lower-level resolution which you can hardly tell there's any problems in there because it's so far. It's using so many pixels anyway. And then you use the higher resolution for when you're looking at something closer. Notice that you can also do it to actually create interactive applications. A classic paper that was presented at SIGGRAPH in 2001, was this cues plot paper, which has scans of Michelangelo statues. And what it actually did is actually created these levels of hierarchy, and it created a much lower resolution version of each of these plots. So if you were interacting with this plot, this is what you would see. You would actually see a lower resolution version, which has been drawn with these big fat points. And in, in OpenGL, the graphics are, points are actually square. And then you could spin it around, and we're talking billions of, of data points. And render to visualize them in over 60 frames per second. Then the moment you started slowing down and the, and the computer was telling you that you had more time to render it will just go one lever deeper in your resolution. If you still didn't move your mouse, then eventually it just started going down and down until it finally reached a higher level of resolution. The same thing can be done for data. Not only can it be done for data but it can actually be done for analyzing and exploring data. Here at CalTech we did a system that actually did the same thing but actually encoded the internal structure in one of these hierarchical trees as well. So you can actually do a CSG probe where we actually moved interactively a probe and are able to cut or even take a sphere through an object and see the inside of an object at any frame rate. It would actually be able to, it was actually able to handle over a million, a million vertices with about 6 million tetrahedra at any frame rate, and eventually when you let the mouse go, it would go to the highest resolution. Of course, you have to program the ability to actually encode these hierarchical models as well as store the data at different levels of resolution. Of course, if you're smart enough, you can use higher algorithms to actually do some data encoding that actually have some lossy encoding to actually make the data sets maybe smaller. However, the alternative which seems simpler, in many cases, is just to do the same thing just in parallel. Here are some results that we did from a parallel volume rendering cluster, in which we visualized a large volume of a really tailored simulation done here at CalTech as well as in Lawrence Livermore National Lab between. And we actually were able not only to visualize this large data set in, in, in real time at interactive speeds, but actually visualize it in a high-resolution screens. Now, obviously the problem is that first of all you need to have access to a parallel system, and then you have to have the software that is actually able to handle this complex environment. And of course, once you develop this thing, it's less likely to be portable. Now let me put a side note right here to say that while parallel can be complicated, there is people working on it. For instance, if you notice, ParaView has the word para at the beginning and the reason is because it was meant to be a parallel system. When you double-click in ParaView, it actually does everything on your workstation, but if you want to, you can use batch tools to actually just deploy the interface in your, your workstation and actually deploy both computational nodes, as well as rendering nodes in the back end. And you can have multiple computational logs, and multiple rendering nodes that will do the best at splitting the data, doing the computation in pieces, then rendering the pieces, and then finally, rendering the pieces, sending the, rendering the cells to your front end, where it will get composed into a single image. The other package we didn't talk about much is called VisIt. That while its interface is not as intuitive as ParaView, is actually even better situated and has been tweaked even better to actually perform in parallel. In fact, ParaView, while it can work on any system, VisIt actually comes with profiles to work on the large clusters and large supercomputers at the national labs.