Chapter 2 has very simple thesis. And it comes from Jock Mackinlay who is one of the I think the trailblazers in modern computational visualization. And he came up with this idea that he called Effectiveness. And he basically says, look, how can we tell if one visualization is better than another one? And he just came up with this heuristic, but I think it's very effective. And it relies on our understanding of human perception. So you can see he says a visualization, it, it's just more effective than another one, if the information it's conveyed is more readily perceived than the information in the other visualization. So in other words, in, if you want to learn how to make visualizations that can communicate powerfully, it's critical to understand how the human mind perceives shape, color, line, space, motion, and interaction. And these elements together, shape how we actually perceive what's on the display. So in this section, I'm going to quickly talk about some of the important qualities of human perception and see how they affect the way that we should structure data when we present it visually. So the first is just a quick walk through of Gestalt's ecology, which is basically a school of psychology that came out of Germany, and it really describes something very interesting, which is that the mind looks at objects in their entirety, before or in parallel with perception of individual parts. And this means that, you know, there's this very intriguing statement at the end here. The whole is other than the sum of its parts. So, let's talk about what this means. Look at this first concept here. The idea is called similarity. And what it means here, is that the mind takes shapes that have similar semantic relationship. Shape, and it brings them together. So what you see here is this eagle shape, but around the eagle are really effectively just a series of triangles. But our mind reads that those triangles create a kind of halo or circle behind the form of the eagle, if you look at the shapes, just by themselves. So don't think of them as a halo. What they are is just a series of triangles, but the mind reads them differently. So what's important to understand is how the mind reads these symbols together. I'll give you another example here. The principle is called Anomaly. An anomaly, very simply, tells the story that if there's continuity, and a break in the continuity, that that shape, that is no longer contiguous is going to stand out more. So what's important in human perception, is putting shapes together and identifying shapes that stand out from that togetherness. This has several interesting consequences. The first, is this principle called Continuity. And what this means, is that we the human mind will look at the cross, the swoop that goes across that H and read that as a continuous path. And while it's just a shape, it's, it's a negative space. What we really read that to mean, is that the leaf has blown through the H. So there's meaning in the con, the connection and continuation of shapes that are next to one another. Another important idea is something called Closure. And what this means, is that the mind makes shapes contiguous that are not formed entirely. And so if you look at the top of this panda bear shape, what you'll see is there's actually no top of the bear drawn, but your mind draws it in for you. So you can see how clearly the human mind is really interesting at forming patterns. And these patterns are not necessarily visible. So it's critical that we understand these patterns when we're putting together things that are more data oriented. So this other concept is called Proximity. And Proximity means that the mind believes that things that are closer together, have more meaning. In other words, the dots, the squares in this particular form, are less related than the forms in this particular shape. So when we see dots that are closer together, we interpret that to mean similar. So it's really important that we understand that, because you'll see that as you create visualization, it's very often easy to accidentally put things together that don't have the same semantic meaning. So we need to be very careful of this because when we see shapes, we infer meaning. Another example of proximity here, is the way that we just look at this distribution of shapes that are the human form. And because of the way that they're approximal to one another, we interpret this as movement and as a formation of people. But if you tried to describe this in computational terms, it's really just the same image, just in different orientations and locations relative to one another. But the human mind reads this as motion. A cloud. So let's talk about how the way that these features in the human perception affect the way that we should design visualizations. So I'll ask a question. Very simple, look at the following shapes and say which of the, which of them is bigger. So it's pretty clear that the there's a tall and short rectangle. Now here's this question. How much bigger is the tall rectangle? Well, the human mind can make this assessment relative much better, than it can this assessment. So interestingly, if we were to say, how much taller is the tall rectangle than the short rectangle, it's far easier for the human mind to see that actually it's three times as tall. Make the same determination of this circle. It's actually significantly harder for us to make that same determination. But interestingly, they are exactly the same heights. Let's talk about color for a moment. Let's ask the question, which of the following squares is brighter? Give everyone a minute, make a determination. Okay. So the answer is the one on the left is slightly brighter, but what you'll see is many people will actually not have been able to tell the difference. Let's look at another one. Look at this pair of images and say which is brighter. Actually, what you can see is the one on the right is brighter, if we describe them in RGB space. And this amount of difference is actually perceivable by a large number of people. So when we look at the way the human mind understands color, what we have to understand is that even though there are differences that we might be able to label in RGB space, the human mind is only capable of perceiving what we call the just noticeable difference. In other words, there're some steps in color space, that are not perceptible to the human mind. So what we have to realize, is that if we don't allow for enough space between individual color values in a scientific visualization, a human mind won't be able to perceive key differences in the data. This is a common, common, common problem you see in scientific visualization, that uses continuous color. It overlooks this very important fact of perception, which is that if you don't allow for humans to perceive the differences, your data will effectively become subsampled. Because people won't be able to see the difference between a value and another value that are close enough in color space, even though their differences are clear in numerical space. So let's think about, what is this mean, if we're going to use how the human mind perceives color, and bring this into a visualization. well, actually color and shape, et cetera. So, what we find important is this idea going back to McKinley. The one of the early visualization thinkers, is he put forth these heuristics. The first being something called the principal of consistency. And that means that the properties of the image should match the properties of the data. In other words if something's big, the data should make it look big. If it's small, the data should look, make it look small. And then the second idea is the importance of the principle of importance ordering. And this is something that we often overlook. The idea is that take the most important variables, and encode them in the most effective way. And our understanding of human perception, is such that we perceive differences of magnitude differently when they're encoded using different encoding methods. And so if you can see that position is much more clear of a difference to the human mind than color. So what this means is that when you are making a visualization, the shape, excuse me, the, the variables that are the most important, are often best to map to position and length rather than color, because it's hard for the human mind to perceive differences in color. And this is something that you see so very often in scientific visualization. That, someone will make an example where they're trying to show difference in volume, where the same numerical difference would be much more clearly demonstrated by differences in length, or in angle, or in slope. So you have to choose the dimensions that you have available, and think about the saliency that they have when people are perceiving what they're looking at, and think about it not you as the scientist, the creator of the visualization, but the human who is not you, who doesn't know the data, will look at something and not be able to perceive differences that are visible to you. So let's try and quantify this a little bit specifically. This'll be something that will probably have occurred in a number of the other lectures. If we take data and map them to different types. So the first is nominal. These are often category labels. So if you have data that describe fruits, the categories would be things like apples and oranges. There might be numerical data, but it's of an ordinal nature, or numerical or non-numerical data that's of an ordinal nature. So the example I put here is the quality of meat, where the example would be that you have A, and double A, and triple A. And what they mean is, that they are relative to one another, but magnitude is not important. Double A is probably not twice as good as A. And then lastly, qualitative, which, excuse me quantitative, which is numerical and often continuous values. That's something like a measure like length. So let's think about, if we have data that have certain types, that if we want to map those particular values to the way that the human mind perceives, you'll see that quantitative values, ordinal values and nominal values all map best to position. But you'll also see that if you, the next best value for nominal is hue. So that means color. So if you want to label categories, the, one of the best ways is differences in color, but I think we'll show in a moment, not continuous color. Whereas if you look at hue under quantitative, it's actually about two thirds of the way down the column. So if you want to show a number, and it has magnitude, and it's a quantitative number, and you're picking hue as the mapping, you should probably think again about that, because it's not going to be easy for the human mind to perceive. Let's pick a more specific example. color. So we've talked previously about this idea, that for, that hue is one of the important ways that we can encode a nominal value. But if you look here, what you can see is that the human mind, based on what we described earlier, and our ability to perceive differences, means that on a, a color scale like black to white, that the human mind is at best able to perceive seven differences. And so, if you look at the continuous values, you can see that co, hue can encode continuous values, but it's far less good than its ability to include ordinal values, which you can use with different magnitude, with hue of a single monotonically changing function. Another thing that I think you can use, is that hue is normally perceived as unordered. And so, this means that hue is a really good way to label nominal variables. And so, when you have things that are, have no ordering, picking color is really valuable, because color, hue does not have an order. Whereas if you look at color value, it does have an order. And so, it effectively is not nearly as good for nominal values, than the unordered color values. And we could show how these are applied if we look at different kinds of the way that color is used. So if we look at for example in this illustration from an anatomy textbook, how different anatomical features are colored using different values. So here, red might mean artery and blue might mean vein and yellow might mean capillary. But they're effectively nominal values. So colors being used nominally. You can look similarly at this presidential election map from the from NBC, and you can see here that we've also used color in a way where red is mapped to the Republicans and blue to the Democrats. So here, we've just simply taken two colors and made a very easily distinguishable map. But if we look at quantitative uses of color, and this is where the most often you most often see mistakes in Earth science data visualizations by quantitative applications of color. So here, what you have is an example of Global Average Temperature and it's rendered from blue to red. Now there're a number of reasons based on what we just learned about color, that described why this is not the best mapping. So first of all, the, there's a, it's very hard for the human mind to perceive more than seven seven different colors. So it's really hard for a hum, for a person to look along the equator and tell me the different average temperature between Panama and Senegal. And they're probably not the same, but in red, they're relatively close. So it's hard for the eye to distinguish values like that. Another example here is this fluid dynamics simulation, where velocity is encoded as color. And what you see here is that these hues, there is no natural ordering. If you look at the yellow in the middle, it's really hard to tell where that lies in color space. And different, our ability to perceive light is emphasizes different kinds of scalar values. So if you look at the way that color can be applied in an ordinal fashion, it's very useful to pick a a single saturation, excuse me, to vary luminance and saturation together, and you get these very nicely marked color scales. Where what you're able to see very clearly are the demarcations between each of the different values. And you can do the same for different types of color, where if color is diverging, in other words, if you want to emphasize the mid point of a color range, you can pick a neutral color in the middle. And when you apply this to a map, you'll see that the midpoint that use saturated colors on the different end points, that they'll be distinct from the midpoint of the map. So just at a high level review, if you're going to be thinking about color in data visualization, you should use only a few colors and that, try to be able to make them as distinct as possible. Strive for harmony, be very cautious of cultural conventions. Be cautious about bad interactions between colors. And one of the best ways to determine if your visualization is going to be effective, is before adding color, see if without it you can distinguish the different variables in black and white. We'll ahead into the next chapter.