[MUSIC]. Last time, we talked about data types, dimensions and how to map data to visual attribute based on the number of data dimensions. Clearly there's a large design space for this visual encoding. Information visualization is about choosing the most effective visual encoding for a given data set. In this lecture, we're going to describe how to do that. We saw earlier that there are different data types. And those can be mapped to a set of visual attributes. And some mapping are better than others. So now we have a challenge. How do we pick the best mapping of all possibilities? So in order to do this we're going to follow three basic principles. Importance ordering, expressiveness and consistency. This was first described by Jock McKinley in 1986. So importance ordering has to do with encoding the most important information in a most perceptually active way. Expressiveness has to do with depicting all the data and only the data. And finally consistency means the property of the image, the visual attributes should match the properties of the data. I'll go into more detail on how to do this on the following slides. We'll start with importance ordering. We've seen this slide before. Now as we discussed earlier, decades of research have taught us that certain visual attributes more accurately represent quantitative data to the human brain. Mackinlay defined expressiveness in the following manner: he said a set of facts is expressible in a visual language if the sentences: i.e the visualizations in the language express all the facts in the set of data and only the facts in the data. So what does this mean? Well, using color hue we cannot express the fact that one item is greater than another. So color hue simply can't represent ordinal data. Ordinal or quantitative data. So color hues simply can't represent ordinal or quantitative data. It's best you use phenomenal data. On the other hand you don't want to express facts that are not in the data. So a length is interpreted as a quantitative value. In this incorrect use of a bar chart, the length of the bars say something untrue about the data. So on the x axis we have the national origin of these cars, now on the y axis the make and model. Now the length of the bars suggests an ordering on the axis. So it implies that somehow Germany is greater than the USA. So therefore using bars for nominal data is an incorrect choice. Number three, we have consistency. The properties of the image, the visual attributes should match the properties of the data. What this means, is you don't want to do something like mapping one dimensional data to two or three dimensional visual representations. This diagram represents why that is not a good idea. So here we have the price per barrel of oil represented as the height of a three dimensional barrel. You can see the obvious problem here. Look at 1974. The price is $10.95 and then for 1979, the price is $13.34. Well this is less than a 30% increase. And yet the 1979 barrel looks significantly larger than the 1974 barrel. So clearly this is an inconsistent representation. The three-dimensional representation is implying something false about the underlying data. Let's move to a data encoding exercise. Here we have a data set of cars. We have the model on the first column. The country of origin, the year, the number of cylinders, the horsepower, the miles per gallon and finally the weight in pounds. I would like you to create a visualization that encodes all seven dimensions in this data set. Make the perceptually appropriate choices and be prepared to explain. I will describe the steps you need to do and show you Bertin's examples once more. So step one the terminal which columns represents nominal, ordinal or quantative data. Step two review the transvisual attributes on the next slide and assign the most perceptually appropriate choices to each data set. Step three create a visualization that encodes as many of the seven dimensions as you can. Don't worry if you can't get all of them but, but do try to do so. To review here are Bertin's visual attributes. Position, size, value, texture, color, orientation and shape. Are you done? Here's one attempt made by one of my students. So this is an attempt at encoding seven variables from the cars dataset in a single visualization. So this student chose to put two of the quantitative variables, horsepower and miles per gallon on the x and y axis. The third ratio data weight he encoded with area of the marks. Not a bad choice since you, you can only put two of those ratio variables on axes. He also chose to encode cylinders, which are quantitative integral data by using the shape of the mark. And then finally he used color hue for the region, Europe, Japan or US. And color value or lightness for the year. Additionally in his visualization you can mouse over each of the marks and the name of the model of the car would pop-up. So, what do you think of this visualization? What works, what doesn't and what could you do better? I'll leave that critique as an exercise for you.