[MUSIC]. Hi Everyone. My name is Cecilia Aragon and I'll be your guest lecturer this week. We're going to be talking about basic principles of information visualization. This is in my opinion one of the most important interesting and fun components of the data science pipeline. We're going to be talking about how you take the data that you've collected and analyzed, and then actually present it in a form that makes sense to the human brain. I'm a professor here at the University of Washington in the Department of Human Centered Design and Engineering. And it'll be my pleasure to introduce you to this exciting subject. My research interests lie in the visual presentation of very large, scientific and social datasets. So, let's start with a few definitions. The word visualization is extremely overloaded. There are many possible meanings to the term so what are we talking about when we say we're studying Information Visualization? Well we want to consider the visual representation of information. The goals of Information Visualization include the effective communication of information or data, with clarity that means that the information should be presented in a way that's easy to understand by a human. With integrity, you present all the data and nothing but the data. And you want to stimulate viewer engagement with the data. Our focus in this course will be on effectiveness. We're talking about the clear and effective presentation of data to the human mind. Data visualization usually falls into three academic categories. Information visualization, scientific visualization and visual analytics. Information visualization typically involves abstract information that may not have a position in space. For examlpe, on the slide you can see a graph of stock prices around the year 2000. We graphed the price, the dollar amount on the y-axis and time on the x axis. Scientific visualization usually involves displaying data that has a physical position in space. For example, here you see a methane combustion data set, where each of the data points has a 3D location. Visual analytics is a combination of highly interactive visual interfaces and sophisticated statistical learning algorithms. In this course we're going to be talking about basic principles of information visualizations. We're not going to be talking about scientific visualization specifically, nor about visual analytics, nor about the artistic field of visual design. The principles, however, of information visualization that we are discussing are applicable to all three domains. We will be discussing specifically this week how to map data variables to visual attributes. So, why is visualization important? Say you have a million numbers. What's the best way to communicate that information to the human brain? Well, imagine a table of a million numbers, it seems almost too vast to comprehend, much less define meaningful trends or information about it. But on the other hand, if you consider the number of pixels on your computer screen it's typically around a million and you have no trouble understanding it very quickly. And being able to, to act on the information you [UNKNOWN], so the human visual system is the highest-bandwith channel to the human brain. And it's the most efficient way to present large amounts of information so they make sense. Lets talk about a few examples, so may be the first use of visualization historically was for eliciting knowledge from data, let's suppose you have a table of numbers. We're, we'll start with a relatively small table with only 50 rows, one for each state in the union, three columns in each row. The name of the state, the percentage of college graduates in that state, and the per capita income. So you might have questions about this data, such as, which state has the highest income? Is there a relationship between income and education? Finally are there any outliers? Are there any trends? Well obviously all the data is present, and you could take your time looking at this data. You could identify the highest and lowest numbers. You could answer all these questions from the data. Everything is presented here. But is this the most efficient way to understand this data? Here's another way of presenting the data in a scatter plot. We have per capita income on the x axis and the percentage of college degrees on the y axis. And, as you can see, the question of which state has the highest income is instantly answered. We can see that Connecticut has the highest per capita income. Similarly, it's clear there's a trend that there's a correlation between per capita income and the percentage of college of degrees and there is a couple of interesting outliers. Nevada has a relatively high per capita income but a very low percentage of college graduates. So, all the information is presented in both of these forms but the, the visual representation is, is much more efficient at presenting meaning to the human mind. Well, so, okay, the first question that might come up is, well, yeah, but all this information could more easily have been accomplished by analyzing the data with a statistical program. You could find the maximum. You can find the minimum. You can find the regression line. However, it turns out the human brain is one of the best pattern matching tools that we have. Actually, it's, we're, the human brain's so far is still better at detecting patterns in data than any of our most sophisticated machine learning algorithms. So an immediate question that arises is. Why not just use a statistical program and identify the minimum, the maximum, the trend lines and go from there? Why bother with a visual representation? The answer is simply, the graphs reveal data that statistics may not. So, here's a famous example, Anscombe's Quartet. Here you have four small data sets that are quite different, as you can see by looking at them, but they have identical linear models. That is, the identical means and the identical trend lines. Again, you could just present the data in, in tabular form, but see how much more information we can learn about these, these individual data sets, just by looking at the data. So even though, we have identical means, identical regression lines. These data sets are quite different. So, so statistics may not tell us everything about a particular data set. I hope I've managed to wet your appetite for this coming weeks set of slides. I'm going to stop here and we'll continue on with more examples in the next lecture.