My name is Ashish Mahabal and we'll be continuing with R and looking at ggplot. So, much of this material is taken from Headley's book by the name ggplot2. We start with the anatomy of a plot. Last time we saw a little bit of what grammar of graphics is, so this shows graphically how the elements in the grammar of graphics can be brought about. So you see at the top, that you have mapping of variables to aesthetics. So overall, what happens, is that you have a data set, a data frame as we saw, and a certain set of mappings, one or more, and then there are multiple layers. And for each of those layer, you again have some data which is very likely a subset of the original data that you have. Some mapping, and associated with that, a geom and a stat and a position where that particular layer or subset should go. Then, for the overall plot, you have got one scale for mapping, and then you've got a coordinate system and a faceting specification. So the coordinate system would take care of whether you have log scale or linear scale and so on. And faceting specification is where you will know what kind of categorical variables are being used to separate the data sets into. Either sublayers or subplots and so on. So this overall picture is very good for allowing you to debug your plot. So if you are trying to build a plot using layers and you're not getting what you want, then you can ask yourself whether you have the correct geom, whether you have got the correct stat, and how they should be interacting with each other, where should the position be, whether that is correct, is one of your coordinate system wrong and son on. So by just keeping this figure in front of you, many times it's fairly easy to debug what is it that you want to look at. Se there it is again a summary of what a plot is made of. Again, there is a data set, a set of mappings besides geom set and so on. Right? And so, example of scale can be this. It can be a continuous variable or a discrete variable. Remember that, that discrete variables are the ones that we use for doing categorization and faceting, but even the continuous variables, you can convert them using the as.factor function into a discrete variable and use that as well. So, what is it that actually happens when you make an object? We saw that you can assign your ggplot command to an object called b and not have anything rendered immediately onto the screen. So what is it that you can do with such an object? If b is the object that you have made, then you can print it. And this is something that have been scored automatically if you're not inside a loop. Say if you make this object and simply say P on the command line. It's going to do the printing of P. Then you can save that object. So if you save it will write out an image for you. So, rather than going to the screen, this is what will allow you to write it out to the disk. The other interesting thing is summary, it is as before or loaded and with this object b if you simply, say summary with empty parentheses, then you're going to see what is hidden in the object and we'll do that to all the objects. Similarly, save instead of ggsave is going to save the object itself, rather than as a plot. The object as a R object will be saved to the disc which you can reload later on. So let's look at the summary of the P object that we have defined at the top here. Where we are again taking the crts data set and simply plotting amplitude versus standard deviation there and factoring it into different colors by the object variable. So you see that though we have used only amplitude and standard deviation in the plot the object knows about all the variables that are in the data frame. So all of them are listed here. It also tells you what is the size of the data frame. There are 1619 rows in total with 26 variables. And then it tells you what the mapping is for the X's or the Y's and what facetting had been used here. Now in addition to p, if you were to do a geom_point, like let's do p + geom_point and look for the summary of that, you'll additionally see what, the stat_identity is, what the geom_point is, and what the position_identity is. In this case we have not specified that explicitly, so we get mostly NULLs, and FALSE, and so on. But all that information in principle is there. So that is another useful thing about the laring aspect. You can keep on piling up various layers into an object, and then those layers become available to you in various different ways. So, you can say save p and file name to save the data file, and then you can reload it again using the load command. And then after that you can add various layers to that, so, that's a very convenient thing to do. This is another way example of adding layers. So we use here ggplot, and we are using a slightly different set here. The set is called diamonds, and it is available with the ggplot2 library. So if you have started ggplot2 library the data diamonds will allow you to access these commands. So here we take that set of diamonds and use the aesthetics of number of carats in that and price, and then use a cut for the colors into the plot there. And so now if we do geom of you can add that as a layer. Notice here that we have used slightly different syntax. Rather than saying gm_point, we say that we want a layer and then geom equal to point. Now, this is where you can start combining different stats in geoms effectively. So, next what we are saying is that we can simply plot diamonds with aesthetics of x equal to carrot, so it's simply going to use that particular variable. But then to that p, to that object that we just defined, we can add a layer with a separate geom and a separate stat. So when you plot a histogram, what is it that is actually happening? You bend the data, so that is what you're, statistics is. And what do you actually plot? You plot bars of it, so that is what your geom is, and so this pairing is automatically done from, for histogram. And then when you say simply p, that is what is going to appear for you on the screen. So here is an example of a histogram as it comes out of ggplot2. In this case we have use a binwidth of 0.3. We can of course use quite interesting and sometimes useless stats, or geoms for a particular data set here. I am simply listing a few different geoms like bar, boxplot, contour, line, point, step, text. There are a total 29 of these, and just for fun here, I have got, used the geom of polygons on this particular data set. This is the same dataset that we used in the last lecture, the CIT six class set. And then you can see that it doesn't make sense. But the shape is roughly the same as what we had got when we had plotted x versus y. Similarly, there are 15 different statistics. I have listed a few here. The book and various websites will have more. We will also have it on the website of the course, so we can have contours there, we can have density2d, and that is what I've used here. I've used density2d, and I've said that let use bandwidth that is very small, only 0.01, and that is why you see that towards the lower parts of xy area, you see that there are very dense contours. As you go out, there is less density in, you can see that in the contours there. So some additional points. You can have the ggplot with geom_point as we saw. When you want to add aesthetics, you can give that as as arguments to the geom_point 2. So, if you say geon.color equals to red, then the points are going to be plotted as red. Similarly, in the layers, you can have various smoothing things. So here are shown that for the point method so that you can provide method equal to lm as your layer, and then smoothing can also be done addition, in addition to that. Similarly, you can do, scales and axes. We have not seen an example of that, but that's homework. You can take a look at that. We are adding that as additional layers. So when we say geom point, it's going to take whatever the default scale is. But in addition, you say that oh, I want log scale for x, that's sure enough add that as a layer and you get that. And this particular example, we have added both x and y as log scale. Then similarly, you can add plot options, so you don't have to be stuck with the defaults. The defaults are reasonable, but normally for making plots into publications, so you want the access and labels et cetera, to be much bigger, so you can specify all of that as a additional layer simply by using the command called opts. Here, I have specified for instance, what the title should be and what the aspect ratio should be. And you can add various other commands with that. So, the homework with respect to this is, take the dataset that we have provided and we saw examples using only two variables, amplitude and standard deviation. But there are several other interesting variables. Try to see if you can use them and separate them in many different ways. This, what we are covered here is fairly basic graphics. We were not trying to visualize in any very different or specific way how to go about things. We were looking at only what is available. There will be additional modules in the course where you'll be looking at visualization, how you should be plotting the data, et cetera. Thank you.