Today, what we are going to do is we are going to look at a package called astRowRap that we have been writing for astronomers specifically. It takes many interesting programs, methods. That are available in R that astronomers should perhaps been using but have not been using. And we started looking at this mainly from the KDD guide, the data mining guide, of the international virtual observatory community. So a long document as been put together there by a few astronomers and section 7 of it in particular. Talks about methods that astronomers should really be using. And for many such methods, there are packages that are already available in R. And with most packages in R there are some examples available. But many times, these examples are. From biology or some other science. And then if we want to initiate astronomers into that, that's not a perfect thing to do. And that is why what we have been doing is that with some astronomy data sets together, which can be used with that. And astRowRap essentially takes these and puts them together. And this work is, in collaboration with the Data Sky in [INAUDIBLE]. And so, the kind of data sets that we have been using here, are, are historical data sets, interesting small data sets and also, some very modern data sets, from large sky surveys like that Catalina Real Time Transient Survey. And with these, we provide some. What are examples so that astronomers can directly jump in take a look at them. So these will also be available on the sister site for this workshop. So using astRowRap like everything else in R is fairly straightforward. In the console, you would load the library, simply by saying Library(astRowRap) here. the, both the Rs are capitalized, the way they are for R. And then, once you have done that, you can use the standard ??astrowrap to figure out, which all. Packages out of a label libel astRowRap. So there is a list of statistical tools available, and one line descriptions for each are given. And don't try to read what you see on the right, because that's, that's too long a list. And so one of the commands for instance, one of the tools that you can use, is lm, something that we had visited in the very first. Doc the linear regression related thing. So for each of them, you can again give greater help by saying something like: ?astrowrap_lm. And that'll give you more details on that particular method. So the documentation on the test includes the relevant atranomical data set, as well as how to use that particular data set. And there will be a few default methods, like the blot method. That'll be available with each of them. And so here are some of the tests that we have covered. There is a regression analysis, simple linear model I, that I mentioned, and generalized linear model as well, ANOVA. And within clustering there is hierarchical clustering and k-means clustering. On the right hand side, the. That you see is the output of one such gamings clustering. The three different types that you see in that plot are three clusters that are returned by gamings. And we'll see the example in slight more detail, in the next couple of slides. Then there are also tasks for dimensionality reduction principle component analysis, which allows you to take. Lots of different dimensions in a given data set and see which of them are dependent on others and reduce them to a more usable set or linear discriminant analysis that is used a great deal also. So, we will see later on brief example of that, another Cmds like Biplot and Bootstrap etc. The data that have been included have CRTS light curves. So these light curves, again as I mentioned, come from the Catalina Real-Time Transient Survey. There are as many as 500 million light curves and those data points are over a period of ten years. And that is a fantastic time series dataset and can be used for most of these tools in a fairly straightforward manner. Then there are some older datasets like the Faber-Jackson dataset which provides absolute magnitude and velocity dispersion of different types of galaxies. That can be used to use the Clean Air modeling. Are the color magnitude data of COMBO-17 galaxies which can be used for H-clustering. Those are the specific examples that we have provided with cluster. Are properties of globular clusters from NGC5128 galaxy. This is the one that we use with gamings and we'll go into on the next slide. Magnitudes with loss of quasars are available in different optical wave-bands. This for generalizing and modeling. But of course each of this data set can be used for any other test you want to run. Because. There are many commonalities in them, although there are differences in terms of the number of rules in the label, the number of columns in the label, and so on. So here is how one would run the Kmeans example, with the dataset that has been provided. The data set is for global clusters in NGC5128. You would lower that by simply sating data NGC5128. And then the columns that you have there are the ones that can be used of our pca. So you get NGC5128_pca, that's what we are calling here. Then you can find out the summary of the astrowrap. And so, astrowrap summary with principle components of that will provide you a summary. I'm not showing the output here fell very deliberately. You can go through those examples and then finally you can get the gamings of that by. Passing the Kmeans mattered to astrowrap. So that is how you would get the plot that we saw a bit earlier, so three different clusters is what it will come up with. So I encourage you to give it a try, see what you find with that. Now, another new thing that has come about just last year, is something called swirl, which stands for statistics with interactive R learning. This allows you to learn R in R, so these are demos, once you start something it'll ask a specific questions and then it'll expect you to give the correct answer and then guide you into giving that answer in case you give it wrong. And then, once you get a correct answer. Then it will go onto the next question again guiding you by a some examples as so on. So we'll be combining AstRowRap with {swirl} as well so there will be some sort of modules that will be a little bit ASTRowRap. So it is for educators and learners where the idea came from. At Johns Hopkins University Nick Carchedi. And then you can find out more information on that using swirlstats.com. So you can also find on GitHub many different examples of swirls. Again, that is another thing that I would encourage you to download and take a look at. Swirlify is another package that is available with Swirl. What it allows you to do is that if you have an interesting idea, then you can swirlify your idea using swirl so that demo can be made for your idea where you provide a set of questions and,. A set of answers for that. So, you can also take a look at what we have swirlified in astrorab. So, the way to load swirl is very straightforward, again, as usual, library(swirl) will do it for you. And simply invoked Swirl with empty parentheses will then tell you what are the different demos available within Swirl for you. And as I said, we have been combining it and the example that you'll find in the sister website is using the linear discriminant analysis of that using CRTS data. Next time, we'll see how object oriented programming handles the two different kind of classes, S3, S4, and some of the other classes that are available with R