1
00:00:01,700 --> 00:00:05,010
Today, what we are going to do is we
are going to look at a package called

2
00:00:05,010 --> 00:00:09,530
astRowRap that we have been writing for
astronomers specifically.

3
00:00:09,530 --> 00:00:13,650
It takes many interesting programs,
methods.

4
00:00:13,650 --> 00:00:18,450
That are available in R that astronomers
should perhaps been using but

5
00:00:18,450 --> 00:00:20,070
have not been using.

6
00:00:20,070 --> 00:00:26,380
And we started looking at this mainly from
the KDD guide, the data mining guide,

7
00:00:26,380 --> 00:00:29,110
of the international virtual
observatory community.

8
00:00:29,110 --> 00:00:35,240
So a long document as been put together
there by a few astronomers and

9
00:00:35,240 --> 00:00:37,070
section 7 of it in particular.

10
00:00:38,160 --> 00:00:42,600
Talks about methods that
astronomers should really be using.

11
00:00:42,600 --> 00:00:47,460
And for many such methods, there are
packages that are already available in R.

12
00:00:47,460 --> 00:00:51,920
And with most packages in R there
are some examples available.

13
00:00:51,920 --> 00:00:54,430
But many times, these examples are.

14
00:00:54,430 --> 00:00:56,840
From biology or some other science.

15
00:00:56,840 --> 00:00:59,470
And then if we want to initiate
astronomers into that,

16
00:00:59,470 --> 00:01:01,470
that's not a perfect thing to do.

17
00:01:01,470 --> 00:01:04,840
And that is why what we have been
doing is that with some astronomy data

18
00:01:04,840 --> 00:01:07,190
sets together,
which can be used with that.

19
00:01:07,190 --> 00:01:11,490
And astRowRap essentially takes these and
puts them together.

20
00:01:11,490 --> 00:01:15,820
And this work is, in collaboration
with the Data Sky in [INAUDIBLE].

21
00:01:15,820 --> 00:01:21,230
And so, the kind of data sets that we
have been using here, are, are historical

22
00:01:21,230 --> 00:01:25,060
data sets, interesting small data sets and
also, some very modern data sets,

23
00:01:25,060 --> 00:01:29,720
from large sky surveys like that
Catalina Real Time Transient Survey.

24
00:01:29,720 --> 00:01:31,800
And with these, we provide some.

25
00:01:31,800 --> 00:01:37,160
What are examples so that astronomers can
directly jump in take a look at them.

26
00:01:37,160 --> 00:01:41,750
So these will also be available on
the sister site for this workshop.

27
00:01:42,750 --> 00:01:47,760
So using astRowRap like everything
else in R is fairly straightforward.

28
00:01:47,760 --> 00:01:50,270
In the console,
you would load the library,

29
00:01:50,270 --> 00:01:52,940
simply by saying Library(astRowRap) here.

30
00:01:53,950 --> 00:01:58,320
the, both the Rs are capitalized,
the way they are for R.

31
00:01:58,320 --> 00:01:59,810
And then, once you have done that,

32
00:01:59,810 --> 00:02:06,200
you can use the standard ??astrowrap
to figure out, which all.

33
00:02:06,200 --> 00:02:09,150
Packages out of a label libel astRowRap.

34
00:02:09,150 --> 00:02:14,650
So there is a list of
statistical tools available, and

35
00:02:14,650 --> 00:02:16,540
one line descriptions for each are given.

36
00:02:16,540 --> 00:02:18,700
And don't try to read what
you see on the right,

37
00:02:18,700 --> 00:02:21,230
because that's, that's too long a list.

38
00:02:21,230 --> 00:02:26,120
And so one of the commands for instance,
one of the tools that you can use,

39
00:02:26,120 --> 00:02:29,840
is lm, something that we had
visited in the very first.

40
00:02:29,840 --> 00:02:32,600
Doc the linear regression related thing.

41
00:02:32,600 --> 00:02:33,870
So for each of them,

42
00:02:33,870 --> 00:02:39,710
you can again give greater help by
saying something like: ?astrowrap_lm.

43
00:02:39,710 --> 00:02:43,500
And that'll give you more details
on that particular method.

44
00:02:43,500 --> 00:02:48,180
So the documentation on the test includes
the relevant atranomical data set,

45
00:02:48,180 --> 00:02:50,910
as well as how to use
that particular data set.

46
00:02:50,910 --> 00:02:55,270
And there will be a few default methods,
like the blot method.

47
00:02:55,270 --> 00:02:57,690
That'll be available with each of them.

48
00:02:59,120 --> 00:03:02,580
And so here are some of
the tests that we have covered.

49
00:03:02,580 --> 00:03:05,550
There is a regression analysis,
simple linear model I,

50
00:03:05,550 --> 00:03:09,330
that I mentioned, and
generalized linear model as well, ANOVA.

51
00:03:09,330 --> 00:03:12,005
And within clustering there is
hierarchical clustering and

52
00:03:12,005 --> 00:03:13,250
k-means clustering.

53
00:03:13,250 --> 00:03:15,410
On the right hand side, the.

54
00:03:15,410 --> 00:03:19,020
That you see is the output of
one such gamings clustering.

55
00:03:19,020 --> 00:03:22,129
The three different types
that you see in that plot

56
00:03:22,129 --> 00:03:24,980
are three clusters that
are returned by gamings.

57
00:03:24,980 --> 00:03:30,270
And we'll see the example in slight more
detail, in the next couple of slides.

58
00:03:30,270 --> 00:03:33,610
Then there are also tasks for
dimensionality reduction

59
00:03:33,610 --> 00:03:36,320
principle component analysis,
which allows you to take.

60
00:03:36,320 --> 00:03:39,220
Lots of different dimensions
in a given data set and

61
00:03:39,220 --> 00:03:44,740
see which of them are dependent on others
and reduce them to a more usable set or

62
00:03:44,740 --> 00:03:48,200
linear discriminant analysis
that is used a great deal also.

63
00:03:48,200 --> 00:03:51,260
So, we will see later on
brief example of that,

64
00:03:51,260 --> 00:03:54,730
another Cmds like Biplot and
Bootstrap etc.

65
00:03:55,830 --> 00:04:01,880
The data that have been included
have CRTS light curves.

66
00:04:01,880 --> 00:04:04,730
So these light curves,
again as I mentioned,

67
00:04:04,730 --> 00:04:07,480
come from
the Catalina Real-Time Transient Survey.

68
00:04:07,480 --> 00:04:11,550
There are as many as 500
million light curves and

69
00:04:11,550 --> 00:04:16,070
those data points are over
a period of ten years.

70
00:04:16,070 --> 00:04:19,560
And that is a fantastic
time series dataset and

71
00:04:19,560 --> 00:04:24,140
can be used for most of these tools
in a fairly straightforward manner.

72
00:04:24,140 --> 00:04:28,630
Then there are some older datasets
like the Faber-Jackson dataset which

73
00:04:28,630 --> 00:04:34,990
provides absolute magnitude and velocity
dispersion of different types of galaxies.

74
00:04:34,990 --> 00:04:38,850
That can be used to use
the Clean Air modeling.

75
00:04:38,850 --> 00:04:43,390
Are the color magnitude data of
COMBO-17 galaxies which can be

76
00:04:43,390 --> 00:04:44,390
used for H-clustering.

77
00:04:44,390 --> 00:04:48,060
Those are the specific examples
that we have provided with cluster.

78
00:04:48,060 --> 00:04:51,220
Are properties of globular
clusters from NGC5128 galaxy.

79
00:04:53,270 --> 00:04:58,200
This is the one that we use with gamings
and we'll go into on the next slide.

80
00:04:58,200 --> 00:05:02,630
Magnitudes with loss of quasars are
available in different optical wave-bands.

81
00:05:02,630 --> 00:05:04,720
This for generalizing and modeling.

82
00:05:04,720 --> 00:05:10,410
But of course each of this data set can be
used for any other test you want to run.

83
00:05:10,410 --> 00:05:11,160
Because.

84
00:05:11,160 --> 00:05:15,270
There are many commonalities in them,
although there are differences in terms of

85
00:05:15,270 --> 00:05:19,520
the number of rules in the label, the
number of columns in the label, and so on.

86
00:05:19,520 --> 00:05:24,310
So here is how one would
run the Kmeans example,

87
00:05:24,310 --> 00:05:26,350
with the dataset that has been provided.

88
00:05:27,380 --> 00:05:30,900
The data set is for
global clusters in NGC5128.

89
00:05:30,900 --> 00:05:35,580
You would lower that by
simply sating data NGC5128.

90
00:05:35,580 --> 00:05:41,790
And then the columns that you have there
are the ones that can be used of our pca.

91
00:05:41,790 --> 00:05:46,470
So you get NGC5128_pca,
that's what we are calling here.

92
00:05:47,740 --> 00:05:51,760
Then you can find out
the summary of the astrowrap.

93
00:05:51,760 --> 00:05:55,110
And so, astrowrap summary
with principle components of

94
00:05:55,110 --> 00:05:56,320
that will provide you a summary.

95
00:05:56,320 --> 00:05:59,650
I'm not showing the output
here fell very deliberately.

96
00:05:59,650 --> 00:06:02,400
You can go through those examples and

97
00:06:02,400 --> 00:06:06,960
then finally you can get
the gamings of that by.

98
00:06:06,960 --> 00:06:09,710
Passing the Kmeans mattered to astrowrap.

99
00:06:09,710 --> 00:06:13,820
So that is how you would get the plot
that we saw a bit earlier, so

100
00:06:13,820 --> 00:06:16,750
three different clusters is
what it will come up with.

101
00:06:16,750 --> 00:06:21,440
So I encourage you to give it a try,
see what you find with that.

102
00:06:21,440 --> 00:06:27,640
Now, another new thing that has come about
just last year, is something called swirl,

103
00:06:27,640 --> 00:06:30,400
which stands for
statistics with interactive R learning.

104
00:06:31,490 --> 00:06:38,150
This allows you to learn R in R,
so these are demos,

105
00:06:38,150 --> 00:06:40,910
once you start something it'll
ask a specific questions and

106
00:06:40,910 --> 00:06:44,230
then it'll expect you to
give the correct answer and

107
00:06:44,230 --> 00:06:49,050
then guide you into giving that
answer in case you give it wrong.

108
00:06:49,050 --> 00:06:52,000
And then, once you get a correct answer.

109
00:06:52,000 --> 00:06:55,590
Then it will go onto the next
question again guiding you

110
00:06:55,590 --> 00:06:58,730
by a some examples as so on.

111
00:06:58,730 --> 00:07:02,670
So we'll be combining AstRowRap
with {swirl} as well so

112
00:07:02,670 --> 00:07:06,240
there will be some sort of modules
that will be a little bit ASTRowRap.

113
00:07:06,240 --> 00:07:10,770
So it is for educators and
learners where the idea came from.

114
00:07:10,770 --> 00:07:14,010
At Johns Hopkins University Nick Carchedi.

115
00:07:14,010 --> 00:07:20,564
And then you can find out more
information on that using swirlstats.com.

116
00:07:20,564 --> 00:07:25,570
So you can also find on GitHub
many different examples of swirls.

117
00:07:25,570 --> 00:07:29,240
Again, that is another thing that I
would encourage you to download and

118
00:07:29,240 --> 00:07:30,440
take a look at.

119
00:07:30,440 --> 00:07:34,250
Swirlify is another package
that is available with Swirl.

120
00:07:34,250 --> 00:07:37,680
What it allows you to do is that
if you have an interesting idea,

121
00:07:37,680 --> 00:07:41,920
then you can swirlify
your idea using swirl so

122
00:07:41,920 --> 00:07:48,300
that demo can be made for your idea where
you provide a set of questions and,.

123
00:07:48,300 --> 00:07:50,270
A set of answers for that.

124
00:07:50,270 --> 00:07:54,470
So, you can also take a look at what
we have swirlified in astrorab.

125
00:07:55,540 --> 00:07:58,440
So, the way to load swirl is
very straightforward, again,

126
00:07:58,440 --> 00:08:00,510
as usual,
library(swirl) will do it for you.

127
00:08:00,510 --> 00:08:05,080
And simply invoked Swirl with empty
parentheses will then tell you

128
00:08:05,080 --> 00:08:09,520
what are the different demos
available within Swirl for you.

129
00:08:09,520 --> 00:08:12,360
And as I said,
we have been combining it and

130
00:08:12,360 --> 00:08:16,350
the example that you'll find
in the sister website is

131
00:08:16,350 --> 00:08:20,990
using the linear discriminant
analysis of that using CRTS data.

132
00:08:23,010 --> 00:08:28,070
Next time, we'll see how object oriented
programming handles the two different kind

133
00:08:28,070 --> 00:08:32,070
of classes, S3, S4, and some of the other
classes that are available with R

