My name is Ashish Mahabal and we'll be looking further at R ding'g. So let's start with the assignments. We saw last time that you can use the less than arrow and dash to do the assignment that is shown in the second line here and that is what is preferred. So here we show that, you are assigning 3.14 to z. As, we said before doing z = 3.14 is also possible but, you should at best avoid that. And there are a couple of good reasons for that and one of the reasons that people feel that using two key strokes is not a good idea but most editors that go with r allow you to use the less than dash including the spaces on the side with the single key stroke so you should. Get into that soon in emags as well as in our studio. There are keystrokes available for that. But the main reason is that, equal to gets used for keywords. Look at the, next line. We saw an example of rnorm before too. Here we are taking an rnorm, 100 numbers in that vector and there we say mean equal to five. So, here the keyword is mean and we are saying that we want the mean of five for that vector. Here you cannot use the less than dash equal to is reserved specifically for that and then people tend to get confuse if you say a equal to rnorm 100 and so on. Plus, there are two other things. If you want to use global variables, and we'll be seeing an example in the next slide, there you have to use less than, less than, dash. So the next line shows how you are assigning seven to the vector y. More than global, it's actually an assignment to, the enclosing scope. So whatever scope encloses that particular statement, that particular variable gets that in that particular. And what I meant. And of course if you are using equal to, you still can not use equal to equal to do the global assignment. Because equal to equal is of course for checking something. So z = = 3 is where you are trying to check for z = 3. So that can get quite confusing, so you should always try to use less than dash, for the assignment. And then for global of course, you have to use less than, less than dash and then for keywords equal to and for checking you have to use the two equal tos. So, get that get into that habit as soon as you can and stay that way. So this is how, when we use in closing scope a variable set. So if at the prompt, you simply say, bar, where you are trying to find out the value of bar. And if bar has not been set before, then it's going to return an error, that, I don't know what bar is. But then, what you can do, is that you can set bar in different ways. Here we see also how to define a simple function. And, the function is defined with the key word function, and then it is followed by parentheses. And, within the parenthesis you can give it the arguments that you wanted to give. And, the return value is the object foo, in this case. So, that is the name of the function that we'll be using. And all that we are doing within the function in this particular case is setting bar to 1, but what we are saying is that set the bar equal to 1 in the enclosing variable, in the enclosing scope. So if you now come out of the function definition and if you run foo. Then it's going to do that setting for you. And then if you ask what varies, the variable's not defined outside the function. You'll get the appropriate value for bar equal to 1. So again less than less than dash will set the, variable in a higher in closing scope. And that is what you should do when you need to do that well. Avoiding something like that is always good, but that is the facility that is available. Then again, we can assign entire vectors using the combine, so we saw this example before, how you can assign several different values using combine to variable x. So what is actually happening here, is a specific function called assign is being called and the function assign text to arguments. The first argument is the variable name that you are giving, x in this case, and the second argument is a combined set based on whatever inputs you have given, so that is what is into a living core. There's another quirk with assignment that you can do in R and that is another thing that you should avoid as best as possible. You can do an assignment on the right hand side also. So you can say combine these five values for me and use a dash and greater than and then a variable in there. That'll work equally well as the first, argument, but again, you shouldn't use that unless you have to for some specific reason. And then, you can easily work on entire vectors as if they were variable, so if you have assigned five values that combine to x, you can do 1 by x and get reciprocals of all those numbers. Similarly, you can combine vectors with each other. Here in the next line we see that we took the vector x, a zero, and the vector x again, combined them so we had five values in x and this combination is going to have 11 values. So those 11 values can be assigned to y. Similarly, you can use the function, repeat. And now I'm trying to assign to v, a combination of things here. I'm saying that repeat x and add to that y and add to that one. So one of course is an atom and y is something that has 11 values, whereas x has five values. So what it's going to do, is, it's going to try to repeat x as many times as necessary. So that its length equals the longest of the three arguments on the right hand side. In this case, y is 11. So it'll repeat x2 point 2 times. And then it'll repeat one 11 times. And so you know have three sets of 11 values. And those would be added as the first one and the first one and so on, and you'll get a vector v, which will, again, have 11 values. [COUGH] And of course you can, combine many of these things very easily so in the next one, we are finding out variance is, we say that, that the mean of x, over five values. Then you subtract the mean from each of the values of x squared divided by the link of x minus 1. So, you can combine those easily, write those as functions but all those things are available as various functions or so. And remember help is available whenever you want. That's simply saying help sum is going to tell you, what it is that sum is actually doing. Another important construct is an R frame. And an R frame, is easy to read into, so you can have a simple space separated file. So here I am talking about a file that contains simply three lines. The first line has num space name, one space, value two. And the next two actually have some values. One space one point one space three and two space four point four and space four. Right? What you see in the table on the right hand bottom. So if you have an ASCII file of that nature, then you can read it [INAUDIBLE] into R, you can simply say read.table. And the filename in this case that I'm using is called foo. So I read in the file, foo. And I set header equal to true meaning that the first line there, is going to define the names of the columns for me. And once that is done, x now contains the entire file and you can do various things with that. Objects is a way to look at the objects that your R current in a location of R knows about and if you give it without any arguments just the empty parentheses then it'll tell you about all objects that R knows about in your current invocation of R. But X is where we have rate our table into, so you can also say that tell me what are the objects available in X. And then, what it'll show you are the names of the columns that you have and if you type simply X then the entire frame will be printed out for you. Name1, is one of the variables there. If I simply say Name1 at the R prompt and each of the commands here that you see are something that you can give on the R prompt. So if you say simply Name1, it's not going to give you anything back because it doesn't know about Name1. The Name1 happens to be inside the object X. And to access objects within another object, you have to use the dollar symbol. So X dollar Name1, is what will give you the Name1 column there. That is where namespaces come in. That's another thing that we had, talked about in the best programming practices so you should try not to modify things by combining namespaces. But, if you have to do that or if you're going to use only one name space, you can do that by saying attach X. When you do attach X, then the objects within X become available to you on the command line or in the enclosing environment that you're using. And after that, if you say Name1, then you'll see the column called Name1. Redirection, is easy to use in R. Just like in Unix, you use the less than to redirect from a file, or greater than to direct to a file. Similarly, in R you can use source and sink. So source myfile.R will read from that file already with the commands that you have given the, and sink outfile will write to file whatever output that you are creating. Just like the unix command T, you can also sink, use sink with a keyword called split=TRUE. So when you do that, you'll see the output on the command line as well as save it to a file. And you can also capture the output, into a variable, instead of a file by using capture.output. So here we are saying that if we want to capture the output of one example a file is going to give you, that will be stored in that particular variable. We talked a little bit of, about Rdata and Rhistory. Rhistory saves the commands that you have been given so that you can easily go back to that. Unless, list for you radius objects. And you can, use this recursively or in tandem with other objects. So the last line here shows how you can say that, assign the output of LS, the objects that you had. To list, then you can remove everything that is in list. Of course, you shouldn't try to do that in a script, because then whatever you have done up to that point in the script, is going to get erased. So in R, everything returns something. In that sense, everything is a function. Then other things to remember is that spaces do not matter, but capitalization does, so if you are using capitals in some cases you have to continue using those. Parameters often have odd names. They can be named, that's a very good thing, but because of the organic nature of how different packages are, some of them can have fairly odd names, so you should get used to those. You can use NA for missing data and is.na with empty parenthesis is a very useful construct. So you should, try to calculate that early on. So it does the corresponding test to check whether a particular point is available or not. We saw X reading the table. There we saw how you can read it with the header information. But you can also specify the column names. In the command line itself. So all these functions you should look up the help. So help read.table will give the details of a large number of key words that are available with it. If you don't give the names through the header R through the command line called read dot table, you can always give names later on by simply by saying names of X equal to and give whatever names you want by combining the, combining function. Then the large number of standard functions are available, the plus, minus power etcetera are their various trigonometric once are available and then there are the function like sort.list and order or so. The pmax and pmin are interesting in that they return vectors. So, if you had a data frame which had numbers and you say the max of that data frame, it won't worry about the different columns and rows, it'll combine all of them as if it was a vector and give you a single number. On the other hand, if you use pmax, then its going to give you a column, y is maximum. And there are various other such functions available also. Square root as you can see here we are getting a square root of a complex number. So it's clearly overloaded. Because you can use the same square root for just ordinary numbers and, get it equally well. Sequences are another important construct. Simply doing 1:30, gives you the numbers one through 30. Again, R starts with one, not with zero so you should remember that. In this case, colon is also a function. So, when you say 1:30, that function is being invoked. And, colon binds in a strong way, so if you do something like assign ten to n and then say that I want segments one going from one, two and minus one, using the colon function, then the colon is going to bind, so you're going to get one colon and first and then one bind will be subtracted from each of them. So, if you want to do something slightly different you may want to enclose n minus in two bracket, similarly 2*1:15 is going to give you 2,4 and all multiples of two, up to 30. So, sequences can also of course take name parameters you can use from and to and by and length, so here you can have sequence of length 50 when starting from minus five with jumps of point two, that's a trivial thing to do. You can also use a repeat with that so sequence of five we are assigning. Add to that a repetition of x times sequence of five so it's going to take, it's going to give you x, and then another copy of x, then another copy of x. Instead of that, if you wanted repeats of individual elements first, then you can do each equal to five rather than times equal to five. So again, doing just help repeat is going to give you all the good details of what is available. Another important construct is the logical vectors. So you can assign to n x, but only where x is greater than 13. So, what this is going to do, is it's going to find out which all values are greater than 13, assign them to TRUE, assign TRUE to them, assign FALSE to the others, and then write the vector n. So the vector n is going to have the same length as the vector x. You can similarly do intersection by using and, union using R, and negation using bank. Remember also that FALSE becomes zero and TRUE becomes one when you coerce them and you say that, okay, you want numbers rather than the logical values. And then similarly, missing values, you can use them wherever you want to. And that's another very important aspect of R. Indexing is also a very trivial thing to use. Just square brackets and 1:10 is going to give you the first ten values of it. Similarly, if you want to leave out something, that's easy also. Just minus and then in round brackets 1:5 will do that. Replacing missing values. That's something that Is needed very frequently when you deal with data sets. So you use the function is.na(x) that finds out for you where that particular thing is TRUE. And then, only for those values. You assign zero. So immediately all the, missing values will be separate, replaced by zero. And then next line shows how absolute y can be defined in a different way also. Now, z assigning to z something like 6,7,8, that vector, and then trying to replace only one value out of that. What it's going to do, is that it'll make the 6, 7, 8 into 6, 7, 5. But when that is done, a copy of z is first made. And in the copy, you replace the third one with 5. And when this is being done, actually, the square bracket less than dash is being used as a function. And, so again many many functions are used in very interesting ways internally. Similarly you can go to arrays and matrices. If you simply say c(3,5,100) you are going to get a vector with those three values. But instead of that if you assign that to dim(z) then you are going to get a 3D array with those sizes. So you will get an array of size (3, 5, 100) Similarly, you can combine radius sections all of 3D arrays and combine them in a preview version. If you want the entire array, you can use commas. So, square brackets and then separated by empty commas is going to give the entire array. Next couple of lines indicate something quite interesting. Here, what we are doing is that we take the numbers one to 20, we make that into an array, and the key word dim tells us that it should be a four by five array, other than, say, a two by ten array. And we assign that to x. Then we make another three by two array. In this case we call it i, and there we are denoting the values that you want within the array by 1:3, 3:1 and those six values now, we say that arrange them in a three by two damages. And then if you say something like x of i and assign zeros to those, then what's going to happen is that the three by two array that is in i is like three pairs. And those three pairs represent three positions, and those three positions are looked up in the array called x and it is those three positions that are set to zero. So in this particular case, you can work that out as an exercise. The positions 9, 6, 3 are what corresponds to what we have set in i, and those values within x will get set to 0. So, you can do fairly complex things in that version. And again, for data massaging and data munching these functionalities are very useful. Now, looking that variant, what you can do is that you can combine different structures into a single structure, also. So, here, we have something called name, Fred. Wife, Mary. Number of children, three. And child ages, a vector. So, we have a vector here, a number here, and some. Strings, those can all be combined into a list and assigned to something called Lst. You should remember that these are always numbered. So if you want, the value of, in the fourth parameter of this particular list, you can say Lst with two square brackets of four. And then it's going to return you. The vector that is there for child.ages. And if you want the age of a particular child, then you will suffix that list with two square brackets of four with two. So that'll give you the second value of seven. If you were to leave out that second pair of square brackets, that is the enclosing second pair of square brackets from the four. It's going to return null to you. And then, if you were to ask only for a list of three is going to give you a number back. And so, remember that if you want to get to one of the specific elements you can use the square brackets combined with individual elements within that particular one. Then there are matrix operations. You can simply do A * B, element by element matrix product or, you can do matrix multiplication. Again, look at the slightly quirky base this is used,% * % for matrix multiplication. And there are a large number of matrix operations that you can do. I mentioned earlier that re, reading a file read the table that we did, you can look at help of that and you can see lots of of various keywords that you can use with that. There are some built in ones to specifically read.csv, read.csv2 and so on, so you should also look at that. Then we'll be, look, looking at more details about various other commands and objects within that. So, attach, we have already seen here. I am simply giving an example of plotting that. You can read a data set. And then you can attach those variables within that and you can simply plot two different variables in there. Pairs is one such example we'll be seeing more details about plotting later. So I just wanted to give you a hint of how to use that within a single command. So let's not go into details of that right now. And, next time we'll be talking about built in datasets. I could just be bugging and get to the basic plotting that I hinted at here.