My name is Ashish Mahabal, and we'll be continuing with the best programming practices module. This will be part III of it. So we'll be looking at some different aspects of programming in this. How a program, in a program you should try avoiding duplication. How various subroutines of a program should be orthogonal in nature, and especially, how you go ahead and try to do refactoring of the program so it becomes more efficient. So, I will try to duplicate myself a lot while talking, telling you how important it is not to do duplication, but that specifically applies to your programs. So don't repeat yourself. In a program, you should have functions that do separate things, that are orthogonal from each other. If you have two functions that seem to be doing similar things, that's definitely a case of duplication. And what you should try to do is isolate the part that's common to both, put it in a third function and make the other two functions call them. That way, your functions and your programs will become modular, and they can be more reusable. And that can be helped by being a bit impatient. Don't try to write the same piece of code again and again, if there's a slight variation. Try to reuse the code that you have. That'll also stop you from reinventing things that have already existed. So, many people have written all kinds of modules for this good languages have been developed, so try to use as much of that as well as possible. So, there are cheat sheets that are available. Those indicate various short cuts that are possible, various fast things that you can do by using built in modules. So, you should try to master them as much as possible. So, try to look at the cheat sheets. To take an example of Python, you can visit what is called the cheese-shop there, which has lot's and lot's of resources available, telling you what are the things that you can use without having to rewrite them yourself. Similarly, there is a website called hitchhiker's guide to Python. Take a look at that. And for all other languages you'll find similar websites. We'll have some of those listed in the additional resources that go with this module. So, what does one mean by orthogonality? What it means is that each of your subroutines should try to do something that is distinct from other subroutines. That they should not be overdependent on each of them. You may be able to call one subroutine from another, but you should not try duplicate what the first one is doing in any case. Especially if you have subroutine a dependent on subroutine b and subroutine b dependent on subroutine a, that is definitely not a good idea. That will lead to chaos, sooner rather than later. So, when you have functions or sub routines that are orthogonal, what that also means is that the changes are going to be localized. If you need to make a particular change somewhere, you'll have to make it only in one place. On the other hand, if you have many functions that look like each other, that do things that are similar, and then you discover that you need to change something. You may need to change it in all of those places. And that's definitely not a good thing, because again, you may forget one of those and then much later, you or someone else will discover that bug. Similarly, when specific functions do only specific things, then unit testing becomes much easier. We can write tests in a more straight forward way testing just that aspect of that function. Then reuse becomes easy as has been already said. And then if requirements change for one function, how many functions or modules should get affected, exactly one. And that is the mantra that should be used here. So, that is how you should try to design your functions. And that also makes things extremely configurable. Let's take a look at simple example. If you are defining a line, and as I put too, that function you're given three in puts. The start, the end point, and the length of the line. And go and implement that. And the same function can be written where it takes only two inputs. Just the start point and the end point. Because after all, you can calculate the length from those two inputs. And clearly the second function is more efficient, more manageable. Because in the first instance, instance there may be users who by mistake give three inputs that are not consistent with each other, and then how is that going to affect your program? We just don't know. So, try to make them consistent in that fashion, don't try to, whatever you can calculate or you need to calculate within the function you do it that way. So, when you are using someone else's libraries, if that means that you have to write some special code to handle that, maybe that is not a good idea. You should try to be as general as possible and try to avoid such cases. Which also brings one to global data. Many times, many programs is global data. So a function, if it relies on some global data, it may not be a good idea. You should be passing the global data as an argument. And you may or may not want to change the global data. But simil you should certainly not rely on that. And in programming languages like Python and all that's quite important. to all of them really, but there are specific cases where the scope decides how things are happening. And unless that has been stated explicitly, there may be grief later on. Avoiding similar functions, I've already stated that. Just restating it to emphasize that. So, that brings us to refactoring. You should try to refactor early and often. What that means, again, is that you should, once the program is running fine, then you want to optimize it if you want to benchmark it, and that is where you start seeing better on the functions are orthogonal. Whether they are maintainable, whether the consistency is being checked, whether the, the syntax is the way I want them, and so on. So you can duplication during this process. You can make sure that that you bring in orthogonality to your design. And if there are some outdated knowledge, this is the time to take them out, take that out too. And you can improve performance in the process. One important thing to remember when you're refactoring is that do not try to add functionality while you are doing that, because the two things are quite separate from each other. When you try to add functionality, you are trying to make changes that which have not been tested. Whereas when you're refactoring, you've already tested things, you have gone through it, and then it's working. You are only trying to make it better, make it perform better. So, if you try to do both things, again, that's likely to lead to sorrow later on. You can add good tests during refactoring. And that's always a good thing, I think, more and more tests. And you should do that in very short deliberate steps. Don't try to reformat the whole thing during refactoring also. Take one step at a time. So, all this can be put together in something that's called Design by contract. And in 1997, Eiffel and Meyer had suggested that. So, what a function should do a function should say that I want these specific inputs, and then this is the way I am going to work on those inputs. And when I am done, this is how the output of the program or what the state of the program is going to be. So, it will work on the input that it has given, there are may be some class invariants which may not be touched. So, in this whole process, what is important is that your function should be very strict in what you accept. Function should be able to say, okay, I want three numbers. Two of them should be integers, and one of them should be float. And the function, right at the outside, if it does not get exactly that, should be able to just come out in a graceful way saying that no, no, no. This is not our, not what I want. I want these specific things. Similarly, a given function should promise as little as possible. It should not try to detain different things. It should try to do one specific thing in a proper way. And then again, you should be lazy, and you should try to do it in a, in the most efficient way that is possible. And if you can do this, then inheritance and polymorphism will result automatically. You will be able to overload your functions, you will be able to use the same function in multiple places much easily. So, try to inculcate that as early as possible, designing by contract. Some more aspects, testing how to write tests, how to use tests. What kind of comments should be used? What kind of arguments should be there in the program, and how debugging can be done. So, let's revisit this. We just briefly touched upon it earlier. One thing that you should remember is that someone is going to test your software. And if you don't do it yourself, your users will at that time, then for some cases it may be too late. So, we can do the test also against the contract that specific functions have. So, if a function is for taking square root, then the function may have told you that I'm going to the square roots but I'll do it only for positive integers. I'm not going to deal with complex numbers. So, I won't allow negative integers. Fine. So, as long the contract is there, as long as the function explicit, you can test specifically against that. But assuming that it can be more generic you, you should make sure that the age cases are also tested. So if you provide zero, does it work with zero. What happens when you go out and you get your number minus four in this case, it's returning zero. Maybe that is the situation that is needed here. But, what about things like scientific notation, many times such things are not tested. Here, the first input that's been given is 10 exponent 12. Does I come back and in fact give you So if it's taking only positive numbers surely that should work. So make sure that there are tests available for that. And then, you don't have to write all these tests by hand. There are various test templates available, so you can start with a test template and just start filling that up in the way that you want. And then, while the testing happens, again there are test harnesses available, where all the logs and errors can be easily put together. So though this may seem to be a lot of work, it is not really too much extra work. But, just that discipline of writing tests is going to take you a long way. And of course what is important is to be able to write test that fail. Because if all test just pass then maybe someone has done something, so that all that the tester doing is simply take different arguments and return true, true, true. So, only when you write a test which you know should fail, and then give it to input such at one input is given and the output don't match and it returns a false, you really know that things are working. So, things to keep in mind when you write tests, is that, try to use long subroutine names. And try to prefix them with the word test so you know that these functions are not really the main functions in your program, but just for testing purposes. Similarly, you should have stand alone code specifically for tests. And they should be stand alone data sets. So again, you are not interfering with the main body of the program. And then you should do clean ups during the start of the tests and once the tests are done. So, that way the entire testing routine is a separate thing and doesn't interfere with your main program. So, in case of Python we saw in the first part how unit test can be written, then we saw in the second part how it can be doctests that can go in the docstrings also. Pytest is another simpler mechanism that works with Python, and you should take a look at that. Similarly, there are specific modules like nose and tox, and mock that can also be used. Then coming to comments. I have already stated that, but I again emphasizing comments, documentation are crucial so that knowledge gets transferred properly. And all kinds of bits can be put in comments. Some people think that if something was difficult to write maybe it was difficult to understand, but that's questionable. And you should also remember that bad code requires more comments. So, if you find yourself writing too many comments, maybe you should take a look at the code, and maybe you should try to improve the code itself. As an example, you shouldn't have trivial things, like when you do x equal to x plus 1, have a comment saying oh, I'm implementing x, that's quite obvious from the statement. On the other hand, if you're implementing x for a specific reason that is nonstandard, like you're compensating for the border as has been shown here. Then perhaps you should state that as a comment. You should say that x equals x plus 1 and that it's compensating for the border. So, what kind of comments or what kind of documentation can go to live to, in the comments. So, you can have a list of functions that are being exported. You can have revision history, when and how has the program changed. Remember, in the first part we talked about source control, the source control will keep that off your comments. But, you can also have a set of comments in either a readme file or in the program itself. You can list, the other files that are being used by your program. And the name of the file itself, which has the program. So the documentation can be algorithmic. You can have full line comments that describe the algorithm that you're using. Or you can just simply list anything where, there are offline comments like we talked about bordering commenting. Or the comments can be defensive. And these are very important sometimes. When a piece of code has been slightly tricky to write, but then figure it out, it is worth mentioning that in the comments. That it had puzzled you before, and this is how you have done that. Or you can even have something indicative that you have done a rather quick and dirty job, you'd want to revisit that. The program is working, but user beware, that it better be rewritten in different ways. So, you can have those comments also. And then you can have discussive comments, where details of what is being done can go into plain old documentation, regular documentation, whatever. And then, arguments. That is another important thing. You should try to have arguments that are meaningful, arguments to functions. You cannot, you should not try to have too many arguments being given to a single function. You can bill your universe with various constants. You can give all the constants to the function subroutine, to the universe subroutine. But then if the order of that is given wrong, what then? Then someone is going to make a mess of the universe that it built. So, that can, one can get around that by, again, having deliberate chart steps for your arguments, are named arguments, then it doesn't matter whether you turn around order in which the arguments are given. Then, in your program you should look for missing arguments, definitely. And then whenever possible, you can set default arguments. So, if you are giving five arguments and two of them normally have the same values, you should just set them and let the user change them through the arguments wish to. And as far as return values are concerned, you should try not to make a sub routine, do things just by the side effects. Many people tend to change variables and leave them like that, and pass them out without the return values, if that is possible, if one can have side effects in the particular programming language that you are using. But, that's again a bad thing, because a user may not realize that that is how the code has been written, and they may look just for specific return values. So try to return everything that you are modifying, and that is going to be significant after that. So, here is in case of arguments where you can have the dash dash type and the dash type, and ordinary types where you have named values or you have values being provided along with the keywords in the command line itself or just position values. And then there are normally models of a level in all good programming languages which make use of that and I've given example of Python here, but you can use the getopt model. And when you use the getopt model ,then it can separate your, a different kinds of arguments in a very meaningful way. And then you can handle or use those arguments within your subroutine quite meaningfully. Then coming to debugging. It's of course one important thing to do because you don't want your code to have bugs. And you should remember that there will be bugs. And the only bug-free program is one that doesn't do anything. So you better make sure that you try to do as best debugging as possible. And again, tests, which have been talked about before, come to your aid. Write lots and lots of unit tests to look at each functionality of your program. So, that all of those cases are caught early. And make sure that the programs compile without warnings. Many times, people go ahead when they see that there is no error during compilation. So, even if there are some warnings or this did not work, or there may be an issue with this, so long as it is not a showstopper they just go ahead. But I don't think that's a good idea. And you should definitely take care of any warnings that are being raised, because they are being raised for some specific reason and you'd better take a look at that. So when dealing with bugs, make them reproducible. Be able to show that this bug crops up every time you had in a particular situation. And if you can not isolate the situation then there are some complex [INAUDIBLE]. So, try to isolate that by going down to one subroutine to another subroutine apart of the subroutine, and so on. So that exactly with the single command, you should be able to get the bug reproduced. And then if needed, you can visualize the data. You can set various break points. And these are the standard ways that you can look for bugs. When you do find a bug, there are some trivial things that you can do. First, check for boundary conditions. Is the bug coming up because of the first case or because of the last case, or is it something else, because those are the simpler bugs to catch. Then, this may seem a bit psychiatric, but describe the problem to someone else. Many times when you are doing that, you usually realize what may be going wrong, and you are able to catch the bug. So definitely you've got to try. Then ask yourself, why wasn't it caught before. So, what assumptions you may have made that were wrong. So you, that's definitely worth investigating. And then ask yourself also whether the bug could be lurking somewhere else. If your code is orthogonal, then if you find a bug somewhere, it's very unlikely that it's somewhere else too. But if it's not orthogonal, it could be somewhere else. So, make sure that you find all instances of the same bug, so you can go after the same bug multiple times. And then, you have tests, right? And you have the tests ran fine, the tests didn't catch the bug. Then maybe the tests are bad. You should have a relook at the tests themselves. So this is what you can do about debugging. Next time we'll be looking at the portfolio building and some meta programming.