[MUSIC] This exercise is going to be on the PsiBlast tool at NCBI. PsiBlast stands for position specific iterated blast, which is a specific form of blast, designed to find distant homologs of a particular protein or DNA sequence. Our objectives for this exercise include setting up the parameters of PsiBlast, specifically to get the right threshold for future iterations of PsiBlast. We will enter a protein sequence to run the first iteration of PsiBlast, and then from those results, we will be able to compile a group of sequences to run a second iteration to find other members of that protein family. Finally, we will be looking in more detail how to interpret those results, which can be a little bit tricky. First, we will retrieve the protein sequence that we were going to use for PsiBlast, and that sequence is available at NCBI. Specifically, we'll be using a transposase from Pyrococcus furiosus. Pyrococcus furiosus is a hyperthermophile that grows at 100 degrees Celsius. It comes from a volcanic vent. The organism was first isolated off the coast of Italy from a volcanic vent on the Mediterranean floor. Now I'm going to click FASTA to get the actual sequence, and here is the sequence. What you might notice is that the first line is descriptive, describes the transposes for Pyrococcus furiosus. And then below that are a series of letters, which represent the amino acids in that particular protein. Okay. You might notice that the top line, which I have highlighted, is the description. The description says that this is a transposase from Pyrococcus furiosus, and then below that are a series of letters, and these letters represent the amino acids in a particular sequence. Now the the FASTA formatted sequence includes both the amino acids sequence and the descriptions, so I'm going to copy and paste that. And now I'm going to go and do the PsiBlast search. To do that, I've got to go to the blast page at NCBI, and you can Google that or type in NCBI, for National Center for Biotechnology Information, .NLM, for National Library of Medicine, .NLH, for National Institutes of Health, .gov, and then slash last. That will take me to the main blast page, and I'm going to click protein blast. And if you notice, by protein blast, you'll see a series of algorithms. The algorithms include, blast p, psi blast, phi blast, and delta blast. So, we're going to click protein blast, and the first thing I'm going to do is paste in the sequence into the window. So there is my, the transposee is from pyrococcus furiosus. Okay. Now I'm going to select PsiBlast. [BLANK_AUDIO] And I'm going to choose my database. My database, in this case, is going to be Reference Proteins. It's a more high quality database than the NR database. The NR database stands for non-redundant, but there are plenty of redundant sequences. It's a huge database with a lot of low quality sequences. We'll switch to the reference proteins, and one thing I want to do before actually running the PsiBlast is you'll notice algorithm parameters in the lower-left corner. We're going to expand that window, and I'm going to change the threshold. Scroll all the way down. You'll see the default PsiBlast threshold is 0.005. I want to be a little more stringent than that. I'm going to use 1 times 10 to the minus 4th to make sure that I include only members of the transposes family when I do the second iteration. So I'm going to change that to 0.0001. Okay, so I have my protein sequence. I have my database. I selected PsiBlast, and I've changed my threshold. One thing I think I want to do is limit to a particular type of organism, so the goal of our exercise might be to find homologs or relatives of this protein in a particular type of bacteria. So I'm going to enter Lacto Bacillus, which will give me lactic acid bacteria. Many different types, and you see that that's taxonomic ID 1578, so I'm going to click that. And now the results should come up shortly. [BLANK_AUDIO] Okay, so here are the results, and you notice that there are a series of alignments better than threshold, and there are a whole bunch of relatives of the pirachocas transposes in Lactobacillus. And then you also see a series of proteins that scored worse than the threshold. Now typically for blast, if a, if an e value, is less than 1 times 10 to the minus 4th, such as this one, 6 times 10 to the negative 5. We can consider that a homolog or a relative of the protein that we entered. So everything above this line we can very, very confidently say, yes, these are homologs of that pyrococcus transposase. Anything below the line we're not quite as sure of. [BLANK_AUDIO] Okay, but we see that there are some transposases there. So what we can do is examine the results. Notice that almost all of these results are from transposases. Even some of the ones that are worse than threshold are transposases, but notice that everything above the line, every protein above the line, has a check. On the right side in that check box, and if you notice the header under worse for threshold says select for side list, anything below that threshold is not checked. So what PsiBlast is about to do in the second iteration is develop a position specific scoring matrix or possum as we commonly call them. The possum is a mathematical representation of the sequence that we started with and all of these relatives that are above threshhold, and if I wanted to manually exclude a protein, I say, oh, yeah, maybe that's not a relative, I could do that. I could also take a sequence that's worse than threshold. I'll actually take the second one here. It's a transposase. Its e value is 1 times 10 to the minus 4, probably right on the border. So maybe I'll include that in the PsiBlast, and now when I hit go, I'm going to run a second version of blast, a second iteration, except not with a query sequence. Now with this possum, that mathematically represents all the of the results in the first iteration of blasts. Okay, and these results should come shortly. Okay, now what we notice is, first of all, a bunch of very, very strong e values. 2 times 10 to the minus 122. I'll talk more about that in a second. You'll also notice that some protein sequences are highlighted in yellow, and the sequences that are in white, without the yellow highlighting, are protein sequences that were found in the first iteration. Any new sequence that was only found in the second iteration is highlighted in yellow. And we see some new sequences with very, very strong e values, and I want to caution you about that because the nature of the second iteration's very different from the nature of the first iteration. In the first iteration, we started with the, transposase sequence for pyrococcus. And, when I ran that first iteration, I said that anything with an e value of less than 1 times 10 to the minus 4 is a homolog. So, now I see e values much, much less than 1 times 10 to the minus 4, but let's remember, in the second iteration. We're not running the query sequence from pyrococcus. We're now running a possum that represents both the pyrococcus sequence and a bunch of other transposes. So these e values say that the possum was matched very, very well, which suggests that it's probably a member of that protein family. However, in some cases, a possum might not quite accurately represent a protein family, maybe because a nonrelative got into that mix, and it's usually best to treat these new yellow sequences as punitive homologs, meaning that they should be further checked for homology. And you notice, in this case, that almost all of the sequences highlighted in yellow say transposase, which tells me that they are probably indeed members of that protein family. So we can see that PsiBlast is a powerful tool for finding potential homologs of protein or DNA sequences that might not be found by the first iteration of blast. Blast is not 100% sensitive, so it will not necessarily find all homologs in an initial database search. Psi-blast is a powerful way to identify potential homologs that may have been missed by the first round of a blast search. This concludes our PsiBlast tutorial. For more information, please contact us in the Center for Biolotechnology Information. Thank you. [MUSIC]