1
00:00:00,000 --> 00:00:08,784
[MUSIC]

2
00:00:08,784 --> 00:00:13,736
This exercise is going to be on the
PsiBlast tool at NCBI.

3
00:00:13,736 --> 00:00:16,667
PsiBlast stands for position specific
iterated

4
00:00:16,667 --> 00:00:18,739
blast, which is a specific form of

5
00:00:18,739 --> 00:00:23,830
blast, designed to find distant homologs
of a particular protein or DNA sequence.

6
00:00:25,330 --> 00:00:30,716
Our objectives for this exercise include
setting up the parameters of PsiBlast,

7
00:00:30,716 --> 00:00:35,800
specifically to get the right threshold
for future iterations of PsiBlast.

8
00:00:37,770 --> 00:00:40,880
We will enter a protein sequence to run
the first

9
00:00:40,880 --> 00:00:44,924
iteration of PsiBlast, and then from those
results, we will

10
00:00:44,924 --> 00:00:47,800
be able to compile a group of sequences to
run

11
00:00:47,800 --> 00:00:52,120
a second iteration to find other members
of that protein family.

12
00:00:53,220 --> 00:00:55,607
Finally, we will be looking in more detail
how to

13
00:00:55,607 --> 00:00:58,590
interpret those results, which can be a
little bit tricky.

14
00:01:00,810 --> 00:01:05,292
First, we will retrieve the protein
sequence that we were going

15
00:01:05,292 --> 00:01:10,090
to use for PsiBlast, and that sequence is
available at NCBI.

16
00:01:10,090 --> 00:01:14,280
Specifically, we'll be using a transposase
from Pyrococcus furiosus.

17
00:01:14,280 --> 00:01:20,280
Pyrococcus furiosus is a hyperthermophile
that grows at 100 degrees Celsius.

18
00:01:20,280 --> 00:01:22,850
It comes from a volcanic vent.

19
00:01:22,850 --> 00:01:25,560
The organism was first isolated off the
coast of

20
00:01:25,560 --> 00:01:28,010
Italy from a volcanic vent on the
Mediterranean floor.

21
00:01:28,010 --> 00:01:36,600
Now I'm going to click FASTA to get the
actual sequence, and here is the sequence.

22
00:01:36,600 --> 00:01:38,758
What you might notice is that the first

23
00:01:38,758 --> 00:01:43,080
line is descriptive, describes the
transposes for Pyrococcus furiosus.

24
00:01:43,080 --> 00:01:46,064
And then below that are a series of
letters,

25
00:01:46,064 --> 00:01:50,270
which represent the amino acids in that
particular protein.

26
00:01:52,630 --> 00:01:53,420
Okay.

27
00:01:53,420 --> 00:01:57,320
You might notice that the top line, which
I have highlighted, is the description.

28
00:01:57,320 --> 00:01:59,428
The description says that this is a

29
00:01:59,428 --> 00:02:03,136
transposase from Pyrococcus furiosus, and
then below that

30
00:02:03,136 --> 00:02:05,751
are a series of letters, and these letters

31
00:02:05,751 --> 00:02:09,045
represent the amino acids in a particular
sequence.

32
00:02:09,045 --> 00:02:13,097
Now the the FASTA formatted sequence
includes both the amino acids

33
00:02:13,097 --> 00:02:17,230
sequence and the descriptions, so I'm
going to copy and paste that.

34
00:02:18,630 --> 00:02:21,170
And now I'm going to go and do the
PsiBlast search.

35
00:02:21,170 --> 00:02:26,206
To do that, I've got to go to the blast
page at NCBI,

36
00:02:26,206 --> 00:02:31,242
and you can Google that or type in NCBI,
for National

37
00:02:31,242 --> 00:02:36,392
Center for Biotechnology Information,
.NLM, for

38
00:02:36,392 --> 00:02:41,543
National Library of Medicine, .NLH, for
National

39
00:02:41,543 --> 00:02:46,710
Institutes of Health, .gov, and then slash
last.

40
00:02:49,220 --> 00:02:54,320
That will take me to the main blast page,
and I'm going to click protein blast.

41
00:02:54,320 --> 00:02:58,670
And if you notice, by protein blast,
you'll see a series of algorithms.

42
00:02:58,670 --> 00:03:04,330
The algorithms include, blast p, psi
blast, phi blast, and delta blast.

43
00:03:04,330 --> 00:03:09,360
So, we're going to click protein blast,
and the first thing I'm

44
00:03:09,360 --> 00:03:13,470
going to do is paste in the sequence into
the window.

45
00:03:14,780 --> 00:03:19,180
So there is my, the transposee is from
pyrococcus furiosus.

46
00:03:23,580 --> 00:03:23,760
Okay.

47
00:03:23,760 --> 00:03:24,964
Now I'm going to select PsiBlast.

48
00:03:24,964 --> 00:03:28,137
[BLANK_AUDIO]

49
00:03:28,137 --> 00:03:29,611
And I'm going to choose my database.

50
00:03:29,611 --> 00:03:33,914
My database, in this case, is going to be
Reference Proteins.

51
00:03:33,914 --> 00:03:37,470
It's a more high quality database than the
NR database.

52
00:03:37,470 --> 00:03:39,642
The NR database stands for non-redundant,

53
00:03:39,642 --> 00:03:41,880
but there are plenty of redundant
sequences.

54
00:03:41,880 --> 00:03:45,310
It's a huge database with a lot of low
quality sequences.

55
00:03:45,310 --> 00:03:49,706
We'll switch to the reference proteins,
and one thing I want to do

56
00:03:49,706 --> 00:03:52,927
before actually running the PsiBlast is
you'll

57
00:03:52,927 --> 00:03:56,798
notice algorithm parameters in the
lower-left corner.

58
00:03:56,798 --> 00:04:01,970
We're going to expand that window, and I'm
going to change the threshold.

59
00:04:01,970 --> 00:04:03,498
Scroll all the way down.

60
00:04:03,498 --> 00:04:06,790
You'll see the default PsiBlast threshold
is 0.005.

61
00:04:06,790 --> 00:04:08,630
I want to be a little more stringent than
that.

62
00:04:08,630 --> 00:04:12,648
I'm going to use 1 times 10 to the minus
4th to make sure

63
00:04:12,648 --> 00:04:14,944
that I include only members of the

64
00:04:14,944 --> 00:04:19,310
transposes family when I do the second
iteration.

65
00:04:19,310 --> 00:04:22,510
So I'm going to change that to 0.0001.

66
00:04:22,510 --> 00:04:25,910
Okay, so I have my protein sequence.

67
00:04:25,910 --> 00:04:27,069
I have my database.

68
00:04:28,500 --> 00:04:31,300
I selected PsiBlast, and I've changed my
threshold.

69
00:04:31,300 --> 00:04:37,332
One thing I think I want to do is limit to
a particular type of organism, so the goal

70
00:04:37,332 --> 00:04:40,574
of our exercise might be to find homologs
or

71
00:04:40,574 --> 00:04:45,190
relatives of this protein in a particular
type of bacteria.

72
00:04:46,200 --> 00:04:56,200
So I'm going to enter Lacto Bacillus,
which will give me lactic acid bacteria.

73
00:04:56,200 --> 00:05:00,065
Many different types, and you see that
that's

74
00:05:00,065 --> 00:05:04,230
taxonomic ID 1578, so I'm going to click
that.

75
00:05:10,530 --> 00:05:11,586
And now the results should come up
shortly.

76
00:05:11,586 --> 00:05:17,295
[BLANK_AUDIO]

77
00:05:17,295 --> 00:05:22,040
Okay, so here are the results, and you
notice

78
00:05:22,040 --> 00:05:27,030
that there are a series of alignments
better than

79
00:05:27,030 --> 00:05:32,386
threshold, and there are a whole bunch of
relatives

80
00:05:32,386 --> 00:05:37,519
of the pirachocas transposes in
Lactobacillus.

81
00:05:37,519 --> 00:05:44,980
And then you also see a series of proteins
that scored worse than the threshold.

82
00:05:44,980 --> 00:05:49,414
Now typically for blast, if a, if an e
value, is less than 1 times

83
00:05:49,414 --> 00:05:54,434
10 to the minus 4th, such as this one, 6
times 10 to the negative 5.

84
00:05:54,434 --> 00:05:59,090
We can consider that a homolog or a
relative of the protein that we entered.

85
00:05:59,090 --> 00:06:02,400
So everything above this line we can very,
very

86
00:06:02,400 --> 00:06:07,880
confidently say, yes, these are homologs
of that pyrococcus transposase.

87
00:06:09,400 --> 00:06:11,169
Anything below the line we're not quite as
sure of.

88
00:06:11,169 --> 00:06:15,678
[BLANK_AUDIO]

89
00:06:15,678 --> 00:06:19,359
Okay, but we see that there are some
transposases there.

90
00:06:19,359 --> 00:06:21,850
So what we can do is examine the results.

91
00:06:21,850 --> 00:06:25,020
Notice that almost all of these results
are from transposases.

92
00:06:26,440 --> 00:06:30,109
Even some of the ones that are worse than
threshold are transposases, but

93
00:06:30,109 --> 00:06:34,329
notice that everything above the line,
every protein above the line, has a check.

94
00:06:35,630 --> 00:06:40,502
On the right side in that check box, and
if you notice the header under

95
00:06:40,502 --> 00:06:43,358
worse for threshold says select for side

96
00:06:43,358 --> 00:06:47,230
list, anything below that threshold is not
checked.

97
00:06:49,360 --> 00:06:54,642
So what PsiBlast is about to do in the
second iteration is develop a

98
00:06:54,642 --> 00:07:00,320
position specific scoring matrix or possum
as we commonly call them.

99
00:07:01,370 --> 00:07:05,461
The possum is a mathematical
representation of the sequence that we

100
00:07:05,461 --> 00:07:09,553
started with and all of these relatives
that are above threshhold,

101
00:07:09,553 --> 00:07:12,780
and if I wanted to manually exclude a
protein, I say,

102
00:07:12,780 --> 00:07:17,300
oh, yeah, maybe that's not a relative, I
could do that.

103
00:07:17,300 --> 00:07:19,574
I could also take a sequence that's worse
than threshold.

104
00:07:19,574 --> 00:07:21,830
I'll actually take the second one here.

105
00:07:21,830 --> 00:07:22,767
It's a transposase.

106
00:07:22,767 --> 00:07:27,430
Its e value is 1 times 10 to the minus 4,
probably right on the border.

107
00:07:27,430 --> 00:07:32,547
So maybe I'll include that in the
PsiBlast, and now when I hit go, I'm going

108
00:07:32,547 --> 00:07:34,829
to run a second version of blast, a

109
00:07:34,829 --> 00:07:38,479
second iteration, except not with a query
sequence.

110
00:07:38,479 --> 00:07:42,653
Now with this possum, that mathematically
represents all the

111
00:07:42,653 --> 00:07:45,890
of the results in the first iteration of
blasts.

112
00:07:54,110 --> 00:07:55,650
Okay, and these results should come
shortly.

113
00:08:04,200 --> 00:08:08,710
Okay, now what we notice is, first of all,
a bunch of very, very strong e values.

114
00:08:08,710 --> 00:08:10,454
2 times 10 to the minus 122.

115
00:08:10,454 --> 00:08:13,730
I'll talk more about that in a second.

116
00:08:13,730 --> 00:08:17,351
You'll also notice that some protein
sequences are

117
00:08:17,351 --> 00:08:20,890
highlighted in yellow, and the sequences
that are

118
00:08:20,890 --> 00:08:24,922
in white, without the yellow highlighting,
are protein

119
00:08:24,922 --> 00:08:28,480
sequences that were found in the first
iteration.

120
00:08:30,140 --> 00:08:32,480
Any new sequence that was only found

121
00:08:32,480 --> 00:08:35,440
in the second iteration is highlighted in
yellow.

122
00:08:36,530 --> 00:08:38,731
And we see some new sequences with very,

123
00:08:38,731 --> 00:08:41,191
very strong e values, and I want to
caution you

124
00:08:41,191 --> 00:08:44,301
about that because the nature of the
second iteration's

125
00:08:44,301 --> 00:08:48,220
very different from the nature of the
first iteration.

126
00:08:48,220 --> 00:08:54,090
In the first iteration, we started with
the, transposase sequence for pyrococcus.

127
00:08:54,090 --> 00:08:57,624
And, when I ran that first iteration, I
said that anything with an e

128
00:08:57,624 --> 00:09:00,860
value of less than 1 times 10 to the minus
4 is a homolog.

129
00:09:00,860 --> 00:09:04,472
So, now I see e values much, much less
than 1 times

130
00:09:04,472 --> 00:09:10,280
10 to the minus 4, but let's remember, in
the second iteration.

131
00:09:10,280 --> 00:09:13,298
We're not running the query sequence from
pyrococcus.

132
00:09:13,298 --> 00:09:15,992
We're now running a possum that represents
both

133
00:09:15,992 --> 00:09:19,930
the pyrococcus sequence and a bunch of
other transposes.

134
00:09:19,930 --> 00:09:23,463
So these e values say that the possum was
matched very, very

135
00:09:23,463 --> 00:09:28,000
well, which suggests that it's probably a
member of that protein family.

136
00:09:29,860 --> 00:09:35,289
However, in some cases, a possum might not
quite accurately represent

137
00:09:35,289 --> 00:09:40,273
a protein family, maybe because a
nonrelative got into that mix,

138
00:09:40,273 --> 00:09:45,524
and it's usually best to treat these new
yellow sequences as punitive

139
00:09:45,524 --> 00:09:50,870
homologs, meaning that they should be
further checked for homology.

140
00:09:50,870 --> 00:09:55,379
And you notice, in this case, that almost
all of the sequences highlighted in yellow

141
00:09:55,379 --> 00:09:57,633
say transposase, which tells me that they

142
00:09:57,633 --> 00:10:00,470
are probably indeed members of that
protein family.

143
00:10:01,760 --> 00:10:07,933
So we can see that PsiBlast is a powerful
tool for finding potential homologs of

144
00:10:07,933 --> 00:10:14,310
protein or DNA sequences that might not be
found by the first iteration of blast.

145
00:10:15,540 --> 00:10:18,564
Blast is not 100% sensitive, so it will
not

146
00:10:18,564 --> 00:10:22,605
necessarily find all homologs in an
initial database search.

147
00:10:22,605 --> 00:10:27,540
Psi-blast is a powerful way to identify
potential homologs that may

148
00:10:27,540 --> 00:10:31,380
have been missed by the first round of a
blast search.

149
00:10:32,510 --> 00:10:35,972
This concludes our PsiBlast tutorial.

150
00:10:35,972 --> 00:10:37,127
For more information, please contact us

151
00:10:37,127 --> 00:10:38,516
in the Center for Biolotechnology
Information.

152
00:10:38,516 --> 00:10:38,566
Thank you.

153
00:10:38,566 --> 00:10:48,566
[MUSIC]

