1
00:00:00,008 --> 00:00:00,715
Hello.

2
00:00:00,715 --> 00:00:04,280
I'm Bob Leske of Johns Hopkins.

3
00:00:04,280 --> 00:00:07,550
We will examine the bacterial genome in
this short lecture.

4
00:00:07,550 --> 00:00:12,862
Be sure to focus on the polycistronic
nature of some transcripts.

5
00:00:12,862 --> 00:00:18,740
Unlike eukaryotes, most bacteria have a
single chromosome.

6
00:00:18,740 --> 00:00:22,030
And in most but not all cases, that
chromosome is circular.

7
00:00:24,590 --> 00:00:28,120
Polycistronic is an important term,
meaning that there can be

8
00:00:28,120 --> 00:00:31,280
more than one protein coding reagent on an
mRNA transcript.

9
00:00:34,950 --> 00:00:37,660
Examining that transcript more closely.

10
00:00:37,660 --> 00:00:42,190
You will find two untranslated regions
that are not protein coated.

11
00:00:42,190 --> 00:00:43,670
Those come at either end of the sequence.

12
00:00:47,310 --> 00:00:52,020
We will look at the lactose operon, made
famous by Jacob and Monod in the 1960s.

13
00:00:54,460 --> 00:00:57,200
And then to bring it back to
bioinformatics, we will look to

14
00:00:57,200 --> 00:01:01,715
identify operons on an NCBI sequence
record and on a gene prediction.

15
00:01:01,715 --> 00:01:08,780
[bbb] First, let's be aware that bacteria

16
00:01:08,780 --> 00:01:11,520
tend to have much smaller genomes than
eukaryotes.

17
00:01:13,390 --> 00:01:16,460
Escherichia coli or what we call E.

18
00:01:16,460 --> 00:01:19,061
coli has a 4.6 megabase genome.

19
00:01:19,061 --> 00:01:20,860
That's 4.6 million base pairs.

20
00:01:25,070 --> 00:01:31,450
One of the simplest eukaryotes budding
yeast has a genome of 12.1 megabases.

21
00:01:31,450 --> 00:01:33,470
Almost three times the size of E coli.

22
00:01:35,190 --> 00:01:37,610
And that's a very small eukaryotic genome.

23
00:01:40,670 --> 00:01:44,140
The human genome is 3.2 gigabases.

24
00:01:44,140 --> 00:01:46,140
That's 3.3 billion based pairs.

25
00:01:46,140 --> 00:01:48,920
And that's not even the largest genome.

26
00:01:48,920 --> 00:01:51,480
Some plants are even bigger in genome size
than humans.

27
00:01:54,260 --> 00:01:57,063
Bacteria usually had a single circular
chromosome.

28
00:01:57,063 --> 00:01:58,240
There are exceptions.

29
00:01:59,480 --> 00:02:03,540
Many also have extrachromosomal plasmids.

30
00:02:03,540 --> 00:02:06,660
Those are small, usually circular pieces
of additional DNA.

31
00:02:11,160 --> 00:02:13,990
Here's a site map of the E coli genome.

32
00:02:13,990 --> 00:02:15,740
Look at the top where you see a zero.

33
00:02:15,740 --> 00:02:21,660
The 0 position represents what's called
the origin of replication.

34
00:02:23,680 --> 00:02:26,166
Then the numbers increase clockwise and
there

35
00:02:26,166 --> 00:02:28,530
are about 4 million 600 thousand or so.

36
00:02:28,530 --> 00:02:29,030
Now

37
00:02:32,560 --> 00:02:35,550
remember, there are still two strands of
DNA.

38
00:02:35,550 --> 00:02:40,090
What we call the plus strand is arranged
in this diagram, clockwise.

39
00:02:42,720 --> 00:02:44,850
However, there are genes on the minus or

40
00:02:44,850 --> 00:02:48,060
complimentary strand and they run in the
other direction.

41
00:02:52,810 --> 00:02:54,780
So, I've mentioned polycistronic.

42
00:02:54,780 --> 00:02:57,720
In eukaryotes, the final splice mRNA

43
00:02:57,720 --> 00:02:59,900
usually only has one protein coding
region.

44
00:03:01,440 --> 00:03:05,720
That means that only one protein is
derived from the mRNA.

45
00:03:05,720 --> 00:03:08,250
Actually many copies of that same single
protein.

46
00:03:11,430 --> 00:03:15,250
In bacteria, while mRNA splicing usually
does not occur.

47
00:03:15,250 --> 00:03:17,840
There are often, but not always, more than

48
00:03:17,840 --> 00:03:20,390
one protein coding region in a prokaryotic
mRNA.

49
00:03:23,320 --> 00:03:27,940
If say, an mRNA has five protein coding
regions, then

50
00:03:27,940 --> 00:03:31,720
translation makes many molecules of each
of those five proteins.

51
00:03:34,620 --> 00:03:36,250
And here's what it looks like.

52
00:03:36,250 --> 00:03:40,140
The diagram's a bit fuzzy but it displays
the concept well.

53
00:03:40,140 --> 00:03:43,100
Prokaryote is at the top and the eukaryote
is at the bottom.

54
00:03:44,820 --> 00:03:47,410
Prokaryotes tend to do it this way for
efficiency.

55
00:03:52,660 --> 00:03:54,540
Here is a similar diagram.

56
00:03:54,540 --> 00:03:58,610
But you can see the untranslated regions
or UTRs at either end.

57
00:04:01,830 --> 00:04:05,122
The one at the left of the mRNA is called
the 5' UTR.

58
00:04:05,122 --> 00:04:08,880
The 5' refers to the chemistry of which

59
00:04:08,880 --> 00:04:11,369
carbon that phosphate attaches to on the
sugar.

60
00:04:12,540 --> 00:04:14,440
Or to put it in another way, protein

61
00:04:14,440 --> 00:04:17,850
sequences go from N terminus to C
terminus.

62
00:04:17,850 --> 00:04:24,702
DNA and RNA sequences go from 5' to 3',
and both involve chemistry.

63
00:04:24,702 --> 00:04:30,590
The 5' UTR is everything to the left of
the first AUG start codon.

64
00:04:33,670 --> 00:04:35,420
The 3' UTR is on the right.

65
00:04:37,820 --> 00:04:41,330
It's everything that is untranslated after
the stop codon.

66
00:04:42,710 --> 00:04:46,644
Both UTRs are found in eukaryotes as well.

67
00:04:46,644 --> 00:04:51,780
So, from left to right on a bacterial
mRNA, you have a 5' UTR.

68
00:04:53,110 --> 00:04:56,430
One or more codon regions then a 3' UTR.

69
00:05:00,420 --> 00:05:03,730
Here is the well-studied lactose operon.

70
00:05:03,730 --> 00:05:09,160
Bacteria prefer glucose as an energy
source, but if there

71
00:05:09,160 --> 00:05:11,380
is no glucose to be found they'll deal
with lactose.

72
00:05:12,700 --> 00:05:16,590
So they have to find away to turn on the
genes involved in lactose synthesis.

73
00:05:18,110 --> 00:05:21,350
This describes some complex regulation and
shows what's called the

74
00:05:21,350 --> 00:05:24,920
regulatory region of the gene, or in this
case the operon.

75
00:05:26,240 --> 00:05:29,040
That's the term for the multiple genes on
a single mRNA.

76
00:05:32,140 --> 00:05:34,732
Look primarily at the numbers below the
first graphic.

77
00:05:34,732 --> 00:05:39,130
The +1 position is called the
transcription start site.

78
00:05:40,190 --> 00:05:42,450
That is no the start codon.

79
00:05:42,450 --> 00:05:46,490
It is the start of the mRNA, transcription
not translation.

80
00:05:48,070 --> 00:05:50,576
The mRNA is to the right of that +1
position.

81
00:05:51,790 --> 00:05:53,658
Or think of it this way.

82
00:05:53,658 --> 00:05:58,500
+1 refers to the beginning of the 5' UTR
of the mRNA.

83
00:06:01,040 --> 00:06:04,060
The region to the left of the +1 is called
the regulatory region.

84
00:06:05,592 --> 00:06:09,180
Here is where proteins bind that help
control the level of transcription.

85
00:06:10,760 --> 00:06:14,170
A need for lactose digestion ultimately
determines whether if

86
00:06:14,170 --> 00:06:17,010
the transcription of the lactose enzymes
is on or off.

87
00:06:17,010 --> 00:06:24,120
The direction to the left, or to the 5'
side, is frequently called upstream.

88
00:06:25,730 --> 00:06:30,719
The minus 35 position, is 35 bases
upstream of the transcription start site.

89
00:06:33,260 --> 00:06:37,440
And to the right or to the 3' side is
called downstream.

90
00:06:42,470 --> 00:06:47,460
In prokaryotes, you can have a single mRNA
that contains more than one coding region.

91
00:06:47,460 --> 00:06:51,029
That would be very rare in eukaryotes, but
very common in bacteria and archaea.

92
00:06:54,090 --> 00:06:58,410
This operon goes from position 70 to 6338.

93
00:06:58,410 --> 00:06:58,910
The

94
00:07:01,030 --> 00:07:04,820
mRNA would be larger since it would
include the untranslated regions of the

95
00:07:04,820 --> 00:07:08,840
mRNA that come before the start codon and
after the last dot codon.

96
00:07:11,050 --> 00:07:14,680
If you look at the annotation, you will
see that this operon has four CDS regions.

97
00:07:14,680 --> 00:07:22,340
[bbb].

98
00:07:22,340 --> 00:07:25,280
Here is an output for
fgenesb@softberry.com.

99
00:07:25,280 --> 00:07:29,639
This is explained in a separate video on
bacterial gene prediction.

100
00:07:31,740 --> 00:07:34,350
Note that there are four transcription
units, meaning that

101
00:07:34,350 --> 00:07:37,220
four mRNAs come out of this region of
genomic DNA.

102
00:07:38,960 --> 00:07:42,750
Of those four transcription unit, two are
operons, which

103
00:07:42,750 --> 00:07:45,150
each have more than one CDS on that mRNA.

104
00:07:47,930 --> 00:07:49,150
The first mRNA.

105
00:07:50,350 --> 00:07:55,180
Has three CDS regions, Op 1, Op 2, and Op
3.

106
00:07:59,650 --> 00:08:04,520
Next are two single CDS transcription
units, Tu 1and Tu 1.

107
00:08:05,760 --> 00:08:07,620
Why not Tu 2?

108
00:08:07,620 --> 00:08:11,670
That would imply a second CDS on an mRNA,
that's how that numbering works.

109
00:08:14,280 --> 00:08:17,460
The last mRNA has 2 CDS regions, Op 1 and
Op 2.

110
00:08:17,460 --> 00:08:20,880
Hopefully that explains the notation.

111
00:08:20,880 --> 00:08:21,380
To

112
00:08:27,560 --> 00:08:30,850
summarize, bacteria usually has circular
genomes.

113
00:08:30,850 --> 00:08:33,620
Eukaryotes tend to have linear chromosomes
and more than one.

114
00:08:36,340 --> 00:08:40,810
With prokaryotic genes, you don't have to
deal with the splicing issue, usually.

115
00:08:42,620 --> 00:08:45,470
However, one complication is that multiple
coding

116
00:08:45,470 --> 00:08:48,105
reagents often occur on a bacterial mRNA
transcript.

117
00:08:50,910 --> 00:08:52,210
Finally, why do it this way?

118
00:08:53,320 --> 00:08:55,940
Bacteria waste little space in their
genome.

119
00:08:55,940 --> 00:09:01,300
Efficiency allows for genes in a certain
pathway to be transcribed onto one

120
00:09:01,300 --> 00:09:05,780
mRNA molecule so that all necessary
enzymes can be translated into one shot.

121
00:09:07,540 --> 00:09:08,060
That's it for now.

