Hello. I'm Bob Leske of Johns Hopkins. We will examine the bacterial genome in this short lecture. Be sure to focus on the polycistronic nature of some transcripts. Unlike eukaryotes, most bacteria have a single chromosome. And in most but not all cases, that chromosome is circular. Polycistronic is an important term, meaning that there can be more than one protein coding reagent on an mRNA transcript. Examining that transcript more closely. You will find two untranslated regions that are not protein coated. Those come at either end of the sequence. We will look at the lactose operon, made famous by Jacob and Monod in the 1960s. And then to bring it back to bioinformatics, we will look to identify operons on an NCBI sequence record and on a gene prediction. [bbb] First, let's be aware that bacteria tend to have much smaller genomes than eukaryotes. Escherichia coli or what we call E. coli has a 4.6 megabase genome. That's 4.6 million base pairs. One of the simplest eukaryotes budding yeast has a genome of 12.1 megabases. Almost three times the size of E coli. And that's a very small eukaryotic genome. The human genome is 3.2 gigabases. That's 3.3 billion based pairs. And that's not even the largest genome. Some plants are even bigger in genome size than humans. Bacteria usually had a single circular chromosome. There are exceptions. Many also have extrachromosomal plasmids. Those are small, usually circular pieces of additional DNA. Here's a site map of the E coli genome. Look at the top where you see a zero. The 0 position represents what's called the origin of replication. Then the numbers increase clockwise and there are about 4 million 600 thousand or so. Now remember, there are still two strands of DNA. What we call the plus strand is arranged in this diagram, clockwise. However, there are genes on the minus or complimentary strand and they run in the other direction. So, I've mentioned polycistronic. In eukaryotes, the final splice mRNA usually only has one protein coding region. That means that only one protein is derived from the mRNA. Actually many copies of that same single protein. In bacteria, while mRNA splicing usually does not occur. There are often, but not always, more than one protein coding region in a prokaryotic mRNA. If say, an mRNA has five protein coding regions, then translation makes many molecules of each of those five proteins. And here's what it looks like. The diagram's a bit fuzzy but it displays the concept well. Prokaryote is at the top and the eukaryote is at the bottom. Prokaryotes tend to do it this way for efficiency. Here is a similar diagram. But you can see the untranslated regions or UTRs at either end. The one at the left of the mRNA is called the 5' UTR. The 5' refers to the chemistry of which carbon that phosphate attaches to on the sugar. Or to put it in another way, protein sequences go from N terminus to C terminus. DNA and RNA sequences go from 5' to 3', and both involve chemistry. The 5' UTR is everything to the left of the first AUG start codon. The 3' UTR is on the right. It's everything that is untranslated after the stop codon. Both UTRs are found in eukaryotes as well. So, from left to right on a bacterial mRNA, you have a 5' UTR. One or more codon regions then a 3' UTR. Here is the well-studied lactose operon. Bacteria prefer glucose as an energy source, but if there is no glucose to be found they'll deal with lactose. So they have to find away to turn on the genes involved in lactose synthesis. This describes some complex regulation and shows what's called the regulatory region of the gene, or in this case the operon. That's the term for the multiple genes on a single mRNA. Look primarily at the numbers below the first graphic. The +1 position is called the transcription start site. That is no the start codon. It is the start of the mRNA, transcription not translation. The mRNA is to the right of that +1 position. Or think of it this way. +1 refers to the beginning of the 5' UTR of the mRNA. The region to the left of the +1 is called the regulatory region. Here is where proteins bind that help control the level of transcription. A need for lactose digestion ultimately determines whether if the transcription of the lactose enzymes is on or off. The direction to the left, or to the 5' side, is frequently called upstream. The minus 35 position, is 35 bases upstream of the transcription start site. And to the right or to the 3' side is called downstream. In prokaryotes, you can have a single mRNA that contains more than one coding region. That would be very rare in eukaryotes, but very common in bacteria and archaea. This operon goes from position 70 to 6338. The mRNA would be larger since it would include the untranslated regions of the mRNA that come before the start codon and after the last dot codon. If you look at the annotation, you will see that this operon has four CDS regions. [bbb]. Here is an output for fgenesb@softberry.com. This is explained in a separate video on bacterial gene prediction. Note that there are four transcription units, meaning that four mRNAs come out of this region of genomic DNA. Of those four transcription unit, two are operons, which each have more than one CDS on that mRNA. The first mRNA. Has three CDS regions, Op 1, Op 2, and Op 3. Next are two single CDS transcription units, Tu 1and Tu 1. Why not Tu 2? That would imply a second CDS on an mRNA, that's how that numbering works. The last mRNA has 2 CDS regions, Op 1 and Op 2. Hopefully that explains the notation. To summarize, bacteria usually has circular genomes. Eukaryotes tend to have linear chromosomes and more than one. With prokaryotic genes, you don't have to deal with the splicing issue, usually. However, one complication is that multiple coding reagents often occur on a bacterial mRNA transcript. Finally, why do it this way? Bacteria waste little space in their genome. Efficiency allows for genes in a certain pathway to be transcribed onto one mRNA molecule so that all necessary enzymes can be translated into one shot. That's it for now.