What you'll learn
Protein synthesis and gene expression explains how the base sequence of a gene becomes the amino acid sequence of a protein, and how cells control which of their genes are used. It is the topic that connects DNA structure to everything an organism actually does, since proteins are enzymes, receptors, transporters, antibodies and structural components. At CAPE level the detail required extends to the roles of the three RNA types, the features of the genetic code, the stages of transcription and translation, and the operon model of gene regulation. By the end of this topic you should be able to describe transcription and translation in full, explain the features of the genetic code, distinguish the RNA types by structure and role, explain how mutations affect the protein produced, and describe how gene expression is controlled.
Key terms and definitions
Gene — a sequence of DNA bases coding for a polypeptide
Transcription — the synthesis of messenger RNA from a DNA template
Translation — the assembly of a polypeptide at a ribosome according to the messenger RNA sequence
Codon — a triplet of bases on messenger RNA coding for one amino acid
Anticodon — a triplet of bases on transfer RNA complementary to a codon
Template strand — the DNA strand that is transcribed, also called the antisense strand
RNA polymerase — the enzyme that synthesises messenger RNA from the DNA template
Degenerate code — a code in which most amino acids are specified by more than one codon
Intron — a non-coding sequence within a gene, removed from the primary transcript
Exon — a coding sequence retained in the mature messenger RNA
Splicing — the removal of introns and joining of exons
Operon — a group of genes transcribed together under the control of a single promoter and operator
Frameshift — a shift in the reading frame caused by an insertion or deletion
Core concepts
The three types of RNA
Messenger RNA is a single-stranded copy of a gene, synthesised in the nucleus and carrying the code to the ribosome. It is comparatively short-lived, which allows the cell to stop producing a protein by ceasing transcription.
Transfer RNA is a small single strand folded into a clover-leaf shape held by hydrogen bonding, with an amino acid binding site at one end and an anticodon of three bases at the other. Each transfer RNA carries one specific amino acid, matched to its anticodon.
Ribosomal RNA combines with protein to form the ribosome itself. The ribosome has a small subunit that binds messenger RNA and a large subunit with two sites that hold transfer RNA molecules during translation, and it catalyses peptide bond formation.
The genetic code
The code is a triplet code: each sequence of three bases on the messenger RNA, called a codon, specifies one amino acid. Three bases are needed because there are twenty amino acids and only four bases, so pairs would give only sixteen combinations while triplets give sixty-four.
The code is degenerate, meaning most amino acids are specified by more than one codon. This is biologically significant because it means many substitution mutations have no effect on the protein produced.
The code is non-overlapping, so each base is part of only one codon, and it is read in one direction from a fixed starting point.
The code is universal, or very nearly so, meaning the same codons specify the same amino acids in almost all organisms. This is what makes genetic engineering across species possible.
Of the sixty-four codons, one acts as a start codon and also codes for methionine, and three are stop codons that specify no amino acid and terminate translation.
Transcription
Transcription occurs in the nucleus and produces messenger RNA from a DNA template.
DNA helicase or the action of RNA polymerase itself breaks the hydrogen bonds between the bases in the region of the gene, unwinding the double helix and exposing the bases.
Only one of the two strands is transcribed: the template strand. The other, the coding strand, has the same sequence as the messenger RNA except that thymine replaces uracil.
Free RNA nucleotides align opposite their complementary bases on the template strand. Adenine on the template pairs with uracil, not thymine, since RNA contains no thymine. Cytosine pairs with guanine as usual.
RNA polymerase joins adjacent RNA nucleotides by forming phosphodiester bonds, building the messenger RNA strand in one direction.
When RNA polymerase reaches a stop signal it detaches, the messenger RNA is released, and the DNA rewinds.
In eukaryotes the initial product is a primary transcript containing both introns and exons. Splicing removes the introns and joins the exons to produce mature messenger RNA, which then passes out through a nuclear pore to a ribosome. Prokaryotes have no introns and no splicing, and because they have no nuclear envelope, translation can begin before transcription is complete.
Translation
Translation occurs at a ribosome in the cytoplasm or on the rough endoplasmic reticulum.
The messenger RNA binds to the small subunit of the ribosome, and the ribosome moves along until it reaches the start codon.
A transfer RNA with an anticodon complementary to the first codon binds, bringing its specific amino acid. Hydrogen bonds form between the codon and anticodon.
A second transfer RNA binds at the adjacent site, bringing the next amino acid. The two amino acids are held close together, and the ribosome catalyses the formation of a peptide bond between them, using ATP.
The first transfer RNA is released and can collect another molecule of its amino acid. The ribosome moves along the messenger RNA by one codon, and the process repeats.
Synthesis continues until a stop codon is reached. No transfer RNA has a complementary anticodon, so no further amino acid is added, the polypeptide is released and the ribosome separates from the messenger RNA.
Several ribosomes may translate the same messenger RNA molecule simultaneously, forming a polysome, which allows many copies of a protein to be made quickly.
The released polypeptide then folds into its secondary and tertiary structure, and may be modified in the Golgi apparatus before becoming functional.
From polypeptide to functional protein
The sequence of bases in the gene determines the sequence of amino acids in the polypeptide. That primary structure determines how the chain folds, because it determines where the R groups lie and therefore where hydrogen bonds, ionic bonds, hydrophobic interactions and disulfide bridges can form.
The resulting tertiary structure determines the shape of the molecule and therefore its function — the shape of an enzyme's active site, of a receptor's binding site, or of an antibody's variable region.
This chain of causation, from gene to base sequence to amino acid sequence to folding to shape to function, is the central explanatory sequence of the topic and is worth being able to state in full.
Mutations and their effects
A substitution replaces one base with another. Because the code is degenerate, the new codon may specify the same amino acid, in which case the mutation is silent and the protein is unchanged. If a different amino acid is specified, the effect depends on where it lies: a change far from the active site may have little effect, while a change within the active site may destroy function. If a stop codon is created, a truncated and usually non-functional protein results.
An insertion or deletion adds or removes a base, shifting the reading frame from that point onwards. Every subsequent codon is altered, so almost the entire remaining amino acid sequence is wrong and the protein is normally non-functional. Frameshift mutations are therefore usually far more damaging than substitutions.
Sickle cell anaemia is the standard example of a substitution with major consequences. A single base change in the gene for the beta chain of haemoglobin alters one amino acid, changing the tertiary structure so that haemoglobin molecules stick together at low oxygen concentrations and distort the red blood cell. It is a particularly relevant example in the Caribbean, where the allele is present at appreciable frequency because heterozygotes have some protection against malaria.
Control of gene expression
Every cell of an organism contains the same genes, yet cells differ because different genes are expressed. Controlling which proteins are made, and when, is therefore fundamental.
The lac operon in the bacterium Escherichia coli is the model studied, and it illustrates the principle economically.
The operon consists of a promoter, where RNA polymerase binds; an operator, a control sequence adjacent to it; and structural genes coding for the enzymes needed to metabolise lactose.
A separate regulator gene produces a repressor protein. In the absence of lactose, the repressor binds to the operator, physically blocking RNA polymerase from moving along the DNA. The structural genes are not transcribed and the enzymes are not made.
When lactose is present, it binds to the repressor protein at a site other than its DNA-binding site. This changes the repressor's tertiary structure so that it can no longer bind the operator. RNA polymerase is then free to transcribe the structural genes and the enzymes are produced.
The system means the bacterium makes lactose-metabolising enzymes only when lactose is available, avoiding the waste of synthesising proteins it cannot use. This is an inducible operon, since the substrate induces its own enzymes.
In eukaryotes control is more complex, involving transcription factors that bind to DNA and either promote or inhibit the binding of RNA polymerase, and hormones such as steroids that enter the cell and act as or activate transcription factors.
Worked examples
Example 1: Deriving a sequence (5 marks)
A section of the template strand of DNA reads TAC GGA TTC AAG. State the messenger RNA sequence, the anticodons of the transfer RNA molecules required, and the number of amino acids in the resulting polypeptide.
The messenger RNA is complementary to the template, with uracil replacing thymine. Adenine on the template gives uracil, thymine gives adenine, guanine gives cytosine and cytosine gives guanine. The messenger RNA is therefore AUG CCU AAG UUC.
The anticodons on transfer RNA are complementary to the messenger RNA codons, so they are UAC GGA UUC AAG. Note that these correspond to the original template sequence with uracil replacing thymine, which is a useful check.
There are four codons, and the first, AUG, is the start codon which also codes for methionine. The polypeptide therefore contains four amino acids, assuming none of the codons is a stop codon.
Example 2: Comparing mutation types (5 marks)
Explain why a deletion of a single base usually has a more severe effect on a protein than a substitution of a single base.
A substitution replaces one base with another, so only one codon is altered. Because the genetic code is degenerate, the new codon may still specify the same amino acid, in which case the protein is unchanged. Even if a different amino acid is specified, only one amino acid in the whole chain differs, and if it lies away from the active site or binding site the protein may still function.
A deletion removes a base, so every subsequent codon is shifted by one position. This causes a frameshift, and all the codons from the point of deletion onwards are read incorrectly. Almost the entire remaining amino acid sequence is therefore wrong, the polypeptide folds differently, and the protein is normally non-functional. A premature stop codon may also be created, truncating the protein.
Example 3: Explaining the lac operon (5 marks)
Explain how the presence of lactose leads to the production of enzymes that metabolise it.
In the absence of lactose, the regulator gene produces a repressor protein which binds to the operator sequence. This prevents RNA polymerase from binding to the promoter and moving along the structural genes, so no messenger RNA is transcribed and the enzymes are not synthesised.
When lactose is present, it binds to the repressor protein at a site away from its DNA-binding region. This alters the tertiary structure of the repressor so that its DNA-binding site changes shape and it can no longer bind to the operator.
With the operator free, RNA polymerase can bind to the promoter and transcribe the structural genes into messenger RNA. The messenger RNA is translated at ribosomes to produce the enzymes needed to metabolise lactose.
This ensures the enzymes are produced only when their substrate is present, so the cell does not waste amino acids and ATP making proteins it cannot use.
Common mistakes and how to avoid them
The most frequent error is confusing codons with anticodons, or deriving the anticodon from the DNA rather than from the messenger RNA. The anticodon is complementary to the codon.
Students often write thymine into an RNA sequence. RNA contains uracil, and this single slip can invalidate an entire derived sequence.
Another common slip is stating that both DNA strands are transcribed. Only the template strand is.
Many candidates describe translation without mentioning that peptide bond formation requires ATP, or without stating that the ribosome moves along one codon at a time.
In operon questions, answers frequently say that lactose binds to the operator. It binds to the repressor protein, changing its shape so that it cannot bind the operator.
Finally, candidates often assert that a substitution always changes the protein. Degeneracy means many substitutions are silent.
Exam technique for "Protein synthesis and gene expression"
When deriving sequences, write the DNA template, the messenger RNA and the anticodons in three aligned rows. Errors become visible immediately and the working earns marks.
Always check that no thymine appears in an RNA sequence before moving on.
Name the enzyme and the bond in every stage: RNA polymerase forming phosphodiester bonds in transcription, and the ribosome catalysing peptide bonds in translation.
For mutation questions, always mention degeneracy when discussing substitutions and the reading frame when discussing insertions and deletions.
For the operon, describe the state of the system both with and without lactose. Answers covering only one state cannot score full marks.
Quick revision summary
Messenger RNA carries the code from nucleus to ribosome, transfer RNA brings specific amino acids matched by its anticodon, and ribosomal RNA forms the ribosome. The genetic code is a triplet, degenerate, non-overlapping and near-universal code, with one start and three stop codons. In transcription, the helix unwinds, free RNA nucleotides pair with the template strand with uracil replacing thymine, and RNA polymerase forms phosphodiester bonds; in eukaryotes introns are then spliced out. In translation, the ribosome binds messenger RNA, transfer RNA anticodons pair with codons, peptide bonds form using ATP, the ribosome moves codon by codon, and a stop codon releases the polypeptide. The base sequence determines the amino acid sequence, which determines folding, shape and function. Substitutions may be silent because of degeneracy, while insertions and deletions cause frameshifts that alter every subsequent codon. The lac operon shows inducible control: the repressor blocks the operator until lactose binds it and changes its shape, allowing RNA polymerase to transcribe the structural genes only when lactose is present.