Molecular Basis of Inheritance

How scientists proved that DNA, not protein, carries heredity, and how that DNA is copied, read and switched on and off inside a cell. This is the chapter where genetics becomes chemistry.

The Search for the Genetic Material

Quick answer Three classic experiments moved the title of hereditary material from protein to DNA: Griffith's transformation in mice, Avery's enzyme test, and the Hershey-Chase blender experiment.

Any molecule that acts as the genetic material has to pass four tests. It must be able to make an exact copy of itself. It must be chemically and structurally stable. It must allow slow changes, that is mutations, so that evolution is possible. And it must be able to express itself as the characters an organism actually shows. For a long time proteins were the favourite candidate, because proteins are built from twenty different amino acids and looked complex enough to store information, while DNA was dismissed as a dull repeat of only four bases. Three experiments settled the question in favour of DNA.

Frederick Griffith worked in 1928 with Streptococcus pneumoniae, a bacterium that causes pneumonia. It grows as two strains. The S strain forms smooth, shiny colonies because every cell is wrapped in a polysaccharide capsule, and it is virulent, so injecting it kills mice. The R strain forms rough colonies, has no capsule, and does not kill mice. Griffith found that heat-killed S bacteria on their own were harmless. But when he injected heat-killed S bacteria together with live R bacteria, the mice died, and living S bacteria could be recovered from their bodies. Something released from the dead S cells had permanently changed harmless R cells into capsule-making S cells, and the change was inherited by their offspring. Griffith called this something the transforming principle and the process transformation. He could not say what the chemical was.

Oswald Avery, Colin MacLeod and Maclyn McCarty spent years purifying that chemical from heat-killed S cells. Their method was to destroy one class of molecule at a time and then check whether transformation still happened. Adding proteases, which digest proteins, did not stop transformation. Adding RNase, which digests RNA, did not stop it either. Adding DNase, which digests DNA, stopped transformation completely. The conclusion was direct: DNA is the transforming principle, and therefore DNA is the hereditary material. Many biologists still hesitated, partly because a purified DNA preparation could carry traces of protein along with it.

Alfred Hershey and Martha Chase removed the last doubt in 1952 using bacteriophage T2, a virus that infects Escherichia coli. A phage is made of only two things, a protein coat and DNA inside it, and the two differ in one very useful way. DNA contains phosphorus but no sulphur, while protein contains sulphur in the amino acids cysteine and methionine but no phosphorus. So they grew one batch of phages in a medium containing radioactive phosphorus, which labelled only the phage DNA, and a second batch in a medium containing radioactive sulphur, which labelled only the phage protein. Each batch was allowed to infect bacteria. The cultures were then whirled in a blender, which shook the empty viral coats off the bacterial surface, and spun in a centrifuge, which pulled the heavier bacterial cells to the bottom. Bacteria infected by the phosphorus-labelled phages were radioactive. Bacteria infected by the sulphur-labelled phages were not. Only DNA had entered the bacterial cell, so DNA and not protein carries the instructions for building new viruses.

DNA won the argument for a chemical reason too. RNA has a hydroxyl group on the 2' carbon of every sugar, and that reactive group makes RNA easy to break down. RNA is also catalytic, and a reactive molecule is an unstable one. DNA lacks that 2' hydroxyl, uses thymine in place of uracil, which adds further stability, and is double stranded, so a damaged strand can be repaired using the intact partner as a reference. RNA still serves as the genetic material in some viruses, and such viruses mutate much faster. There is one test, though, on which RNA does better than DNA. RNA can be translated into protein directly, whereas DNA has to be transcribed into RNA first before any protein can be made from it. So RNA expresses itself more readily, while DNA is the safer long-term store. That is the division of labour we actually see in cells: DNA keeps the information and RNA carries and uses it.

S strain vs R strain S forms smooth colonies, has a polysaccharide capsule and is virulent; R forms rough colonies, has no capsule and is non-virulent. Heat kills the S cells but does not destroy their DNA.
Griffith vs Avery-MacLeod-McCarty Griffith proved that transformation happens; Avery's team proved which molecule does it. Do not credit either result to the other.
32P labels DNA, 35S labels protein DNA contains phosphorus but no sulphur; protein contains sulphur (cysteine, methionine) but no phosphorus. Fix the direction of each label before you answer, because the two are easy to swap.
Why DNA is more stable than RNA No reactive 2' hydroxyl group, thymine instead of uracil, and a complementary second strand that allows repair.
Remember
  • Griffith showed that a heat-stable substance from dead S bacteria could transform live R bacteria into S bacteria, but he never identified it.
  • Avery, MacLeod and McCarty showed that DNase alone destroyed the transforming ability, while protease and RNase did not, naming DNA as the transforming principle.
  • Hershey and Chase labelled phage DNA with radioactive phosphorus and phage protein with radioactive sulphur, because DNA has phosphorus but no sulphur and protein the reverse.
  • Radioactivity entered the bacteria only from the phosphorus-labelled phages, proving DNA is the material passed into the host cell.
  • Genetic material must replicate, stay stable, allow slow mutation and express itself; DNA is the better store because it is far more stable and repairable, while RNA expresses itself more directly since it can be translated without being transcribed first.

Structure of DNA and Its Packaging

Quick answer Chargaff's base ratios plus X-ray diffraction data gave Watson and Crick the antiparallel double helix, and histones then fold two metres of that helix into a nucleus a few micrometres wide.

A nucleotide has three parts: a nitrogenous base, a pentose sugar and a phosphate group. The sugar is deoxyribose in DNA and ribose in RNA. A base joined to the sugar alone is a nucleoside; add the phosphate and it becomes a nucleotide. The bases come in two chemical types. Purines, which are adenine and guanine, have two fused rings. Pyrimidines, which are cytosine and thymine in DNA and cytosine and uracil in RNA, have a single ring. Nucleotides are linked into a chain by phosphodiester bonds, each joining the 3' carbon of one sugar to the 5' carbon of the next through a phosphate. This gives every strand a direction, called polarity: one end has a free 5' phosphate and the other a free 3' hydroxyl.

Erwin Chargaff analysed DNA from many organisms and found a rule hiding in the numbers. In any double-stranded DNA the amount of adenine always equals the amount of thymine, and the amount of guanine always equals the amount of cytosine. It follows that total purines equal total pyrimidines, so the ratio (A+G)/(T+C) works out to 1. Note that the ratio (A+T)/(G+C) is not fixed; it changes from one organism to another, and that is exactly why it can be used to tell species apart.

In 1953 James Watson and Francis Crick, using the X-ray diffraction photographs produced by Maurice Wilkins and Rosalind Franklin, proposed the double helix. Two polynucleotide chains wind around a common axis in a right-handed spiral. The sugar-phosphate backbones lie on the outside and the bases point inward, stacked nearly flat one above the other like the steps of a spiral staircase. The two chains are antiparallel: where one runs 5' to 3', its partner alongside runs 3' to 5'. The chains are held together by hydrogen bonds between the bases, and the pairing is fixed. Adenine pairs with thymine through two hydrogen bonds, and guanine pairs with cytosine through three. Because a two-ring purine always faces a one-ring pyrimidine, the helix keeps a uniform width of about 2 nm all along its length. Successive base pairs sit 0.34 nm apart, there are roughly ten base pairs in one complete turn, and so one turn, called the pitch, measures about 3.4 nm. DNA rich in G and C needs more heat to separate, because each of those pairs is held by an extra hydrogen bond.

Now the packing problem. A human cell carries about 6.6 x 109 base pairs of DNA, and at 0.34 nm per base pair that stretches to roughly 2.2 metres of thread. It has to fit inside a nucleus only a few micrometres across. The solution is to wind the DNA on protein spools. DNA is negatively charged because of its phosphate groups, and histones are basic proteins carrying a positive charge, because they are rich in the positively charged amino acids lysine and arginine, so the two stick to each other. Eight histone molecules, two copies each of H2A, H2B, H3 and H4, form a histone octamer. About 200 base pairs of DNA wrap around one octamer, and the DNA-plus-octamer unit is called a nucleosome. Histone H1 sits on the outside and seals the DNA where it enters and leaves the bead. Under an electron microscope a gently stretched chromatin fibre looks like beads threaded on a string, each bead a nucleosome and the string the linker DNA joining them. These nucleosomes then coil and fold on themselves repeatedly, helped by a set of additional proteins called non-histone chromosomal proteins, until the whole length fits. In an interphase nucleus the loosely packed, lightly stained regions are euchromatin and are transcriptionally active, while the tightly packed, darkly stained regions are heterochromatin and are inactive.

Nucleoside vs nucleotide Nucleoside = base + sugar. Nucleotide = base + sugar + phosphate. Only nucleotides can be joined into a chain.
0.34 nm vs 3.4 nm vs 2 nm 0.34 nm is the rise between two adjacent base pairs, 3.4 nm is one full turn (the pitch, about 10 base pairs), and 2 nm is the diameter of the helix.
(A+G)/(T+C) = 1 but (A+T)/(G+C) varies The first ratio is fixed by base pairing in every double-stranded DNA; the second differs between species and is not a Chargaff rule.
Histone octamer vs nucleosome The octamer is protein only. The nucleosome is that octamer with roughly 200 base pairs of DNA wound around it.
Euchromatin vs heterochromatin Loosely packed and transcribed versus densely packed and silent. Staining intensity tells them apart in a stained nucleus.
Remember
  • Nucleotides join through 3'-5' phosphodiester bonds, giving each strand a 5' end and a 3' end.
  • The two strands are antiparallel and right-handed; backbones outside, bases paired inside.
  • A pairs with T through two hydrogen bonds and G pairs with C through three, so G-C rich DNA is harder to melt apart.
  • Chargaff's rules: A equals T and G equals C, hence purines equal pyrimidines and (A+G)/(T+C) equals 1.
  • One histone octamer (two each of H2A, H2B, H3 and H4) plus about 200 base pairs of DNA makes one nucleosome; H1 seals it.
  • Euchromatin is loosely packed, lightly stained and active; heterochromatin is condensed, darkly stained and inactive.

Replication: Semi-conservative and Fast

Quick answer Each old strand acts as a template for a new one, so every daughter molecule keeps one parental strand. Meselson and Stahl proved this with heavy nitrogen and a density gradient.

Watson and Crick pointed out in their own paper that the fixed base pairing suggests an obvious copying mechanism. If the two strands separate, each one specifies the sequence of its new partner. This predicts that every daughter DNA molecule will contain one old strand and one newly made strand, which is what semi-conservative replication means. Two rival ideas had to be ruled out: conservative replication, in which the original molecule stays whole and a completely new molecule is made alongside it, and dispersive replication, in which old and new stretches are mixed through both strands.

Matthew Meselson and Franklin Stahl tested this in 1958. They grew Escherichia coli for many generations in a medium where the only nitrogen source was ammonium chloride made with the heavy isotope of nitrogen, so every nitrogen atom in the bases was heavy and the whole DNA was denser than normal. The cells were then shifted to a medium containing ordinary light nitrogen, and samples were taken at regular intervals. The DNA from each sample was spun in a caesium chloride density gradient, a technique that lets each molecule settle at the level matching its own density, so molecules of different weight form separate bands. E. coli divides about every twenty minutes in these conditions. After twenty minutes, that is one round of replication, all the DNA formed a single band of intermediate density, exactly halfway between heavy and light. That result alone kills conservative replication, which would have produced one heavy band and one light band with nothing in between. After forty minutes, two rounds, there were two bands in equal amounts, one hybrid and one light. That pattern is what semi-conservative replication predicts and dispersive replication does not. Similar experiments by Taylor and his colleagues on root tips of Vicia faba, the broad bean, using radioactive thymidine, showed that whole chromosomes replicate the same way.

Replication does not begin just anywhere. It starts at a defined sequence called the origin of replication. A bacterial chromosome has a single origin; a long eukaryotic chromosome has many, so the copying can be shared out. Helicase unwinds the double helix, single-strand binding proteins hold the separated strands apart so they do not snap back together, and topoisomerase (DNA gyrase in bacteria) releases the twisting strain that builds up ahead of the unwinding point. The Y-shaped junction where the two strands part is the replication fork.

The central enzyme is DNA-dependent DNA polymerase. Its raw materials are deoxyribonucleoside triphosphates, and these do double duty: they supply the nucleotide and, when their two terminal phosphates are cut off, they supply the energy for joining it on. The enzyme has two firm limitations that shape the whole process. First, it can add a nucleotide only to a free 3' hydroxyl group, so a new strand can grow only in the 5' to 3' direction. Second, it cannot start a chain on bare template, so a short RNA primer laid down by primase provides the first 3' end for it to build from.

Because the two template strands run in opposite directions, only one new strand can be built continuously in the same direction that the fork is opening. That one is called the leading strand. On the other template the enzyme has to work away from the fork, so the new strand is built in short pieces called Okazaki fragments, each started fresh as more template is exposed; this is the lagging strand and its synthesis is discontinuous. DNA ligase then seals the nicks between the fragments to give one continuous strand. Notice that ligase does not make new strand, it only joins ends. Speed matters here: the E. coli genome is about 4.6 x 106 base pairs, and it is copied at roughly 2000 base pairs per second. Accuracy matters just as much, because every uncorrected mistake becomes a mutation. In a eukaryotic cell, replication is confined to the S phase of interphase, and it is tightly linked to cell division; if the DNA replicates but the cell then fails to divide, the chromosome number doubles and polyploidy results.

Semi-conservative vs conservative vs dispersive One old plus one new strand per molecule, versus an untouched old molecule beside a fully new one, versus old and new patches scattered through both strands.
Leading vs lagging strand Leading is built continuously towards the fork. Lagging is built in Okazaki fragments pointing away from the fork. Both are still made 5' to 3'.
DNA polymerase vs DNA ligase Polymerase adds nucleotides to a growing strand. Ligase adds no nucleotides at all; it only seals the gap between two adjacent fragments.
Helicase vs topoisomerase Helicase breaks the hydrogen bonds and separates the two strands. Topoisomerase relieves the supercoiling that this unwinding creates ahead of the fork.
Hybrid DNA after n generations Only the two original heavy strands survive, so exactly two molecules stay hybrid however long you wait; after four generations that is 2 hybrid out of 16 molecules, or 12.5 per cent.
Remember
  • Semi-conservative means each daughter DNA molecule has one parental strand and one newly synthesised strand.
  • Meselson and Stahl: after one generation in light nitrogen all DNA was of hybrid density; after two generations, half hybrid and half light.
  • The all-hybrid band after one generation rules out conservative replication; the appearance of a distinct light band after two rules out dispersive.
  • DNA polymerase works only in the 5' to 3' direction and needs a free 3' hydroxyl, which is why a primer is required.
  • Leading strand is made continuously; lagging strand is made as Okazaki fragments that DNA ligase later joins.
  • In eukaryotes replication occurs in the S phase of interphase and starts from many origins on each chromosome.

Transcription and RNA Processing

Quick answer Only one strand of one gene is copied into RNA. Bacteria manage with a single RNA polymerase and no editing; eukaryotes use three polymerases and then cap, tail and splice the transcript.

Before the details, the shape of the whole process. Francis Crick set out the direction in which genetic information moves inside a cell as DNA to RNA to protein, and this statement is called the central dogma. DNA is copied into RNA by transcription, and the RNA is then read into a chain of amino acids by translation. In some viruses the first arrow runs backwards, so that RNA is copied into DNA, but for the rest of life the order is fixed. This section deals with the first arrow.

Transcription is the copying of one segment of DNA into RNA. It differs from replication in two ways: only a small part of the DNA is copied, and only one of the two strands is used as the template. There is a good reason for using only one strand. If both were transcribed, they would produce two different RNA molecules coding for two different proteins, which would make one gene mean two things; worse, the two RNAs would be complementary to each other, so they would pair up into a double-stranded molecule and translation would stop.

The stretch of DNA that is transcribed as one piece is called a transcription unit, and it has three parts. The promoter comes first and is the landing site for RNA polymerase; it lies towards the 5' end of the coding strand, which is why it is described as upstream. The structural gene comes next. The terminator comes last, towards the 3' end of the coding strand, and it is where transcription stops.

Of the two DNA strands, the one actually read by the enzyme is the template strand, and it runs 3' to 5' in the direction the polymerase travels. The other strand is called the coding strand. It is never copied, yet it is the more useful one to look at, because its sequence is identical to the RNA that is made, with thymine replaced by uracil everywhere. So if you are given the coding strand and asked for the mRNA, you simply rewrite it and change every T to a U. RNA polymerase, like DNA polymerase, can only build in the 5' to 3' direction, and that is exactly why it has to read the template from the template's 3' end towards its 5' end.

In prokaryotes a single RNA polymerase makes all three kinds of RNA, that is mRNA, tRNA and rRNA. It cannot start on its own: an accessory protein called the sigma factor associates with it for initiation and leaves once transcription is under way. Termination needs another accessory protein, the rho factor. Because a bacterial cell has no nuclear membrane, ribosomes can begin translating an mRNA while its far end is still being transcribed, so transcription and translation are coupled in time and place. Bacterial mRNA is often polycistronic, meaning one mRNA molecule carries several genes read one after another, and it needs no processing before use.

In eukaryotes the work is divided among three nuclear RNA polymerases. RNA polymerase I transcribes the ribosomal RNAs 28S, 18S and 5.8S. RNA polymerase II transcribes the precursor of mRNA, called heterogeneous nuclear RNA or hnRNA. RNA polymerase III transcribes tRNA, the 5S rRNA and the small nuclear RNAs. Eukaryotic mRNA is monocistronic, that is one mRNA for one polypeptide.

Eukaryotic genes are also split genes. The coding stretches, called exons, are interrupted by non-coding stretches called introns, whose other name, intervening sequences, says exactly what they are. Only exons survive into the mature message. The primary transcript therefore has to be processed in three ways inside the nucleus. Splicing removes the introns and joins the exons together in their original order. Capping adds an unusual nucleotide, methyl guanosine triphosphate, to the 5' end. Tailing adds a stretch of about 200 to 300 adenylate residues to the 3' end, and this poly-A tail is put on without any template to copy from. Only after capping, tailing and splicing does the hnRNA become mature mRNA and move out of the nucleus to the ribosomes. Because different combinations of exons can sometimes be joined, one split gene can yield more than one protein, which is part of why humans have far fewer genes than proteins.

Template strand vs coding strand Template is 3' to 5' and is actually copied. Coding strand is 5' to 3', is not copied, but reads the same as the mRNA except that T replaces U.
Exon vs intron Exon is expressed and stays in the mature RNA. Intron is an intervening sequence that is cut out during splicing.
Sigma factor vs rho factor Sigma is needed to start transcription in bacteria, rho to terminate it. Neither is part of the core enzyme permanently.
Capping vs tailing Capping puts methyl guanosine triphosphate on the 5' end; tailing adds a template-independent poly-A stretch to the 3' end.
Polycistronic vs monocistronic mRNA Bacterial mRNA commonly carries several genes in one transcript; eukaryotic mRNA carries one.
Remember
  • A transcription unit has a promoter, a structural gene and a terminator; the promoter is upstream, near the 5' end of the coding strand.
  • The template strand runs 3' to 5' and is read by the enzyme; the coding strand has the same sequence as the RNA with T in place of U.
  • Prokaryotes have one RNA polymerase; sigma factor is needed for initiation and rho factor for termination.
  • Eukaryotes have three polymerases: I for 28S, 18S and 5.8S rRNA, II for hnRNA, III for tRNA, 5S rRNA and snRNAs.
  • Exons are expressed and appear in mature mRNA; introns are removed by splicing.
  • hnRNA is capped at the 5' end with methyl guanosine triphosphate and tailed at the 3' end with 200 to 300 adenylate residues before export.

The Genetic Code and the Adapter Molecule

Quick answer Sixty-four triplets, sixty-one of which name an amino acid, read without commas or overlaps. tRNA is the molecule that physically connects a codon to the amino acid it stands for.

The genetic code is the rule that converts a base sequence into an amino acid sequence. The size of the code word can be worked out by counting. There are four bases. If one base named one amino acid, only four amino acids could be specified. If two bases were read together, there would be four times four, that is sixteen combinations, still short of the twenty amino acids that proteins use. Three bases give four times four times four, which is sixty-four, comfortably more than enough. George Gamow argued on exactly these mathematical grounds that the code must be a triplet code. It was established experimentally through Marshall Nirenberg's cell-free protein synthesising system and Har Gobind Khorana's chemically synthesised RNA molecules with defined repeating sequences, helped by the enzyme discovered by Severo Ochoa, which could polymerise RNA of known composition.

Of the sixty-four codons, sixty-one specify amino acids. The remaining three, UAA, UAG and UGA, do not code for any amino acid at all and act as stop or termination signals. AUG has two jobs at once: it is the initiation codon, and it also codes for the amino acid methionine wherever it appears inside a message. The code is degenerate, which means that most amino acids are specified by more than one codon. Degenerate does not mean sloppy: the code is also unambiguous and specific, because any one codon codes for one amino acid and no other. It is read continuously with no punctuation between successive codons, which is what comma-less means, and it is non-overlapping, so each base belongs to exactly one codon. Finally it is nearly universal: the same codon means the same amino acid in a bacterium and in a human being, with a small number of exceptions such as the code used inside mitochondria.

Because the code has no commas, the reading frame is fixed entirely by where translation starts. That has a consequence worth understanding. If one or two bases are inserted into or deleted from a gene, every codon after that point is read in the wrong grouping, and the rest of the protein comes out as nonsense. This is a frameshift mutation. If three bases are added or lost together, the frame is restored after that point, so only one amino acid is gained or lost and the rest of the protein is normal. A different kind of change is the substitution of a single base, which alters only one codon. In sickle cell anaemia, a single base substitution changes the sixth codon of the gene for the beta chain of haemoglobin from GAG to GUG, so glutamic acid at the sixth position of that chain is replaced by valine, and that one swap changes the behaviour of the whole haemoglobin molecule.

There is a gap in this story that has to be filled. A codon is a sequence of bases and an amino acid is a completely different sort of molecule; the two cannot recognise each other directly. Francis Crick predicted that some adapter molecule must exist with two ends, one that reads the code and one that carries the amino acid. That molecule turned out to be transfer RNA. Each tRNA has an anticodon loop carrying three bases complementary to the codon it must recognise, and an amino acid acceptor end at which the correct amino acid is attached. There is a separate tRNA for each amino acid, plus a special initiator tRNA for the start codon. Importantly, there are no tRNAs for the three stop codons, which is precisely why the ribosome cannot go past them. Drawn flat on paper, a tRNA looks like a clover leaf, but in the cell it folds up into a compact shape resembling an inverted L, with the anticodon at one tip and the amino acid at the other.

64 = 61 coding + 3 stop Only three codons, UAA, UAG and UGA, are non-coding. Do not count AUG among the non-coding ones; it codes for methionine.
Degenerate vs ambiguous Degenerate means one amino acid has several codons. Ambiguous would mean one codon has several amino acids, and the genetic code is not ambiguous.
Codon vs anticodon Codon is the triplet on mRNA. Anticodon is the complementary triplet on tRNA. The codon is read 5' to 3' on the message.
Frameshift vs point substitution Adding or removing one or two bases shifts every codon downstream; replacing one base changes a single codon, as in the GAG to GUG change in the haemoglobin beta chain gene.
Remember
  • The code is a triplet code: 4 x 4 x 4 gives 64 codons, of which 61 specify amino acids.
  • UAA, UAG and UGA are stop codons; AUG is both the initiation codon and the codon for methionine.
  • The code is degenerate (several codons per amino acid) but unambiguous (one codon means one amino acid).
  • It is comma-less, non-overlapping and nearly universal.
  • Insertion or deletion of one or two bases causes a frameshift; three together add or remove a single amino acid.
  • tRNA is Crick's adapter: an anticodon at one end reads the codon and the amino acid is attached at the other end.

Translation and the lac Operon

Quick answer Ribosomes read mRNA three bases at a time and stitch amino acids together. Bacteria then save effort by switching whole gene sets on only when needed, as the lac operon shows.

Translation is the building of a polypeptide from the order of bases in an mRNA. The first step happens before any ribosome is involved. Each amino acid is joined to its own tRNA in a reaction that consumes ATP, a step called activation of the amino acid or charging of the tRNA, also known as aminoacylation. Only a charged tRNA can be used, and once two charged tRNAs are held side by side on the ribosome, a peptide bond can form between the amino acids they carry.

The ribosome is the factory where this happens. It is made of two subunits which lie apart from each other when the ribosome is inactive; when the smaller subunit meets an mRNA the two subunits come together and translation begins. A bacterial ribosome is a 70S particle made of a 50S and a 30S subunit, while a eukaryotic cytoplasmic ribosome is 80S, made of a 60S and a 40S subunit. The larger subunit has grooves that can hold two tRNAs next to each other, so that the growing chain on one can be transferred onto the amino acid on the next. The peptide bond is formed not by a protein enzyme but by one of the ribosomal RNA molecules acting as a catalyst, which is why it is called a ribozyme.

An mRNA carries more than just the coding message. Between its 5' end and the start codon lies a stretch that is never translated, and a second untranslated stretch follows the stop codon before the 3' end. These untranslated regions carry no amino acid information but are needed for translation to run efficiently. Translation itself has three phases. In initiation, the ribosome assembles on the mRNA at the start codon and only the initiator tRNA is accepted there. In elongation, charged tRNAs arrive one after another, each anticodon pairing with the next codon, and the polypeptide grows by one amino acid per codon while the ribosome moves along. In termination, a stop codon arrives, no tRNA can pair with it, a release factor binds instead, the finished polypeptide is set free and the two subunits separate.

A cell does not need every protein all the time, and making an enzyme for a food that is not present is pure waste. In bacteria the main control point is transcription, and the classic example is the lac operon, worked out by Francois Jacob and Jacques Monod. An operon is a set of structural genes transcribed together as a single unit, together with the switches placed in front of them. Along the DNA, in order, the lac system has the regulatory gene i, then the promoter, then the operator, and then three structural genes named z, y and a in that sequence. Gene z codes for beta-galactosidase, the enzyme that splits lactose into galactose and glucose. Gene y codes for permease, which increases the permeability of the cell to lactose. Gene a codes for transacetylase. All three are transcribed together into one long polycistronic mRNA, so they are made and stopped as a group.

The i gene is transcribed at a low level all the time and produces a protein called the repressor. When lactose is not available, the repressor binds tightly to the operator, which sits between the promoter and gene z. RNA polymerase can still settle on the promoter but cannot move forward past the occupied operator, so no mRNA for the three enzymes is produced. When lactose is supplied, a small amount enters through the few permease molecules the cell always carries, and lactose then acts as the inducer. It binds to the repressor and changes its shape so that the repressor can no longer hold on to the operator. The operator is now free, RNA polymerase moves through, the polycistronic mRNA is made and the three enzymes appear. Because the substrate itself turns the genes on, the lac operon is described as an inducible operon, and because the controlling protein is a repressor whose job is to block transcription, this arrangement is called negative regulation. When the lactose has all been used up, there is nothing left to hold the repressor in its inactive shape, it binds the operator again, and the operon shuts down.

70S vs 80S ribosome 70S in prokaryotes, made of 50S and 30S. 80S in the eukaryotic cytoplasm, made of 60S and 40S. The numbers do not add up arithmetically because S is a sedimentation value, not a mass.
z, y and a genes of the lac operon z makes beta-galactosidase (splits lactose into galactose and glucose), y makes permease (lets lactose in), a makes transacetylase.
i gene vs operator The i gene is a stretch of DNA that codes for the repressor protein. The operator is a stretch of DNA that the repressor protein sits on. Do not confuse the maker with the target.
Inducer vs repressor The repressor is a protein that blocks transcription. The inducer, here lactose, is the small molecule that switches the repressor off and thereby switches the operon on.
Negative regulation The default state is off because a repressor blocks the operator; removing the block turns transcription on. Contrast this with positive regulation, where a protein must be present to allow transcription.
Remember
  • Amino acids are first attached to their own tRNAs using ATP; this charging step is called aminoacylation.
  • Bacterial ribosomes are 70S (50S plus 30S); eukaryotic cytoplasmic ribosomes are 80S (60S plus 40S).
  • The peptide bond is catalysed by ribosomal RNA acting as a ribozyme, not by a protein enzyme.
  • mRNA has untranslated regions before the start codon and after the stop codon that are needed for efficient translation.
  • The lac operon order along the DNA is i gene, promoter, operator, then z, y and a.
  • Lactose acts as the inducer by inactivating the repressor, so the lac operon is an inducible operon under negative regulation.

The Human Genome Project and DNA Fingerprinting

Quick answer One project read the 99.9 per cent of the human genome we all share; one technique reads the 0.1 per cent where we differ, and that difference is enough to identify a person.

The Human Genome Project ran from 1990 to 2003. In the United States it was coordinated by the Department of Energy and the National Institutes of Health, and the Wellcome Trust in Britain joined as a major partner. The goals were to determine the order of all roughly 3 x 109 base pairs in a human genome, to identify the genes within it, to store the information in databases and to improve the tools for analysing such data. It was called a megaproject for good reason. At a sequencing cost of about three US dollars per base, the total estimate came to about nine billion US dollars, and the sequence printed as books of a thousand pages each would fill a shelf 3300 volumes long. Because the information touches individuals so directly, the ethical, legal and social questions it raises were made part of the project from the start.

Two broad approaches were used. The first identified only the genes that are actually expressed as RNA, using short sequences read from them called Expressed Sequence Tags. The second, called sequence annotation, sequenced the whole genome, coding and non-coding regions alike, and only afterwards assigned functions to the different parts. The practical procedure was to break the DNA into small fragments, insert each fragment into a vector so that it could be multiplied inside a host cell, sequence the fragments in automated sequencers based on the method developed by Frederick Sanger, and then use computer programs to spot overlaps and assemble the fragments back into a continuous sequence. The vectors used for the very large pieces were bacterial artificial chromosomes and yeast artificial chromosomes. The assembled sequences were then assigned to particular chromosomes using genetic and physical maps built from polymorphism in restriction enzyme recognition sites and from repetitive microsatellite sequences.

The findings were full of surprises. The human genome contains about 3164.7 million base pairs. An average gene is roughly 3000 bases long, but the largest known human gene, dystrophin, stretches to 2.4 million bases. The number of genes came out at about 30000, far below earlier estimates of 80000 to 140000, and less than 2 per cent of the genome actually codes for proteins, with repeated sequences making up a very large part of the remainder. Chromosome 1 carries the most genes, 2968 of them, and the Y chromosome the fewest, 231. About 99.9 per cent of the bases are exactly the same in every human being, and around 1.4 million sites were found where individuals differ by a single base; these are called single nucleotide polymorphisms, or SNPs, and they are what makes it possible to track down chromosomal regions linked to disease. The function of more than half the genes discovered is still not known. Alongside the human genome, the genomes of several model organisms were sequenced, including rice, Arabidopsis thaliana, the fruit fly Drosophila melanogaster and the free-living non-pathogenic nematode Caenorhabditis elegans.

DNA fingerprinting, developed by Alec Jeffreys, works on the small fraction of the genome that differs from person to person, and most of that variation lies in non-coding repetitive DNA. When total genomic DNA is spun in a density gradient, these repeated stretches separate out as small peaks beside the main band, which is why they are called satellite DNA. They are classified by the length of the repeating unit into microsatellites and minisatellites. A particular class of minisatellite is the VNTR, short for variable number of tandem repeats. Its repeat unit is short, but the number of copies varies enormously from one individual to another, so the fragment produced can be anywhere from roughly 0.1 to 20 kilobases long. A person inherits one set of these repeats from each parent, so the overall pattern is effectively unique except in identical twins, and it is the same in DNA taken from any tissue of that person, whether blood, hair follicles, skin, bone, saliva or semen.

The technique runs in a fixed order. DNA is isolated from the sample. It is cut into fragments using restriction endonucleases. The fragments are separated according to size by gel electrophoresis. The separated fragments are transferred, or blotted, from the gel onto a synthetic membrane of nitrocellulose or nylon, a step named Southern blotting after its inventor. The membrane is then treated with a labelled VNTR probe, which sticks only to fragments containing that repeat, a process called hybridisation. Finally the positions where the probe has bound are revealed by autoradiography. The result is a pattern of bands of different sizes, and that pattern is characteristic of the individual. It is used in forensic investigation, in resolving questions of parentage, and in studying genetic diversity within and between populations.

EST vs sequence annotation ESTs find only the genes that are expressed as RNA. Sequence annotation reads the entire genome, coding and non-coding, and works out the functions afterwards.
BAC vs YAC Bacterial artificial chromosome and yeast artificial chromosome. Both are cloning vectors for very large DNA inserts; the difference is the host cell they are maintained in.
Microsatellite vs minisatellite (VNTR) Both are tandemly repeated satellite DNA; minisatellites have longer repeat units. The VNTR used as the fingerprinting probe is a minisatellite.
Southern blotting Transfer of DNA fragments from the electrophoresis gel to a nitrocellulose or nylon membrane. It is a transfer step only; the separation was done by electrophoresis and the visualisation is done by autoradiography.
99.9 per cent same, 0.1 per cent different The Human Genome Project describes the shared 99.9 per cent; DNA fingerprinting reads the polymorphic 0.1 per cent.
Remember
  • The Human Genome Project ran from 1990 to 2003 and read about 3 x 10^9 base pairs at an estimated cost of around nine billion US dollars.
  • Expressed Sequence Tags target only expressed genes, while sequence annotation sequences everything first and assigns function later.
  • Fragments were cloned in BAC and YAC vectors and sequenced by Sanger's method, then assembled by computer.
  • About 30000 genes, less than 2 per cent coding DNA, dystrophin the largest gene at 2.4 million bases, chromosome 1 with 2968 genes and the Y chromosome with 231.
  • Roughly 99.9 per cent of bases are identical between people; about 1.4 million single-base differences (SNPs) were catalogued.
  • DNA fingerprinting, developed by Alec Jeffreys, compares VNTR minisatellite patterns and is identical only in identical twins.

The formula sheet

Every formula in this chapter, in one place — screenshot it before your exam.

S strain vs R strain
Griffith vs Avery-MacLeod-McCarty
32P labels DNA, 35S labels protein
Why DNA is more stable than RNA
Nucleoside vs nucleotide
0.34 nm vs 3.4 nm vs 2 nm
(A+G)/(T+C) = 1 but (A+T)/(G+C) varies
Histone octamer vs nucleosome
Euchromatin vs heterochromatin
Semi-conservative vs conservative vs dispersive
Leading vs lagging strand
DNA polymerase vs DNA ligase
Helicase vs topoisomerase
Hybrid DNA after n generations
Template strand vs coding strand
Exon vs intron
Sigma factor vs rho factor
Capping vs tailing
Polycistronic vs monocistronic mRNA
64 = 61 coding + 3 stop
Degenerate vs ambiguous
Codon vs anticodon
Frameshift vs point substitution
70S vs 80S ribosome
z, y and a genes of the lac operon
i gene vs operator
Inducer vs repressor
Negative regulation
EST vs sequence annotation
BAC vs YAC
Microsatellite vs minisatellite (VNTR)
Southern blotting
99.9 per cent same, 0.1 per cent different

Test yourself

Tap an answer to check it instantly — you'll see why it's right, and what to revise if it isn't.

0 correct · 0/12 answered
Q1

In Griffith's experiment, the mice died in which of the following injections?

Q2

In the Avery, MacLeod and McCarty experiment, transformation was blocked by adding which enzyme?

Q3

In the Hershey-Chase experiment, radioactivity was found inside the bacterial cells when the phages had been labelled with which isotope?

Q4

The pitch of the B-form DNA double helix, that is the length of one complete turn, is about

Q5

A histone octamer is made of two molecules each of which set of histones?

Q6

In the Meselson and Stahl experiment, what was observed after two generations of growth in the light nitrogen medium?

Q7

Which enzyme joins the Okazaki fragments of the lagging strand into one continuous strand?

Q8

In a eukaryotic nucleus, the precursor of mRNA, that is hnRNA, is transcribed by

Q9

Which of the following is NOT a stop codon?

Q10

Which molecule was predicted by Francis Crick to act as the adapter between a codon and its amino acid?

Q11

In the lac operon, the z gene codes for which product?

Q12

Classical DNA fingerprinting as developed by Alec Jeffreys uses which sequences as the probe?

NCERT solutions & previous-year questions

Step-by-step model answers — tap a question to reveal the full solution.

NCERT questions 8

1 Why is DNA considered a better genetic material than RNA?

Both DNA and RNA can act as genetic material, but DNA does the job better on three counts. First, stability. Every ribose sugar in RNA carries a hydroxyl group at the 2' position, and this group is chemically reactive, which makes RNA easy to break down. Deoxyribose in DNA lacks that group, so DNA is far more stable. Second, RNA is itself catalytic, and a catalytic molecule is by nature a reactive one, so RNA is more easily degraded. DNA also uses thymine instead of uracil, and that substitution adds further chemical stability. Third, repair. DNA is double stranded with complementary partners, so if one strand is damaged the cell can rebuild it correctly using the other as a reference; a single-stranded RNA molecule has no such backup and mutates much faster. This is why RNA viruses show high mutation rates. One point should be stated honestly the other way round: RNA is better than DNA at expressing itself, because RNA can be translated into protein directly while DNA must first be transcribed into RNA. So DNA is the better material for storing genetic information over a lifetime and over generations, and RNA is the better material for carrying and using it, which is exactly the division of labour cells settled on.

2 Describe the Hershey-Chase experiment and state what it proved.

Hershey and Chase used bacteriophage T2, which infects Escherichia coli and is built from only a protein coat and DNA. Their experiment turns on one chemical fact: DNA contains phosphorus but no sulphur, while protein contains sulphur but no phosphorus. They grew one set of phages on a medium containing radioactive phosphorus, so only the phage DNA became radioactive, and another set on a medium containing radioactive sulphur, so only the protein coat became radioactive. Each set of phages was then allowed to infect bacteria. The infected cultures were agitated in a blender to knock the empty viral coats off the bacterial surface, and then centrifuged so that the heavier bacterial cells settled as a pellet.

The bacteria infected with the phosphorus-labelled phages were radioactive, showing that DNA had entered the cell. The bacteria infected with the sulphur-labelled phages were not radioactive, showing that the protein coat had stayed outside. Since only the DNA went in and yet complete new phages were produced, DNA and not protein must carry the genetic instructions.

3 What is semi-conservative replication, and how did Meselson and Stahl prove it?

Semi-conservative replication means the two strands of the parent DNA separate, each acts as a template, and each daughter molecule therefore ends up with one old parental strand and one newly synthesised strand.

Meselson and Stahl grew E. coli for many generations in a medium in which the only nitrogen source contained the heavy isotope of nitrogen, so all the DNA became heavy. The cells were then transferred to a medium with ordinary light nitrogen, and samples were taken at intervals. DNA from each sample was spun in a caesium chloride density gradient, in which each molecule settles at a level matching its own density. After one round of replication, about twenty minutes, the whole of the DNA formed a single band of intermediate density, exactly halfway between heavy and light. Conservative replication would have given a heavy band and a light band with nothing between them, so it was ruled out at once. After two rounds, about forty minutes, there were two bands present in equal amounts, one of intermediate density and one fully light. That is exactly what semi-conservative replication predicts and what dispersive replication does not. Taylor and colleagues obtained the same conclusion for whole chromosomes using radioactive thymidine in root tips of Vicia faba.

4 Differentiate between the template strand and the coding strand. If the coding strand of a transcription unit reads 5'-ATGCATGCATGCATGCATGCATGCATGC-3', write the sequence of the mRNA.

The template strand is the one actually read by RNA polymerase. It runs 3' to 5' in the direction the enzyme moves, and the RNA made is complementary to it. The coding strand is the other strand. It is never copied, but because it is complementary to the template it has the same sequence as the RNA, except that thymine stands wherever the RNA has uracil. It is called the coding strand only because it lets us read off the message directly.

So to get the mRNA from the coding strand you simply rewrite the sequence and replace every T with a U. The coding strand given is 5'-ATGCATGCATGCATGCATGCATGCATGC-3', so the mRNA is:

5'-AUGCAUGCAUGCAUGCAUGCAUGCAUGC-3'

The template strand for the same unit would be the complement written in the opposite direction, that is 3'-TACGTACGTACGTACGTACGTACGTACG-5'.

5 Describe the processing that a eukaryotic primary transcript undergoes before it can be translated.

The primary transcript made by RNA polymerase II is called heterogeneous nuclear RNA, or hnRNA. It cannot be translated as it is, because eukaryotic genes are split genes containing non-coding introns in between the coding exons. Three changes are made inside the nucleus.

Splicing removes the introns and joins the exons together in their original order, so that the coding message becomes continuous. Capping adds an unusual nucleotide, methyl guanosine triphosphate, to the 5' end of the transcript. Tailing adds a stretch of about 200 to 300 adenylate residues to the 3' end, and unlike everything else in transcription this poly-A tail is added without any template to copy from. Only after all three steps is the molecule mature mRNA, and only then is it transported out of the nucleus to the ribosomes. Because different combinations of exons can sometimes be spliced together, one gene may give rise to more than one polypeptide.

6 Explain how the lac operon is regulated in the presence and in the absence of lactose.

Along the DNA, the lac system is arranged in this order: the regulatory gene i, then the promoter, then the operator, and then the structural genes z, y and a. Gene z codes for beta-galactosidase, which splits lactose into galactose and glucose; gene y codes for permease, which increases the entry of lactose into the cell; gene a codes for transacetylase. The three are transcribed together as one polycistronic mRNA.

The i gene is expressed at a low level all the time and makes the repressor protein. In the absence of lactose, this repressor binds to the operator. RNA polymerase can still occupy the promoter but cannot move past the blocked operator, so the three enzymes are not made and the cell wastes no resources. In the presence of lactose, a small quantity enters the cell through the few permease molecules already present, and lactose acts as the inducer. It binds the repressor and alters its shape so that the repressor can no longer attach to the operator. The operator is now free, RNA polymerase transcribes z, y and a, and the enzymes are produced. Since the substrate itself induces the genes, this is an inducible operon, and since control is exercised by a blocking repressor, it is an example of negative regulation. Once the lactose is exhausted, the repressor is released from the inducer, binds the operator again and the operon switches off.

7 List the salient features of the human genome as revealed by the Human Genome Project.

The human genome contains about 3164.7 million nucleotide base pairs. The average gene is roughly 3000 bases long, but sizes vary enormously; the largest known human gene, dystrophin, has 2.4 million bases. The estimated number of genes came to about 30000, far fewer than the earlier guesses of 80000 to 140000. Less than 2 per cent of the genome codes for proteins, and repeated sequences make up a very large proportion of the rest. Repeated sequences have no direct coding function but tell us a great deal about chromosome structure and dynamics, and about evolution. Chromosome 1 carries the largest number of genes, 2968, while the Y chromosome carries the fewest, 231. Nearly 99.9 per cent of the nucleotide bases are exactly the same in all people, and about 1.4 million locations were identified where single-base differences occur; these single nucleotide polymorphisms help in tracing disease-associated regions of chromosomes. The function of over 50 per cent of the genes discovered is still unknown.

8 What is DNA fingerprinting? Outline its steps and mention its applications.

DNA fingerprinting, developed by Alec Jeffreys, is a technique for identifying differences in specific regions of DNA between individuals. It works on repetitive non-coding DNA. When genomic DNA is spun in a density gradient these repeats form small satellite peaks beside the main band, and are therefore called satellite DNA. The particular class used is the VNTR, or variable number of tandem repeats, a minisatellite whose copy number differs greatly from one individual to another, giving fragments of roughly 0.1 to 20 kilobases.

The steps are: isolate DNA from the sample; digest it with restriction endonucleases; separate the fragments by size using gel electrophoresis; transfer the separated fragments from the gel to a nitrocellulose or nylon membrane, a step called Southern blotting; hybridise the membrane with a labelled VNTR probe; and detect the bound fragments by autoradiography. The result is a set of bands whose pattern is characteristic of the individual and the same in DNA taken from any tissue of that person. The only people who share the same pattern are identical twins, who developed from a single fertilised egg.

It is used in forensic investigation to match samples from a scene with a person, in settling questions of parentage, and in studying genetic diversity within and between populations.

Previous-year board questions 6

Q1 Explain the process of transcription in a eukaryotic cell. Why does the primary transcript require processing? 3 marks mark

In a eukaryote, transcription happens in the nucleus and uses three different RNA polymerases. RNA polymerase I makes the 28S, 18S and 5.8S ribosomal RNAs, RNA polymerase II makes hnRNA, the precursor of mRNA, and RNA polymerase III makes tRNA, 5S rRNA and the small nuclear RNAs. The enzyme binds the promoter, which lies upstream of the structural gene, opens the helix and reads the template strand from its 3' end towards its 5' end while building the RNA in the 5' to 3' direction. Transcription stops at the terminator.

Processing is needed because eukaryotic genes are split genes: their coding exons are interrupted by non-coding introns, which would be translated as nonsense if left in. Splicing removes the introns and joins the exons in order. In addition, capping adds methyl guanosine triphosphate to the 5' end and tailing adds about 200 to 300 adenylate residues to the 3' end, both of which are required before the mature mRNA leaves the nucleus for translation.

Q2 In the Meselson and Stahl experiment, what proportion of the DNA would be of hybrid density and what proportion fully light after 80 minutes of growth in the light nitrogen medium? Justify your answer. 3 marks mark

Since E. coli divides roughly every 20 minutes, 80 minutes allows four rounds of replication. Starting from one heavy DNA molecule, four rounds give 2 to the power 4, that is 16 molecules.

The key point is that only the two original heavy strands survive from the parent, and each of them ends up paired with a newly made light strand. So exactly two molecules are of hybrid density, however many generations pass. The remaining 14 molecules are made entirely of new light strands.

Therefore after 80 minutes, 2 out of 16 molecules, that is 12.5 per cent, band at hybrid density, and 14 out of 16, that is 87.5 per cent, band at light density. No heavy DNA remains after the very first round. Working the same way, one round gives 100 per cent hybrid, two rounds give 50 per cent hybrid and 50 per cent light, and three rounds give 25 per cent hybrid and 75 per cent light.

Q3 State any five features of the genetic code. 3 marks mark

First, the code is a triplet code: three bases on the mRNA specify one amino acid, giving 64 possible codons, of which 61 code for amino acids. Second, it is degenerate, meaning that most amino acids are specified by more than one codon. Third, it is unambiguous and specific, because one codon codes for one amino acid and no other; degeneracy runs in only one direction. Fourth, it is comma-less and non-overlapping: the bases are read continuously three at a time with no punctuation, and no base is shared between two codons, so the reading frame is fixed by the start codon. Fifth, it is nearly universal, the same codons meaning the same amino acids from bacteria to human beings, with a few exceptions such as the codes used in mitochondria. In addition, UAA, UAG and UGA act as stop codons and code for no amino acid, while AUG serves both as the initiation codon and as the codon for methionine.

Q4 Why is the lac operon described as an inducible operon under negative control? Explain with reference to the role of lactose. 3 marks mark

It is called inducible because the genes are normally kept off and are switched on only when their substrate appears. Lactose is that substrate and acts as the inducer. A little lactose enters through the permease molecules already present in the cell, binds to the repressor protein made by the i gene, and changes its shape so that it can no longer sit on the operator. With the operator free, RNA polymerase moves out of the promoter and transcribes z, y and a as a single polycistronic mRNA, producing beta-galactosidase, permease and transacetylase.

It is called negative control because the controlling protein is a repressor whose normal action is to block transcription. In the absence of lactose the repressor occupies the operator, which lies between the promoter and gene z, and RNA polymerase cannot move past it, so the operon stays off. Regulation therefore works by removing a block rather than by supplying an activator.

Q5 Explain how a length of DNA over two metres long is packaged inside a nucleus only a few micrometres wide. 3 marks mark

A human cell contains about 6.6 x 109 base pairs of DNA, and since consecutive base pairs are 0.34 nm apart this comes to roughly 2.2 metres of thread. Packing depends on charge attraction. DNA is negatively charged because of its phosphate groups, and histones are basic proteins rich in the positively charged amino acids lysine and arginine, so the DNA binds tightly to them.

Eight histone molecules, two copies each of H2A, H2B, H3 and H4, form a histone octamer, and about 200 base pairs of DNA are wound around it to form a nucleosome. Histone H1 lies on the outside and seals the DNA where it enters and leaves. A chain of nucleosomes joined by linker DNA gives chromatin its beads-on-a-string appearance in an electron micrograph. The chromatin fibre then coils and folds on itself repeatedly, assisted by a set of non-histone chromosomal proteins, producing the highly condensed chromosome. In an interphase nucleus, loosely packed lightly stained euchromatin is transcriptionally active while densely packed darkly stained heterochromatin is inactive.

Q6 State the exact role played at a replication fork by each of the following: helicase, primase, DNA polymerase and DNA ligase. 3 marks mark

Helicase unwinds the double helix by breaking the hydrogen bonds between the paired bases, creating the Y-shaped replication fork; the separated strands are then held apart by single-strand binding proteins.

Primase lays down a short RNA primer on each template. This is necessary because DNA polymerase cannot begin a new chain on bare template; it can only extend from an existing free 3' hydroxyl group.

DNA-dependent DNA polymerase is the main enzyme. It adds deoxyribonucleoside triphosphates one at a time to the growing 3' end, so synthesis always proceeds in the 5' to 3' direction. The incoming nucleotides also supply the energy for the reaction when their two terminal phosphates are removed. On the leading strand it works continuously towards the fork; on the lagging strand it produces short Okazaki fragments pointing away from the fork.

DNA ligase adds no nucleotides at all. It seals the nicks between adjacent Okazaki fragments so that the lagging strand becomes one continuous molecule.

Part of Priodemy for School

Interactive Maths & Science — free with every school on Priodemy EduSuite. Explore more chapters and labs on the Priodemy for School hub.

Ask AI