The perfect substrate for sequencing of a complex genome would be to have each chromosomal DNA molecule represented by a set of DNA clones containing large inserts that have been linearly ordered according to the subchromosomal origins of their inserts, with neighboring clones having partially overlapping sequences, a tiling path. A series of such clones where the insert of each clone partially overlaps that of its neighbors with no gaps is known as a clone contig: the clones collectively represent a contiguous (continuous) DNA sequence from a subchromosomal region (Figure 1), or even a whole chromosome.

Fig1. Assembling clones in a clone contig. Because partial digestion of genomic DNA is employed to construct a genomic DNA library, some DNA clones from a genome region will have inserts that partially overlap with each other (see Figure 2). Here, the chromosomal DNA sequence from positions A to B is represented by a linear series of overlapping DNA inserts, a tiling path. The collection of clones whose inserts produce a tiling path is known as a clone contig.

Fig2. Generating clones with overlapping DNA inserts when constructing a genomic DNA library. Because all nucleated cells of an individual have essentially the same genomic DNA content, easily accessible cells (e.g., white blood cells) can be used as source material. Given that the cells contain the same sets of DNA molecules, the starting DNA will contain very many copies of each type of DNA molecule. Here, we imagine zooming in on a short region present on four copies of the same chromosomal DNA molecule (for example, paternal chromosome number 1 contributed by each of four different cells). The four copies of this sequence (#1 to #4) will have identical restriction sites for a specific restriction endonuclease (short vertical blue bars). However, because partial digestion is used, the DNA will be cleaved at only a small subset of the available restriction sites (indicated by yellow darts). Because the choice of which restriction site is cleaved is essentially random, the enzyme cuts the different copies of the same DNA sequence at different places, so that restriction fragments with overlapping sequences are produced (for example, fragment F from copy #2 partially overlaps fragments B and C from copy #1, fragments I and J from copy #3, and fragments M and N from copy #4).
Recall from Figure 2 that DNA library construction produces DNA fragments with overlapping sequences, but when they are cloned in cells the original order is lost; the problem for genome assembly is to put the clones back together with their inserts in the same order as on the chromosome of origin. To establish clone contigs, clones with over lapping inserts can be identified using clone fingerprinting methods. To do that, the standard procedure in genome projects is to establish a high-density DNA marker map for each chromosome, and then screen DNA clones for the presence or absence of DNA markers known to map to that chromosome. Clones are then identified that test positive for the same marker or markers. As we show below, the DNA markers needed to meet only two requirements: they should have a unique subchromosomal location, and be able to be conveniently assayed by PCR.