How does the genome specify the functional complexity and diversity evident in Fig. 1? As we saw in the previous chapter, genetic information is contained in DNA in the chromosomes, within the cell nucleus. However, protein synthesis, the process through which information encoded in the genome is used to specify cellular functions, takes place in the cytoplasm. This compart mentalization reflects the fact that the human organism is a eukaryote. This means that human cells have a nucleus containing the genome, which is separated by a nuclear membrane from the cytoplasm. In contrast, in prokaryotes like the intestinal bacterium Escherichia coli, DNA is not enclosed within a nucleus. Because of the compart mentalization of eukaryotic cells, information transfer from the nucleus to the cytoplasm is a complex process that has been a focus of much attention among molecular and cellular biologists.

Fig1. The amplification of genetic information from genome to gene products to gene networks and ultimately to cellular function and phenotype. The genome contains both protein-coding genes (blue) and noncoding RNA (ncRNA) genes (red). Many genes in the genome use alternative coding information to generate multiple different products. Both small and large ncRNAs participate in gene regulation. Many proteins participate in multigene networks that respond to cellular signals in a coordinated and combinatorial manner, thus further expanding the range of cellular functions that underlie organismal phenotypes.
The molecular link between these two related types of information—the DNA code of genes and the amino acid code of protein—is ribonucleic acid (RNA). The chemical structure of RNA is similar to that of DNA, except that each nucleotide in RNA has a ribose sugar component instead of a deoxyribose; in addition, uracil (u) replaces thymine as one of the pyrimidine bases of RNA (Fig. 2). An additional difference between RNA and DNA is that RNA in most organisms exists as a single-stranded molecule, whereas DNA, as we saw in Chapter 2, exists as a double helix.

Fig2. The pyrimidine uracil and the structure of a nucleotide in RNA. Note that the sugar ribose replaces the sugar deoxyribose of DNA.
The informational relationships among DNA, RNA, and protein are intertwined: genomic DNA directs the synthesis and sequence of RNA, RNA directs the syn thesis and sequence of polypeptides, and specific proteins are involved in the synthesis and metabolism of DNA and RNA. This flow of information is referred to as the central dogma of molecular biology.
Genetic information is stored in the DNA of the genome by means of a code (the genetic code, discussed later) in which the sequence of adjacent bases ultimately determines the sequence of amino acids in the encoded polypeptide. First, RNA is synthesized from the DNA template through the process of transcription. The RNA, carrying the coded information in a form called messenger RNA (mRNA), is then transported from the nucleus to the cytoplasm, where the RNA sequence is decoded, or translated, to determine the sequence of amino acids in the protein being synthesized. The process of translation occurs on ribosomes, which are cytoplasmic organelles with binding sites for all of the interacting molecules, including the mRNA, involved in protein synthesis. Ribosomes are themselves made up of many different structural proteins in association with specialized types of RNA known as ribosomal RNA (rRNA). Translation involves yet a third type of RNA, transfer RNA (tRNA), which provides the molecular link between the code contained in the base sequence of each mRNA and the amino acid sequence of the protein encoded by that mRNA.
Because of the interdependent flow of information represented by the central dogma, one can begin discussion of the molecular genetics of gene expression at any of its three informational levels: DNA, RNA, or protein. We begin by examining the structure of genes in the genome as a foundation for discussion of the genetic code, transcription, and translation.