From the earliest stages of genome and gene analysis, central computer repositories were established for storing mapping data and sequence data produced in laboratories throughout the world. After major genome mapping and sequencing centers developed, a parallel data storage effort began when the individual genome centers developed dedicated in-house databases to store mapping and sequencing data produced in their own laboratories. The data from the publically funded genome projects were made freely available through the Web.
As genome data began to be produced in very large quantities, strenuous efforts were devoted to developing new genome databases (Table 1) and designing new software that would permit the huge amounts of mapping and sequence data and associated information to be searched in a systematic and user-friendly way. A major new focus on in-silico (computer-based) analyses made vital contributions to our understanding of the structure of genes and genomes.

Table1. SOME OF THE MAJOR EUKARYOTIC GENOME BROWSERS AND GENOME DATABASES
An important advance was the development of genome browsers with graphical user interfaces to portray genome information for individual chromosomes and sub chromosomal regions. Users of genome browsers can quickly navigate the sequence of a selected human chromosome moving from large scale to nucleotide scale, identifying genes and associated RNA transcripts in regions of interest, with exon-intron organization revealed as the user zooms in (see Figure 1 for an example). Thereafter, the user can click on features of interest to allow numerous connections to other databases and programs (permitting, for example, amino acid sequences to be obtained for selected transcripts, or evolutionary conservation of selected sequences by comparison with homologs in other genomes). As more and more information is obtained for genes and other functional units, more informative and precise gene annotation will be available in frequent, periodic updates of the genome browsers and databases.

Fig1. An example to illustrate using the Ensembl genome browser. Here the October 2016 version of Ensembl was queried with the human CFTR gene (the cystic fibrosis transmembrane regulator gene spans nucleotides 117,465,784–117,715,971 on chromosome 7). The two frames with the title “Genes” show exon–intron organizations of the different transcripts from the two DNA strands: the upper frame shows the CFTR sense transcripts, and the frame below the “Contigs” bar (which gives GenBank accession numbers for indicated DNA sequence contigs) shows CFTR antisense transcripts, plus transcripts from a neighboring, partially overlapping, protein-coding gene, CTTNBP2 (transcribed from the opposite DNA strand). In each case, exons are represented by short vertical bars that are connected by introns (flattened chevrons); protein-coding transcripts are represented by red- and gold-colored lines, noncoding transcripts by blue lines (according to the gene legend at bottom). There are five CFTR protein-coding transcripts (two full-length isoforms and three smaller isoforms), six noncoding sense transcripts, and two antisense transcripts. The “Regulatory Build” shows colored vertical bars representing the positions of indicated regulatory sequences. Clicking over individual items brings up additional information, as shown here by clicking on the CTTNBP2-011 transcript at bottom right (clicking on underlined items in blue font allows access to further information, often presented in further graphical frames).