SELECTION OF A REPRESENTATIVE SET OF STRUCTURES FROM BROOKHAVEN PROTEIN DATA-BANK
Reliable structural and statistical analyses of three dimensional protein structures should be based on unbiased data. The Protein Data Bank is highly redundant, containing several entries for identical or very similar sequences. A technique was developed for clustering the known structures based on their sequences and contents of alpha- and beta-structures. First, sequences were aligned pairwise.
