Unsupervised Learning on 250m Proteins, Biological Structure and Function Emerge

Unsupervised Learning on 250m Proteins, Biological Structure and Function Emerge

To this end we use unsupervised learning to train a deep contextual language model on 86 billion amino acids across 250 million sequences spanning evolutionary diversity. The resulting model maps raw sequences to representations of biological properties without labels or prior domain knowledge. Learning on full sequence diversity rather than individual protein families increases recoverable information about secondary structure.

Source: www.biorxiv.org