Biological Function Emerges from Unsupervised Learning on 250M Protein Sequences
To this end we use unsupervised learning to train a deep contextual language model on 86 billion amino acids across 250 million sequences spanning evolutionary diversity. The resulting model maps raw sequences to representations of biological properties without labels or prior domain knowledge. Learning on full sequence diversity rather than individual protein families increases recoverable information about secondary structure.
Source: www.biorxiv.org