Learning Latent Protein Languages for Autoregressive Generation
Learned latent protein languages improve autoregressive sequence and structure generation versus amino-acid tokens.
Autoregressive transformers remain comparatively weak at protein sequence and structure generation on raw amino-acid tokens or backbone coordinates. The authors introduce Protein Latent Language, a 4,096-state contextual alphabet from a frozen ESM-2 encoder, and Structure Latent Language, adapted from GCP-VQVAE Lite. Matched training gives PLLM a compute-scaling exponent of 0.038 versus 0.020 for an amino-acid model, and cuts low-entropy samples by 54%. SLL reduces sequence-to-structure validation perplexity by 34%, and latent-token sampling is about 1,000 times faster than MSA-based AlphaFold2.
- PLL maps each residue to one of 4,096 contextual tokens from frozen ESM-2.
- PLLM's fitted compute-scaling exponent is 0.038 versus 0.020 for amino-acid autoregression.
- Unconditional generation cuts sub-1.5-bit entropy samples by 54% versus the amino-acid model.
- Replacing GCP-VQVAE Lite with SLL lowers best validation perplexity by 34%.
- Latent-token sampling measured about 1,000 times faster than MSA-based AlphaFold2.
Full article248 words · extracted from huggingface.co · click to collapse
Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework. We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates. We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model. In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training. For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample. These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.03978