Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
IDiom, trained on 54 million predicted disordered regions, designs IDR sequences with feature-guided reinforcement learning.
IDiom is an autoregressive protein language model trained on IDiom-DB, 54 million predicted intrinsically disordered regions from the AlphaFold Database. It generates sequences that match natural composition, patterning, motifs, and predicted disorder. Reinforcement learning over sparse autoencoder features activates, on average, 90% of 30 targeted features across eight tasks, versus 24% for activation steering, and improves predicted localization and transcriptional activity. Code is public.
- IDiom-DB holds 54 million predicted IDRs curated from AlphaFold.
- RL-SAE activates about 90% of 30 targeted features versus 24% for steering.
- Generated IDRs improve predicted localization and transcriptional activity.
- Features tied to different functions can be combined in one sequence.
Full article220 words · extracted from arxiv.org · click to collapse
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02189