Does Learning Protein Folding Generalize to Broader Reasoning?
Fold2Reason post-trains language models on protein-structure QA, lifting reasoning across 10 benchmarks and raising macro-average accuracy from 45.09% to 48.33%.
The authors build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a post-training recipe combining discrete structural answers from the model's language head with continuous 3D geometry decoded from shared representations. On FoldBench, Fold2Reason reaches structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond proteins, it improves all 10 evaluated benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy by 3.23 points, while random, synthetic, and shuffled-structure controls yield smaller or negative gains.
- FoldingCorpus turns solved protein structures into exactly checkable QA supervision
- Fold2Reason scores 2.7-3.5x Qwen3.5-9B on FoldBench structure prediction
- Macro-average accuracy across 10 reasoning benchmarks rises 45.09% to 48.33%
- Matched random/synthetic/shuffled controls show much smaller or negative gains
Full article187 words · extracted from huggingface.co · click to collapse
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.38879