Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
Synthetic document perturbations train LLMs to reason from context, lifting Qwen3.6-35B-A3B to 24.6% on CL-bench.
The paper builds a synthesis pipeline that rewrites public documents, generates document-dependent questions and rubrics, answers them in context, and keeps only samples that genuinely depend on the document. About 10,000 samples from 3,500 documents train a student without human annotators. Supervised fine-tuning raises Qwen3.6-35B-A3B from 13.7% to 22.8% on CL-bench, and rubric-reward reinforcement learning reaches 24.6%, near Qwen3.8-2.4T at 23.9%. Gains transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
- Pipeline rewrites documents to reduce memorization and require context.
- About 10,000 samples come from 3,500 documents without human annotators.
- SFT lifts Qwen3.6-35B-A3B from 13.7% to 22.8% on CL-bench.
- Rubric-reward RL reaches 24.6%, near Qwen3.8-2.4T at 23.9%.
Full article259 words · extracted from huggingface.co · click to collapse
Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33642