arXiv cs.AI / cs.LG / cs.CL·6d agoLESSER: Post-Training Data Selection with Output-Layer Gradients#llm#post-training#data-selectionAI research1
Hugging Face daily papers·12d agoSelecting Diverse SFT Traces Improves Post-RL Generalization#sft#reinforcement-learning#data-selection