Hugging Face daily papers·3d agoAV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation#diffusion#reinforcement-learning#audio-video-generation
Hugging Face daily papers·3d agoAgent-Editing World Model: Rethinking World Modeling for LLM Agents#llm-agents#world-model#state-editing 2 sources
arXiv cs.AI / cs.LG / cs.CL·3d agoFine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following#llama#fine-tuning#machine-translationAI research
The Decoder·3d agoAnthropic engineer explains why Claude's writing got worse although the model got smarter#anthropic#claude#opus 2 min
Hugging Face daily papers·5d agoHarness-Zero: Harness Distillation via Agent-as-Harness#harness-zero#agents#harness-distillation 2 sources
Hugging Face daily papers·6d agoCircuit Hypernetworks for Quantum-Augmented Diffusion Language Models#hyperq#quantum#diffusion-language-model
Latent Space·8d ago[AINews] Here are 6 Clones of Jev in 2 days#agents-md#claude-code#coding-agents 11 min1
arXiv cs.CR·9d agoOrigin Is All You Need: Provenance-Aware Transformers for Structural Trust-Boundary Separation#indirect-prompt-injection#provenance#transformersAI safety & security
Hugging Face daily papers·9d agoFrom Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention#fine-tuning#foundation-models#long-horizon-tasks
arXiv cs.AI / cs.LG / cs.CL·9d agoHow Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?#cfd#distribution-shift#fine-tuningAI research1
arXiv cs.AI / cs.LG / cs.CL·9d agoSTR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite Networks#fine-tuning#intent-understanding#leoAI research
SecurityWeek·9d agoAI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals#agentic-self-modification#ai-safety#fine-tuning 4 min1
The Register · Security·10d agoAI agents can modify themselves without humans telling them to do so#agentic-self-modification#agent-security#ai-agents 3 min1
arXiv cs.AI / cs.LG / cs.CL·10d agoScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments#agents#code-repair#environmentsAI research1
Latent Space·11d agoCan Skills Learned in Games Transfer to Real-World Work?#agents#benchmarks#diplomacy 7 min
Hugging Face daily papers·11d agoBI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence#benchmark#business-intelligence#fine-tuning
arXiv cs.AI / cs.LG / cs.CL·12d agoInoculation Midtraining with Learned Neologisms#alignment#fine-tuning#llmAI safety & security
OpenAI News·12d agoHow Fyxer built an AI executive assistant people trust#ai-assistant#dpo#fine-tuning 5 min
arXiv cs.CR·13d agoPick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks#backdoor#data-poisoning#fine-tuningAI safety & security1
Hugging Face daily papers·15d agoTraining Specialist Models without Reasoning Trajectories for Domain Expert Distillation#domain-adaptation#fine-tuning#knowledge-distillation
Hugging Face daily papers·15d agoDrift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models#fine-tuning#instruction-tuning#optimization
Hugging Face Blog·17d agoAsync GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL#distributed-training#fine-tuning#grpoAI tools & infra
Hugging Face daily papers·17d agoThe information geometry of large language models is shared, learned, and controllable#fine-tuning#information-geometry#interpretability1