ZeroHour

Search: “diffusion-policy”

30 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

Researchers present Diffusion TV, a CRT-based installation where antenna manipulation lets audiences physically experience diffusion model denoising.

Diffusion TV is an interactive installation built around a modified CRT television where turning the antenna controls the clarity of AI-generated images and sounds, mirroring the denoising process of diffusion models. Three channels present AI-generated animals from the past, present, and future within a temporal and ecological narrative. The authors frame the work as an embodied, non-verbal alternative to explainable AI that highlights intermediate generative states rather than final outputs.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

Diffusion Models and Concept Formation

Paper argues diffusion models implicitly form Cobweb-like concept hierarchies, with a basic level emerging at intermediate noise levels.

The authors draw a formal correspondence between diffusion models and Cobweb, a classic incremental concept-hierarchy learner, noting both are hierarchical Bayesian density models with Gaussian prototypes. Modes of the diffusion model's noisy marginals form a hierarchy whose basic level sits at intermediate noise, where class identity commits. The correspondence is tested on MNIST and Fashion-MNIST via mode-finding. Diffusion is reframed as a cognitive model of concept formation.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Researchers propose Mask Forcing, a dual-noise masking rollout that mitigates mode collapse in autoregressive video diffusion distillation.

The paper targets over-saturation and over-smoothing in autoregressive video diffusion models distilled via Distribution Matching Distillation, attributed to reverse-KL mode-seeking behavior. Mask Forcing perturbs the student self-rollout with random masks along spatial and temporal axes, injecting cleaner tokens that act as denoising guidance for noisier tokens. Experiments show improved visual quality across multiple distillation methods without using real video data or extra post-training stages.

Hugging Face daily papers · 9d agoAI research

CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models

CanvasAnneal injects teacher reasoning traces into diffusion canvases during curriculum RL, improving diffusion LLMs on MATH500, Countdown, and Tau2.

CanvasAnneal is a curriculum-guided reinforcement learning framework for diffusion language models that addresses exploration bottlenecks in standard RL. It warm-starts exploration by injecting teacher-generated reasoning traces into the initial diffusion canvas, then gradually removes this guidance so the model generates reasoning trajectories independently. Across mathematical reasoning and tool-use benchmarks, it improves over standard diffu-GRPO on MATH500, Countdown, and Tau2 and accelerates reward improvement, though gains are task-dependent.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research1

Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport

Researchers introduce model-aware diffusion schedules via fiberwise optimal transport, cutting flow-matching FID on CIFAR-10 by 38.6% at 16 function evaluations.

The paper proposes constructing diffusion and flow-matching sampling schedules from a fiberwise prediction risk defined via optimal transport, combined with coefficient-path kinetic action, yielding a closed-form time allocation. Across DDPM and flow-matching experiments spanning targets, datasets, and architectures, the schedules beat model-agnostic baselines, including a 38.6% relative FID reduction for flow matching on CIFAR-10 at 16 function evaluations. Normalized fiberwise-risk profiles from independently trained models align closely, suggesting empirical universality, and a frozen analytic allocation template retains most of the gains.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

OptiFlow learns one-step multimodal flow policies for offline RL via state-wise entropic optimal transport, avoiding critic overestimation and mode collapse.

The paper introduces OptiFlow, a framework that frames one-step flow policy learning as a structured sample-allocation problem in offline reinforcement learning. It jointly trains a value-aware reference flow policy and a one-step policy, coupling action samples through state-wise entropic optimal transport where critic values set distillation priority and action-distance cost ensures geometrically compatible pairings. By avoiding direct critic maximization, it anchors the policy to high-value dataset-supported modes without out-of-distribution divergence. Code is released on GitHub and the method performs strongly across diverse offline RL benchmarks.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

OPRD distillation enables weak-to-strong generalization by amplifying verifier-supported policy updates, outperforming existing RL and distillation methods with fewer student updates.

On-Policy Reverse Distillation (OPRD) evaluates a weak teacher's policy shift relative to its reference policy on student rollouts and amplifies the verifier-supported component of the student's policy gradient. This rescaling preserves the stationary points of policy optimization while letting the student learn beyond the teacher's capacity ceiling. In successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches, and response-style analysis shows students remain closer to verifier-RL-trained models than to their weak teachers.

Hugging Face daily papers · 9d agoAI research

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

Introduces DRIFT, a black-box attack removing diffusion watermarks by deflecting generative trajectories, achieving 98-100% success across nine watermarking schemes.

Researchers propose DRIFT, a black-box watermark removal attack combining partial forward diffusion with stochastic reverse resampling to break trajectory-dependent verification. The paper derives information-theoretic and Wasserstein source-dependence bounds and shows the first verifier-rejected rung is least distorted among rejected rungs. Across nine watermarks spanning three paradigms, DRIFT achieves 98-100% attack success with the best image quality among compared attacks, without secret keys, verifier internals, or per-image gradient optimization.

arXiv cs.CR · 9d agoResearch

Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning

Review connects control theory, optimal transport, probabilistic inference, thermodynamics, and machine learning via free-energy optimization under constraints.

The review unifies five fields: control theory, optimal transport, probabilistic inference, non-equilibrium thermodynamics, and machine learning. The common conceptual thread is optimization of free-energy-like functionals under dynamical or statistical constraints. Selected applications are presented in reinforcement learning, variational inference, and generative modeling. The tutorial-style text assumes no prior familiarity and begins from physics principles.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

MInTRL: Off-policy Intervention can boost On-policy RL

MInTRL injects sparse judge corrections into on-policy RL rollouts, expanding exploration beyond on-policy sampling while preserving learnability on math and code benchmarks.

Minimal Intervention Reinforcement Learning periodically has a judge-intervention policy replace erroneous suffixes of the current policy's output with short corrections, then returns control, keeping trajectories largely on-policy. Training uses a sequence-level advantage-regression objective that removes the need for importance sampling. Across math and code benchmarks it consistently beats standard on-policy and off-policy baselines, remains effective with self-intervention, and performs best at moderate intervention intensity.

Hugging Face daily papers · 6d agoAI research

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement

Generalized Agent Iteration formally unifies iterative policy improvement and recursive self-improvement, defining axes that distinguish anchored, goal-drifting, and self-referential agents.

The paper proposes Generalized Agent Iteration (GAI), a formal framework that models learning as a cycle of agent evaluation and agent improvement, defining the agent as a configuration of modifiable components. Two dials—whether the improving mechanism is part of the agent and whether the evaluation standard is grounded outside it—separate generalized policy iteration (GPI) from recursive self-improvement (RSI) and classify systems as anchored, goal drift, or fully self-referential. The framework places existing systems on shared axes and makes defects of recursive self-improvement statable one condition at a time.

Hugging Face daily papers · 6d agoAI research1

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Attention-DP3 adds spatially object-aware attentional conditioning to 3D diffusion policies, improving robotic manipulation by up to 31% under heavy clutter.

Attention-DP3 injects object-level geometric cues into the unchanged DP3 diffusion policy via Tri-field Attentional Conditioning, using targetness, intra-target saliency, and backgroundness fields. Open-vocabulary 2D segmentation masks are lifted to 3D with calibrated camera geometry to build object-centric priors. Experiments on Adroit, DexArt, MetaWorld, and a real-world SO101 platform show state-of-the-art results, outperforming DP3 by up to 31% under heavy distractor clutter; the code is publicly available on GitHub.

Hugging Face daily papers · 7d agoAI research

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

A review paper frames on-policy self-distillation collapse as governed by three levers: token weighting, privileged information, and guidance decay.

The paper critically reviews On-Policy Self-Distillation (OPSD), where a language model trains on its own generations scored token-by-token by a teacher conditioned on privileged information such as reference solutions or environment feedback. It identifies collapse, the progressive narrowing of producible reasoning paths, as the dominant failure mode and analyzes it through three levers: signal weighting, the nature of privileged information, and teacher dynamics. The review is restricted to mathematical reasoning, reports no new experiments, and offers a shared vocabulary separating settled findings from disputed ones.

Hugging Face daily papers · 22d agoAI research

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

LLaDA-UI, a 16.7B block-wise diffusion vision-language GUI agent, outperforms Qwen2.5-VL-7B and beats Qwen3-VL-8B on four of six GUI benchmarks.

LLaDA-UI is a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent built on the LLaDA2.0-mini-base diffusion language backbone with a native-resolution vision encoder. It uses a two-stage pipeline: general multimodal pre-training followed by GUI-agent supervised fine-tuning on mobile, desktop, web, and grounding data. It substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks, establishing block-wise diffusion as a practical paradigm for latency-sensitive multimodal agents.

Hugging Face daily papers · 8d agoAI research

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Uno pairs autoregressive LLMs with lightweight diffusion weights to draw multiple tokens in parallel, delivering up to 3x lossless speedup without a draft model.

The paper introduces diffusion-augmented LLMs: autoregressive weights trained with the standard next-token objective plus lightweight diffusion weights trained via a Diffusion Distillation phase to emit multiple tokens in parallel. Psi-Spec samplers enable lossless acceleration without the separate draft model required by speculative decoding. The 8B Uno model outperforms the 26B open DiffusionGemma and proprietary Mercury 2 on agentic tool use, coding, and long-context reasoning benchmarks, with up to 3x throughput gains over the base model at all evaluated batch sizes. Code and checkpoints are released publicly.

Hugging Face daily papers · 14d agoAI research

Register Tokens for Bounded-State Reasoning in Diffusion Language Models

Register tokens let diffusion language models like LLaDA and Dream carry reasoning state across cleared chunks, gaining up to 19.5 points on code.

Researchers propose register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks in masked diffusion language models. After decoding and clearing a chunk, the model continues from the prompt and the carried register state instead of retaining earlier text. On LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation and can be further refined with reinforcement learning on long-horizon reasoning tasks.

Hugging Face daily papers · 3d agoAI research

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

EvolveTrade lets LLM trading agents self-refine their tool-use policy from realized portfolio feedback, improving Sharpe ratios.

EvolveTrade treats a tool-using trading agent's system prompt as a text-parameterized policy that a Policy Agent revises after each update interval using accumulated decision traces and realized portfolio feedback, keeping the backbone LLM fixed. Experiments across multiple market regimes and two LLM backbones show improved Sharpe Ratio and Cumulative Return over fixed-policy baselines in most settings. Behavioral analyses show evolved policies increase code-mediated analysis and activate regime-relevant computations, with case-level attributions linking policy changes to returns.

Hugging Face daily papers · 2d agoAI research

Double descent is the principle of least action

A statistical mechanics analysis explains double descent: finite-time diffusion induces effective weight decay that regularizes models as parameters grow.

The paper models stochastic gradient-based training as a particle diffusing over the training-loss energy landscape at an induced temperature, sampling parameters via a Boltzmann distribution. Finite training time carries an effective weight decay, making every parameter a quadratic degree of freedom governed by the equipartition theorem. Adding parameters at fixed training loss lowers the temperature and the L2 norm of the stationary path, increasing effective regularization and explaining the double descent phenomenon.

arXiv cs.AI / cs.LG / cs.CL · 12h agoAI research

MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling

MCRL2 augments reinforcement learning with multi-resource cross-attention representations to improve cloud microservice scheduling and load balancing.

MCRL2 combines a multi-resource cross-attention representation learning module (MCRL) with an actor-critic architecture and maximum entropy objective for microservice scheduling. The approach captures interdependencies among nodes, resources, and microservices in data centers. Experiments on real production cluster traces show improvements in load balancing, scheduling success rate, and average completion time versus baselines.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research1

Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging

LP-BTS uses graph proposal policies, learned critics, and budgeted PUCT search to plan mobile charging across dynamic action spaces up to 2,813 stops.

LP-BTS is a learning-guided planning architecture for one-to-many mobile charging, where N=250 sensors induce roughly 1,125 initial candidate charging stops. A graph proposal policy concentrates candidate support, a learned value critic evaluates leaves, and edge-budgeted PUCT compares simulated futures, letting a single frozen checkpoint cover action universes from 736 to 2,813 stops. On a sealed 30-scenario confirmatory bank it attains the highest observed survival (0.4545) and alive-AUC (0.8031), though its +0.0066 survival edge over the strongest engineered comparator is statistically unresolved.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

VC-Attention is a training-free low-bit attention method for diffusion transformers, achieving 1.46-1.59x kernel speedups on datacenter GPUs with higher fidelity.

VC-Attention is a training-free low-bit attention framework for diffusion transformers that pairs V-Smooth value smoothing via lightweight online clustering with ExpCast-FP8, which maps log-domain scores directly to E4M3 FP8 probability codes and eliminates the FP32 softmax exponential. It is implemented for B200, B300, H200, RTX PRO 6000, and RTX 5090 GPUs. Across Wan2.2, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3, it improves fidelity over low-bit baselines and speeds attention 1.46-1.59x over BF16 FlashAttention-4 on datacenter Blackwell and Hopper GPUs and 2.3-3.6x on workstation cards, with 1.13-1.70x faster end-to-end clip generation.

Hugging Face daily papersupdated · 4h agofirst · 3d agoAI research 2 sources

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

Researchers prove off-policy evaluation under history-dependent logging requires exponentially many episodes, resolving a hardness question for model-based POMDP evaluation.

The paper constructs POMDPs with at most two latent states per stage, three actions, and a three-memory-state logger where evaluating a known deterministic target policy to accuracy 1/8 requires Θ((3/2)^H log(1/δ)) episodes for any horizon H≥3. Coverage and outcome-revealing conditions hold with constants independent of H, yet a reset erases the unknown transition that determines the target value. The authors characterize the resulting statistical experiment exactly, derive a matching optimal estimator, and validate predictions on a two-lane gridworld. This settles the history-dependent-logging, model-based case posed by Zhang and Jiang (arXiv:2503.01134).

arXiv cs.AI / cs.LG / cs.CL · 11h agoAI research

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

Researchers propose ERPO, enabling test-time reinforcement learning for code generation via probe-executed consensus rewards, rank masking, and entropy regularization.

The paper introduces probe-driven test-time reinforcement learning (TTRL) for code generation, where output-free probe inputs are constructed from problem statements and candidate programs are executed on them to compute a Probe Consensus Reward (PCR). Because PCR can be gamed through spurious consensus, the authors propose Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which turns low-PCR outcomes into conservative negative updates via rank masking and constrains policy drift with an entropy ceiling. On coding benchmarks, ERPO substantially improves pass@1 and pass@k in both in-domain adaptation and zero-shot transfer.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research1

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Research shows on-policy expert correction, not imitation fine-tuning, lets weaker agent models catch up under evolved harnesses.

Researchers study how to combine automated agent-harness evolution with lightweight fine-tuning across seven enterprise agent tasks. Naively training weaker models (Qwen3-Coder, Gemma 4) on expert trajectories under an evolved harness regressed performance by 4 to 30 points on all tasks. They propose an on-policy correction pipeline, automated by a meta-level MLE agent, where an expert rewrites only the failing turn of the weaker model's rollout, preserving model-harness fit.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research

Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

Theorists prove multi-step lookahead RL planning is NP-hard for every fixed rational discount factor yet give a randomized polynomial-time approximation scheme.

The paper resolves open questions about reinforcement learning with multi-step transition lookahead. It shows exact planning remains NP-hard for every fixed rational discount factor in (0,1), not just discounts arbitrarily close to one, and introduces a randomized polynomial-time approximation scheme for every fixed lookahead depth. Extending to unknown transitions and stochastic rewards via optimism and variance-adaptive confidence bounds, the algorithm achieves cumulative regret matching classical tabular discounted RL up to logarithmic factors.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research1

Rebuilding AUTOMATIC1111 with Gradio Workflow

Hugging Face demonstrates rebuilding the AUTOMATIC1111 Stable Diffusion web interface using its Gradio Workflow framework.

Hugging Face published a post showing how to rebuild the AUTOMATIC1111 Stable Diffusion WebUI experience with the Gradio Workflow framework. The article body was unavailable, so details beyond the title are limited, but the piece appears to be a tutorial on composing interactive AI interfaces with Gradio Workflow components.

Hugging Face Blog · 7d agoAI tools & infra

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

Marigold V2 adapts diffusion transformers for monocular depth estimation, improving AbsRel 16-26% over the previous best on KITTI and ETH3D.

Huawei's Bayer lab revisits the Marigold approach to repurpose image generation and editing models built on the diffusion transformer (DiT) architecture into monocular depth estimators. The recipes target single-step inference from pretrained multi-step flow-matching models, with remedies including alignment to ground-truth semantic features and a two-stage fine-tuning protocol using a Sinkhorn-based loss. The resulting model produces crisper depth maps that generalize out-of-distribution and also achieves state-of-the-art results on surface normals estimation and intrinsic image decomposition.

Hugging Face daily papers · 9d agoAI research

Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

Drift-Constrained Optimization reformulates fine-tuning as update-direction selection, letting Qwen3 models improve target tasks within a behavioral drift budget.

The paper specifies a behavioral drift budget before optimization and shows that update direction is the remaining degree of freedom, reformulating fine-tuning as a direction-selection problem. In a stringent QA-only setting where instruct models must still generate multi-step reasoning at inference, a coarse layer-selective probe reverses the failure of QA-only fine-tuning. Across Qwen3-8B and Qwen3-14B, these directions substantially improve scientific reasoning and multilingual translation, matching or outperforming dedicated translation systems over 100+ languages and giving stronger initialization for reinforcement learning.

Hugging Face daily papers · 5d agoAI research

Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

Hugging Face blog describes running async GRPO reinforcement learning with LoRA across HF Jobs using a storage bucket and proxy instead of NCCL.

A Hugging Face blog post titled 'Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL' explains an asynchronous Group Relative Policy Optimization training setup using LoRA adapters distributed across Hugging Face Jobs workers. The architecture coordinates training through an object storage bucket and a proxy server, removing the need for NCCL collective communication. No full article text was available at classification time.

Hugging Face Blog · 7d agoAI tools & infra

Epsilon-Nash Equilibria in History-Dependent SA-MDPs

Researchers give the first algorithm for computing epsilon-approximate history-dependent equilibria in state-adversarial Markov decision processes with observation-perturbing adversaries.

The paper studies state-adversarial Markov decision processes (SA-MDPs) where an adversary knowing the true state perturbs observations within state-dependent proximity sets each step. The authors prove universal history-dependent equilibrium policies do not exist and reduce SA-MDPs to a strategically equivalent constrained zero-sum one-sided partially observable stochastic game, enabling the first algorithmic route to epsilon-approximations of initial-state dependent equilibria. The algorithm is validated on small analytically verifiable games and scales to larger benchmarks, including Atari Freeway rollouts with a 12-period-ahead horizon.

arXiv cs.CR · 13h agoAI safety & security