ZeroHour

Search: “zero-shot”

3 stories in the last 24h

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

Pipeline pairs generated video with audio-derived force profiles to enable zero-shot, force-aware robot manipulation on Franka Panda for contact-rich tasks.

The paper leverages generated video and audio jointly: loudness of generated contact sounds shapes a bounded, time-varying desired-force profile from a natural-language task prompt. Trajectories execute on a Franka Panda robot with a closed-loop force regulator tracking the audio-shaped profile, succeeding where a kinematic-only baseline fails. The pipeline also serves as a data generation engine to train closed-loop manipulation policies.

arXiv cs.AI / cs.LG / cs.CL · 14h agoAI research

Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out

Google Research introduced R4T, an RL-trained fan-out pipeline distilled into a 53.9M-parameter diffusion retriever achieving 12x-20x faster query fan-out.

Google Research introduced Retrieve-for-Train (R4T), which trains a fan-out language model with GRPO plus soft PPO regularization, then distills query fan-out into a 53.9M-parameter diffusion transformer that generates all retrieval embeddings in a single non-autoregressive pass. A three-term reward (groundedness 0.6, diversity 0.2 via Vendi Score, alignment 0.2) prevents paraphrastic collapse and reward hacking during training. On the Polyvore dataset, Gemma3-4B R4T-FOLM averaged 49.1 versus 40.9 for Best-of-N, and the diffusion retriever cut fan-out latency from 1.46s to 0.07s at batch size 8, a consistent 12x-20x speedup over autoregressive methods.

MarkTechPost · 1h agoAI research

Robot Visions: Breaking reCAPTCHA at Zero Cost and Zero Shot

Researchers defeat Google reCAPTCHA using free local models CLIP and OWLv2, achieving 92.6% per-session success at zero cost.

The paper taxonomizes Google reCAPTCHA challenges into Type A (independent tiles) and Type B (4x4 grid) and builds zero-shot, training-free solvers from open-source local models. CLIP solves 58% of Type A challenges and OWLv2 43.5% of Type B, while an end-to-end automated solver succeeds on 92.6% of 500 real-world sessions. The authors also show a non-technical adversary can solve challenges using natural-language instructions to a commodity AI assistant, collapsing the attacker skill floor and suggesting visual challenge CAPTCHAs have reached the end of their useful life.

arXiv cs.CR · 20h agoAI safety & security