LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
LynnReal-Omni unifies controllable video generation tasks in a 32B multimodal diffusion transformer, with a 27B Flash variant rendering 540p clips in 377 ms.
LynnReal-Omni is a native multimodal video generation framework built on a 32B shared multimodal diffusion transformer unifying text-to-video, image-conditioned generation, reference guidance, structural control, editing, restoration and long-video generation, accepting heterogeneous inputs like 3D renders and game recordings for agentic visual workflows. A dedicated 27B Flash model enables real-time rendering, producing a 22-frame 540p video in 377 ms on one H100 versus 843 ms for the full model. The work introduces a curated multi-shot audiovisual data pipeline and MSAVP, a 100-prompt, 20-metric evaluation design covering instruction following, plausibility, visual quality, temporal behavior and audio coordination.
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
AgenticGen applies DPO and GRPO reward-guided reasoning to ad video generation, improving TikTok CTR 2.72%, CVR 2.63%, and Advv 9.61%.
AgenticGen decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, supervised by online business feedback. It learns a performance-based reward from accumulated online feedback plus a rubric-based reward aligned with human quality standards, then optimizes policies with DPO followed by GRPO using process and outcome rewards. Online A/B experiments in the TikTok advertising system show CTR up 2.72%, CVR up 2.63%, and Advv up 9.61% over an SFT baseline.
The Republican Nominee for New York Governor Made a Creepy, AI-Generated Video of Mamdani and Hochul
NY gubernatorial nominee Bruce Blakeman posted an unlabeled AI-generated video of Mamdani and Hochul, prompting a Board of Elections complaint.
Bruce Blakeman, the Republican nominee for New York governor, posted an AI-generated video depicting New York City Mayor Zohran Mamdani and Governor Kathy Hochul gardening and cycling together. The New York State Democratic Party filed a complaint with the Board of Elections alleging the video violates state election law requiring disclosure of materially deceptive AI-generated political media. Blakeman argues the video is protected satire or parody under the law's exceptions.
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
Vidu S2 delivers real-time 720p interactive avatar generation plus real-time video editing with style, clothing, character, and background replacement.
Vidu S2 comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. S2-Avatar supports real-time 720p generation, dynamic references updatable at any moment, and stronger instruction following such as dancing, while S2-Editing performs real-time style rendering, clothing replacement, character replacement, and background replacement. The work also explores real-time spatial video generation for both models, reports outperforming all baselines, and offers a playable online demo at vidu.com.
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
MovieGrid arranges long videos on spatial grids during post-training, generating 6.05x more shots than temporal packing with state-of-the-art cross-shot consistency.
MovieGrid is a multi-grid post-training paradigm that decomposes long videos into temporally ordered chunks arranged on a spatial grid for joint modeling, enabling cross-chunk information exchange. The authors build the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos, producing 54K grid videos paired with character-aware story prompts. Under the same token budget, MovieGrid generates 6.05x more shots than Temporal Packing in a 1,616-frame video. It achieves state-of-the-art intra-shot consistency of 0.9131 versus 0.8086 for HoloCine and inter-shot consistency of 0.5914 versus 0.5384 for StoryMem.
[AINews] Fal’s H3 Max Live breaks the infinite videogen barrier
Fal post-trained MiniMax H3 into a 'Max' variant with 35x-faster inference, enabling faster-than-realtime AI video generation and infinite streams.
Fal post-trained MiniMax's H3 model into a 'Max' variant and optimized it for its in-house inference engine, achieving roughly 35x the speed of the official endpoint. The optimization enables faster-than-realtime video generation, demonstrated by an infinite interactive AI-generated stream productized by levels.io. The roundup also notes Meta Muse Code's general availability with an SDK, open DeepSeek-V4-Flash-Vision-Exp weights, GLM-5.3-Flash's strong agentic cost/performance rankings, and Tencent's 770B-parameter Hy4 Preview MoE with 49B active parameters.
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
PhysStream enables mid-generation interactive control of physics-grounded video via structured scene memory and velocity-increment signals, reducing motion distribution distance 33%.
PhysStream is an autoregressive physics-grounded image-to-video model that maintains structured scene memory—positional maps and object tracking maps derived online from previously generated frames—and accepts fine-grained motion control via sparse velocity-increment signals encoding physical quantities. Training runs in two stages: a bidirectional model finetuned with motion-control conditioning, then a causal autoregressive model with structured scene memory. It supports interactive mid-generation control over multi-object tabletop rigid-body scenes, reducing motion distribution distance (FVMD) by 33% and trajectory error by 12% over the strongest baselines. Human evaluators preferred it in over 85% of in-the-wild comparisons.
DF26: We Cannot Tell Fake From Real Anymore
DF26 benchmark shows humans and state-of-the-art deepfake detectors perform near chance on videos generated by seven modern text-to-video models.
Researchers introduce DF26, a benchmark of 271 real and 2,420 fully synthetic videos created by seven modern video generation models, all depicting single-person public-speaking scenarios such as direct-to-camera recordings, official statements, and studio interviews. Human viewers and state-of-the-art deepfake detectors scored close to random chance at distinguishing fakes from real footage. The authors argue current evaluation protocols are insufficient and call for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
WarmBloodAban/Minimax-h3_Singularity — new model trending #22 on Hugging Face
Community fine-tune Minimax-h3_Singularity enhances MiniMax-H3 video generation with HDR quality, distant face restoration, and improved motion, trending #22 on Hugging Face.
Minimax-h3_Singularity is a community fusion fine-tune of the MiniMax-H3 multimodal video generation model, built from multiple checkpoints and refined with pruning and weight optimization. It supports Text-to-Video, Image-to-Video, Reference-to-Video, and Video-to-Video workflows in ComfyUI, and claims improvements in HDR clarity, distant face restoration, motion fluidity, and fantasy VFX. The authors recommend pairing it with the minimax_h3_ref2v_turbo_4step_v0.1 LoRA for four-step accelerated inference, and an online demo is available via RunningHub.
Meta expands subscription push with new AI-focused plans
Meta launched Meta One subscriptions ($7.99–$499/month) bundling expanded Muse AI image and video generation across Facebook, Instagram, and WhatsApp.
Meta introduced Meta One with consumer Core ($7.99/mo) and Premium ($19.99/mo) tiers plus business plans ranging from Essential ($14.99/mo) to Max ($499/mo). Subscriptions unlock expanded Muse Image, Muse Video, Restyle editing, and Meta Business Agent usage, following Meta's $14.3B investment in Scale AI. Appfigures data shows Instagram's daily subscription revenue averaging $1.2M and Facebook's $528K after the March Plus-tier launches, up 475% and 143% respectively. BNP Paribas forecasts $13.5B added revenue by 2028; Truist estimates $20B by 2030.
Adobe is trying to make its AI generators idiot-proof in Premiere
Adobe adds in-timeline generative media to Premiere, letting editors generate video and audio clips using Firefly, Veo, Runway, Luma, and Kling models.
Adobe's new Generative Media tool lets Premiere editors highlight empty gaps in the timeline and generate context-aware, editable video, sound effects, music, and soundscapes without leaving the project. Editors can choose among underlying models including Adobe Firefly, Google Veo, Runway, Luma, and Kling. Beta AI audio tools can separate overlapping speakers and duck music under speech, and an AI Assistant is coming to After Effects for plain-language project commands.
Man gets 15 years for extorting women with AI-generated porn videos
An Ohio man was sentenced to 15 years in prison for sextortion and cyberstalking numerous victims using AI-generated sexually explicit videos.
An Ohio man received a 15-year prison sentence for multiple cybercrimes. The offenses included sextortion and cyberstalking of numerous victims, with AI-generated sexually explicit content used against the victims. The case illustrates criminal justice outcomes for offenders using AI-generated imagery in extortion schemes.
OpenVDN/vdn-minimax-h3 — new model trending #12 on Hugging Face
OpenVDN releases VDN-H3, an open hybrid-attention video model on MiniMax H3 that renders a 14.4-second 768p clip in 11.23 seconds on 8 B200 GPUs.
VDN-Minimax-H3 (VDN-H3) adds a frame-wise linear attention branch plus two LoRA adapters to MiniMax H3, distilled into 8-step and 50-step variants. It generates 768p, 14.4-second clips in 11.23 seconds on 8 B200 GPUs (90.5 seconds on one H200) using 8 denoising steps. Weights (about 82 GB total, including the 72 GB H3 base), the optimized inference stack, and training code are fully open-source under the MiniMax H3 Community License, which excludes the EU, UK, Korea, and US.
Programmable World Model
Programmable World Model decouples executable world-state evolution from video generation, reaching 94% Count Accuracy and 98% State Accuracy on new CombatStateBench.
An agent translates natural-language instructions into executable programs specifying entity states and transition rules, executed by a lightweight engine that maintains an explicit, persistent global world state including off-screen entities. State-augmented 3D oriented bounding boxes are deterministically compiled into pixel-aligned spatiotemporal conditioning signals for a pretrained video model acting as the generative renderer. On the new CombatStateBench benchmark it achieves 94% Count Accuracy and 98% State Accuracy, substantially outperforming existing interactive video world models.