Generative Cinematographer: Composing Camera and Object Motion in 3D
GenCine lets artists jointly control 3D camera and object motion in video via guidance maps on a Wan model.
Generative Cinematographer (GenCine) turns a single image into an editable 3D scene scaffold so artists can author a camera path and move foreground regions with local 3D handles. Multiple handles provide a piecewise-rigid approximation of non-rigid motion without a physics simulator or category-specific prior. Controls are projected into guidance maps that keep object motion in the same world coordinates as the background while the camera moves. A lightweight guidance branch and LoRA adapters trained on a pretrained Wan model, using recovered real-video controls and synthetic geometry, improve camera-relative motion and geometric consistency.
- GenCine lifts one image into an editable 3D scene for joint camera and object motion.
- Local 3D handles approximate non-rigid motion without a physics simulator.
- Guidance maps encode camera-relative positions in a shared world coordinate system.
- A lightweight guidance branch and LoRA adapters are trained on a pretrained Wan model.
Full article232 words · extracted from arxiv.org · click to collapse
Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02180