4Director

Controlling Video World Models with Rigid 3D Geometry

1Stability AI 2University of Illinois at Urbana-Champaign
Stability AI University of Illinois at Urbana-Champaign
Object Motion Rigid 3D trajectories
New Object Insertion Reference-guided composition
Joint Camera and Object Control Jointly authored trajectories

TL;DR: 4Director uses complete object meshes and rigid trajectories in a shared 3D space to jointly control the camera, existing objects, and inserted objects in video world models.

Abstract

Precise control over camera and object motion is essential for professional video production. Existing methods support either camera control alone or coarse object control, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

Method

4Director pipeline: an input image becomes rigid 3D geometry with user-defined object and camera trajectories; depth rendering drives a Motion Adapter alongside a frozen Wan-DiT to produce the generated video.

Qualitative Comparison

What happens if the object leaves and returns?

Input Image
Input image of a camel in an enclosure
Point CloudInteractive
4Director (Ours)

Previous Works Native control above, generated video below

MotionControl
Perception as Control
SymphoMotion
VerseCrafter