Abstract
Recent advances in 2D and 3D generative modeling have paved the way for 4D generation. However, existing methods predominantly focus on text-to-4D generation, emphasizing temporal smoothness and spatial consistency while often neglecting precise motion control. Furthermore, due to the inherent ambiguity of text descriptions, current approaches fail to generate specific and controllable motion patterns. To address this limitation, we propose Motion4D, a rigging-free framework that animates static humanoid meshes with arbitrary initial poses by replicating the motion pattern of a reference sequence. Crucially, Motion4D adopts a generate-then-animate paradigm: it first synthesizes multi-view videos to capture precise spatiotemporal dynamics, and then explicitly uses these videos to drive the input 3D mesh without altering its topology or vertex-face connectivity. To facilitate training, we curate Motion4D-Video, a large-scale multi-view dataset featuring a paired structure where diverse characters perform identical motions under consistent camera settings. Leveraging this dataset, we develop MoMuDiff, the first motion-controlled multi-view video diffusion model. MoMuDiff is conditioned on reference videos for motion cues and static images for appearance and geometry. MoMuDiff employs a two-stage training strategy to disentangle motion from appearance, enabling the generation of multi-view videos that faithfully replicate reference dynamics while effectively preserving the target's geometry. Building upon these generated videos, we further introduce a Gaussian-Proxied Mesh Animation algorithm, which utilizes a dense, topology-aware Gaussian field to robustly transfer the synthesized dynamics onto the target mesh while leaving its topology and vertex-face connectivity unchanged. Extensive experiments demonstrate that Motion4D achieves robust, controllable motion transfer from arbitrary starting poses, outperforming prior approaches in both fidelity and controllability.