Heterogeneous pre-training
Joint video and action learning from heterogeneous datasets with complementary body, root, and dexterous hand annotations.
Joint video–action learning with explicit whole-body supervision,
from heterogeneous motion data to humanoid task execution.
Paper link coming soon



WB-WAM incorporates body, root, and dexterous hand supervision into generative video pre-training. A three-stage training framework combines heterogeneous pre-training, retargeted PICO demonstrations, and robot task adaptation in a shared physical action space. We evaluate the model in simulation and on a physical humanoid, with additional studies of data efficiency and language-conditioned manipulation.
WB-WAM incorporates explicit whole-body action supervision into generative video pre-training, jointly learning visual dynamics and action trajectories in a shared physical action space.

Top: Progressive training integrates heterogeneous supervision, retargeted PICO motion, and real-robot demonstrations. Bottom: Video and action experts jointly model visual dynamics and whole-body actions conditioned on vision, language, and proprioception. Body and root references are executed through SONIC, while hand references directly control the finger joints.
View full-size figureJoint video and action learning from heterogeneous datasets with complementary body, root, and dexterous hand annotations.
Motion transfer from egocentric human demonstrations retargeted to G1, with complete whole-body action supervision.
Task-specific adaptation using robot demonstrations and auxiliary forward kinematics supervision.
Heterogeneous pre-training, human demonstrations,
and real-robot teleoperation.
Nine external datasets provide complementary supervision through video paired with body and hand motion, video paired with hand motion, and text-conditioned body motion.
Hover or select a source to explore the corpus.
Egocentric recordings pair video with body motion retargeted to G1 through GMR. Hand keypoints reconstructed with MINT are retargeted to Wuji joint configurations.
Task-execution data only; setup and reset periods excluded.
Demonstrations are collected through SONIC whole-body teleoperation, using PICO tracking for body and root references and MANUS gloves for Wuji hand control.
Task-level videos accompany
the quantitative evaluations.
Video unavailable.
HumanoidArena evaluation with SONIC as the execution backend.
Two camera views of each simulation rollout. Switching views preserves playback progress.
Simulation benchmarks, physical execution,
and task-aligned PICO transfer.
WB-WAM exceeds the strongest reported SONIC baseline on all seven HumanoidArena tasks, spanning locomotion, posture adjustment, and object interaction.
Select a legend label to show or hide a model.
Each task is compared with its strongest reported SONIC baseline. Mean SR lists the individual model averages across all seven tasks. Values are shown as mean ± standard deviation.