WB-WAMHeterogeneous Body–Hand Pre-training
for Humanoid Loco-Manipulation

  1. Chuan Qin*
  2. Shaoting Zhu*
  3. Siyuan Luo
  4. Siqiao Huang
  5. Hongyu Zhao
  6. Hang Zhao†

Joint video–action learning with explicit whole-body supervision,
from heterogeneous motion data to humanoid task execution.

Paper link coming soon

Tsinghua University
Xiong’an Institute of Artificial Intelligence
The University of Melbourne

Overview

WB-WAM incorporates body, root, and dexterous hand supervision into generative video pre-training. A three-stage training framework combines heterogeneous pre-training, retargeted PICO demonstrations, and robot task adaptation in a shared physical action space. We evaluate the model in simulation and on a physical humanoid, with additional studies of data efficiency and language-conditioned manipulation.

1,880.2 h heterogeneous pre-training72-D physical action space3 stages from pre-training to adaptation
01 / Method

Whole-body pre-training.
A three-stage framework.

WB-WAM incorporates explicit whole-body action supervision into generative video pre-training, jointly learning visual dynamics and action trajectories in a shared physical action space.

WB-WAM architecture and three-stage training: heterogeneous pre-training, retargeted PICO mid-training, and robot post-training. Video and action experts model visual dynamics and physical references, with SONIC for body and root control and direct commands for the hands.

Top: Progressive training integrates heterogeneous supervision, retargeted PICO motion, and real-robot demonstrations. Bottom: Video and action experts jointly model visual dynamics and whole-body actions conditioned on vision, language, and proprioception. Body and root references are executed through SONIC, while hand references directly control the finger joints.

View full-size figure
Stage I

Heterogeneous pre-training

Joint video and action learning from heterogeneous datasets with complementary body, root, and dexterous hand annotations.

Stage II

PICO mid-training

Motion transfer from egocentric human demonstrations retargeted to G1, with complete whole-body action supervision.

Stage III

Robot post-training

Task-specific adaptation using robot demonstrations and auxiliary forward kinematics supervision.

02 / Datasets

WB-Datasets.
Data for each training stage.

Heterogeneous pre-training, human demonstrations,
and real-robot teleoperation.

Stage I

Heterogeneous pre-training data

1,880.2hours

Nine external datasets provide complementary supervision through video paired with body and hand motion, video paired with hand motion, and text-conditioned body motion.

Pre-training dataset composition totaling 1,880.2 hours: video with body and hand annotations, video with hand annotations, and text-conditioned body motion.
Composition of the nine-source pre-training corpus, grouped by available supervision.View original figure
Stage II

PICO whole-body motion data

Egocentric recordings pair video with body motion retargeted to G1 through GMR. Hand keypoints reconstructed with MINT are retargeted to Wuji joint configurations.

hours
22
tasks
73
episodes
13,396

Stage III

Real-robot teleoperation data

Demonstrations are collected through SONIC whole-body teleoperation, using PICO tracking for body and root references and MANUS gloves for Wuji hand control.

hours
3.37
tasks
8+2
episodes
1,011
frames
242,668
03 / Task demonstrations

Whole-body skills.
In simulation and the real world.

Task-level videos accompany
the quantitative evaluations.

Simulation tasks

Two camera views of each simulation rollout. Switching views preserves playback progress.

04 / Evaluation

Performance.
Beyond the demo.

Simulation benchmarks, physical execution,
and task-aligned PICO transfer.

HumanoidArena · SONIC
81.9%

Mean success across seven tasks.

WB-WAM exceeds the strongest reported SONIC baseline on all seven HumanoidArena tasks, spanning locomotion, posture adjustment, and object interaction.

Baseline results are reported by HumanoidArena under the SONIC setting.
Simulation task videos

Success rate by task

0–100% · higher is better

Select a legend label to show or hide a model.

HumanoidArena success rates for five modelsSeven task axes show mean success rates from zero to one hundred percent for ACT, DP, FM, pi0.5, and WB-WAM using SONIC. Exact values are listed below.
Task-wise comparison

WB-WAM and the strongest reported baselines

Success rate (%)
HumanoidArena SONIC setting. WB-WAM compared with the best reported baseline for each task.

Each task is compared with its strongest reported SONIC baseline. Mean SR lists the individual model averages across all seven tasks. Values are shown as mean ± standard deviation.

BibTeX

Method