Humanoid Adaptation Framework

HAF

Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Langzhe Gu1,2*, Chengkai Hou1,2*, Meng Li2*, Xinhua Wang2, Jiaming Liu1, Xinyuan Lv2,3, Bowei Zhang2,3, Shuanghao Bai2,4, Guangrun Li1,2, Jingyang He1,2, Gaole Dai1, Ziluo Ding2, Zhiyuan Xu2, Kuan Cheng1, Jian Tang2✉, Zhengping Che2✉, Shanghang Zhang1✉

1 State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University 2 Beijing Innovation Center of Humanoid Robotics 3 Nankai University 4 Xi'an Jiaotong University

* Equal contribution    ✉ Corresponding authors

HAF-VLAHAF-Steer

Explore the framework

01 / Overview

Adapt Generalist VLA
To Humanoids.

Abstract

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action foundation models are not readily applicable to humanoid whole-body loco-manipulation. High-dimensional and interdependent whole-body motion makes it difficult for conventional single-stage architectures to coordinate locomotion, waist posture, and dual-arm manipulation, while policies trained purely through offline behavior cloning can remain suboptimal during real-world deployment.

HAF, the Humanoid Adaptation Framework, combines two complementary components. HAF-VLA splits full-body action denoising into three sequential stages and preserves kinematic dependencies with stage embeddings and cross-stage KV caches. On top of the frozen generator, HAF-Steer uses flow reversal and DCT compression to train a regularized SAC policy in a compact noise subspace, enabling efficient offline-to-online refinement without updating the large VLA backbone.

Overview of HAF real-world experiments, HAF-VLA, HAF-Steer, and their performance
Figure 1. HAF couples hierarchical whole-body generation with spectral latent policy adaptation.

02 / Method

Two Components:
HAF-VLA &
HAF-Steer

Part A

HAF-VLA

Hierarchical Action Flow

A shared action expert progressively expands the active whole-body action space. Clean actions from earlier stages are encoded as KV caches, conditioning later stages before a single final action chunk is deployed.

01

Locomotion + Head

02

+ Waist

03

+ Manipulation

HAF-VLA architecture with three hierarchical action generation stages
HAF-VLA. Cumulative action masks and cross-stage clean-action KV caches preserve the dependency from base stabilization to fine manipulation.

Part B

HAF-Steer

Spectral Latent RL

Expert action chunks are reversed through the frozen flow field, compressed to low-frequency DCT coefficients, and used to initialize a regularized offline-to-online SAC policy.

HAF-Steer flow reversal, DCT compression, reinforcement learning, and reconstruction pipeline
HAF-Steer. Flow reversal constructs expert spectral targets; mixed offline-online learning adjusts the low-frequency noise modes while the flow generator remains fixed.

03 / Real-world tasks

Seven tasks.
Whole-body coordination throughout.

Each task combines navigation, posture adjustment, and physical interaction. All task clips below intentionally play without sound.

01

Laundry Loading

Carry + bimanual loading

Silent
02

Clothes Retrieval

Navigate + retrieve

Silent
03

Table Tidy

Reach + rearrange

Silent
04

Basket Transfer

Lift + transport

Silent
05

Toy Storage

Pick + relocate

Silent
06

Ball Tossing

Align + throw

Silent
07

Box Transfer

Squat + whole-body carry

Silent

04 / Experiments

Clear gains in both VLA & RL adaptation.

HAF-VLA Performance

HAF-VLA enable long-horizon humanoid loco–manipulation beyond existing SOTA methods

Main results on seven long-horizon humanoid loco-manipulation tasks
TaskHAF-VLApi0.5GR00T N1.7CosmosACT
Laundry Loading66.753.340.00.010.0
Clothes Retrieval53.353.333.326.723.3
Table Tidy80.070.040.016.723.3
Basket Transfer63.350.043.333.316.7
Toy Storage80.053.330.040.023.3
Ball Tossing56.733.336.73.330.0
Box Transfer93.360.043.373.350.0
Average70.553.338.127.625.2

OOD robustness

HAF-VLA generalize to unseen visual and positional disturbances

With an unseen chair near the Laundry Loading path, HAF-VLA scores 40.0% versus 26.7% for pi0.5. With a 20 cm backward start shift in Clothes Retrieval, it scores 43.3% versus 36.7%.

Object distraction and position disturbance generalization experiments
Controlled visual and positional disturbances used for the HAF-VLA generalization study.

Offline to online

HAF-Steer improve real-world deployment performance through offline-to-online adaptation.

Across Toy Storage and Basket Transfer with both pi0.5 and HAF-VLA backbones, offline noise behavior cloning provides a strong initialization and offline-to-online SAC delivers further adaptation.

HAF-Steer success rates on Toy Storage and Basket Transfer in in-distribution and out-of-distribution settings
Successful trials out of 10. DSRL training was terminated in unsafe exploration settings, as noted in the paper.