Project page

FWBC-VLA: Force-Aware
Whole-Body Compensation
for Contact-Rich Loco-Manipulation

A force-aware VLA framework that connects sensorless contact estimation with whole-body compensation for wheel-legged robots.

Sensorless force estimation Force-conditioned VLA Whole-Body Loco-Manipulation
Yutian Zhang1,* Siyuan Ma4,* Liwen Yang1 Yang Li2 Ce Hao5 Haozhen Chi6 Dong Wei3,† Qiaojun Yu2,† Dibo Hou1,†
1Zhejiang University 2Shanghai Artificial Intelligence Laboratory 3Deep Robotics 4Tsinghua University 5Zhongguancun Academy 6Zhejiang University of Science and Technology

*Equal contribution †Corresponding authors

Demo

Abstract

FWBC-VLA real-world door opening and blackboard cleaning task setup.
Contact-rich evaluation scenarios covering door opening with a closer, handle pushing, continuous blackboard pressure, and eraser pickup.

Contact-rich loco-manipulation requires a bridge that connects semantic action generation and physical interaction regulation. Existing VLA models generate task-level actions from visual and linguistic observations, but lack interpretation of how generated actions interact with the physical world. While the whole-body control policy can stabilize the robot, it is unable to distinguish between forces originating from task-relevant contacts and those caused by external disturbances during manipulation.

We propose FWBC-VLA, a force-aware loco-manipulation VLA framework for wheel-legged robots that bridges task-level action generation and compensation for WBC. HSR-Force estimates contact intensity and temporal contact trends without dedicated force/torque sensors. These estimates are encoded as force tokens for action decoding and are also used by a compensation generator to produce corrective whole-body actions. Real-world experiments on whiteboard wiping and door opening with a door closer demonstrate the effectiveness of FWBC-VLA in contact-rich loco-manipulation.

Motivation

Contact Perception Is Missing State in Whole-Body Control

Comparison of kinematic VLA, arm-centric force-aware VLA, and FWBC-VLA paradigms for loco-manipulation.
Figure 1. FWBC-VLA turns contact from an unobserved consequence of motion into a shared physical signal for both task action and whole-body control.

In contact-rich loco-manipulation, contact is not an auxiliary sensing modality. It is a latent physical state that determines both whether the task is progressing and whether the robot can remain stable while executing it. The same end-effector load that enables wiping or overcomes a door closer also propagates through the arm and perturbs the mobile body.

Yet existing whole-body VLA stacks communicate mainly through kinematic commands. The VLA can prescribe where to move without knowing how contact is evolving, while the WBC can reject disturbances without knowing which forces are task-relevant. A robot may therefore follow the intended trajectory and still fail physically by losing contact, applying the wrong load, or drifting at the base.

The central challenge is not simply to add force as another policy input, but to recover a causal interaction state and make it actionable across the control hierarchy. FWBC-VLA estimates this state without dedicated F/T sensors and uses it to close two coupled feedback loops: one adapts task-level action generation, while the other compensates the body-level consequences of the same contact.

  • One interaction, two consequences: contact drives task progress and body disturbance at the same time.
  • One shared signal, two closed loops: action adaptation and whole-body compensation remain physically coordinated.
  • A necessary correction under sustained load: body compensation adds 24.5 percentage points on average, including +52 points for door pushing with a closer and +44 points for board cleaning.

Overview

FWBC-VLA builds a sensorless force-aware interface between the VLA policy and whole-body controller. The same interaction representation is shared by two pathways: force-conditioned task action generation and task-relevant body compensation. This keeps the pretrained VLA backbone focused on semantic action decoding while giving the controller physical feedback for sustained contact.

FWBC-VLA framework connecting sensorless interaction estimation, VLA action generation, and whole-body compensation.
FWBC-VLA introduces a sensorless force-aware interface between the VLA policy and the whole-body controller.

Method

The method has three tightly connected parts. HSR-Force estimates residual interaction from robot proprioception at high rate, the VLA action expert receives short force-token histories, and a compensation generator converts body-frame load cues into bounded base corrections. The downstream WBC still handles balance and low-level tracking.

Interaction estimation

HSR-Force

A 200-Hz sensorless estimator predicts free-motion torque, preserves directional residual joint torque, and summarizes contact intensity and temporal trends.

Action generation

Force-conditioned VLA

Short histories of interaction estimates are encoded as force tokens and late-fused into the action expert during loco-manipulation action decoding.

Whole-body control

Body compensation

Projected load proxies and base proprioception feed a compensation sidecar that generates bounded corrective actions for sustained physical contact.

WL&Arm Dataset

WL&Arm captures contact-rich wheel-legged loco-manipulation through Pico-based force-intent teleoperation. It pairs high-rate proprioceptive recordings for contact estimation with low-rate real-robot demonstrations for VLA action learning. The dataset covers bottle pick-and-place, whiteboard wiping, and door opening, with force-intent annotations kept as metadata rather than deployed policy inputs.

Dataset coming soon
WL&Arm dataset composition and Pico-based teleoperation controls.
Dataset composition and data collection interface. Force-intent commands are retained as metadata and inspection references, not as deployed policy inputs.
5,000+

teleoperation episodes

200 Hz

joint data for contact estimation

15 Hz

loco-manipulation demonstrations

Force intent

Pico-based tangential and normal command metadata

Results

Experiments evaluate sensorless force estimation, real-world loco-manipulation, and the contribution of force-conditioned compensation. HSR-Force responds near contact onset and remains active during sustained door loading. In whiteboard wiping and door opening, FWBC-VLA improves final-stage success most strongly during contact-intensive execution.

Real-world contact-rich loco-manipulation experiments.
Real-world door-opening and whiteboard-wiping experiments.
Video-aligned contact estimation and force prediction results.
Video-aligned HSR-Force contact estimation.
Force-aware compensation effect on contact-rich execution.
Effect of force-aware feedback and body compensation.

BibTeX

@misc{zhang2026fwbcvla,
  title        = {FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation},
  author       = {Zhang, Yutian and Ma, Siyuan and Yang, Liwen and Li, Yang and Hao, Ce and Chi, Haozhen and Wei, Dong and Yu, Qiaojun and Hou, Dibo},
  year         = {2026},
  eprint       = {2609.03889},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url          = {https://arxiv.org/abs/2609.03889}
}