Tackling Sim-to-Real Mismatch Through Sampling-Based Disturbance Observers: From Analytical Models to Learned World Models

Feedback compensation through state rollouts and cost queries

1Southeast University 2The Hong Kong University of Science and Technology (Guangzhou) 3Shanghai University of Electric Power

2 state and cost disturbance channels
5 simulation studies
3 real-robot demonstrations

Project video

SDOB at a glance

A visual introduction to disturbance observation through rollout and cost-query interfaces.

Overview

One observer principle, modern model interfaces

SDOB extends disturbance observation to analytical models, simulators, and learned predictors accessed through state rollouts or cost queries.

Abstract

Robotic controllers increasingly rely on analytical models, simulators, cost-query interfaces, and learned world models. However, physical deployment can deviate from nominal assumptions, and additional disturbances may arise even when the model itself is accurate. In control systems, disturbance observers (DOB) are widely used to estimate such unmeasured effects from nominal models and measured feedback. Classical DOB formulations are generally built around explicit plant models.

This paper develops the sampling-based disturbance observer (SDOB), extending the DOB principle to a broader range of models, including simulators and learned world models, through state-rollout or cost-query interfaces. SDOB separates two observable channels: state-effect disturbances, for which the feedback state differs from its prediction, and cost disturbances, for which the same query state receives different costs as the perceived environment changes. Diverse simulation and real-robot experiments across traditional and learned models demonstrate the effectiveness of SDOB in compensating for model mismatch and improving control performance.

Overview connecting classical control models, robot cost models, and world models to state-effect and cost disturbance channels in SDOB.
SDOB connects different model classes through state-rollout and cost-query interfaces. Each task supplies the correction model and queries needed to estimate its disturbance effects.

Contributions

Disturbance compensation through model queries

SDOB targets locally persistent, task-relevant mismatch through the interfaces used for prediction and planning.

  1. 1
    Two predictive-control interfaces

    We formulate disturbance compensation at two predictive-control interfaces. The state channel represents feedback mismatch through a task-defined correction model, while the cost channel represents local motion of an environment-dependent cost field.

  2. 2
    Forward-evaluation implementation

    A forward-evaluation implementation estimates these corrections from successive feedback and incorporates them into future state propagation and cost evaluation. Given the required correction and query interfaces, SDOB does not require analytic model inversion or retraining of the nominal model.

  3. 3
    Evaluation and reproducibility

    Five simulation studies and three representative real-robot demonstrations evaluate complete SDOB-equipped controller configurations across analytical and learned models. Source code, configurations, and simulation scripts are released to support reproducibility.

Two channels

What changed: the realized state or its evaluation?

The state channel fits feedback transitions within a task-defined disturbance model. The cost channel estimates local motion of an environment-dependent cost field.

01

State channel

State-effect disturbance

The feedback state differs from its nominal prediction. SDOB samples corrections within a task-defined disturbance model, evaluates them through forward prediction, and scores how well they explain the completed transition.

Same command Different state
02

Cost channel

Cost disturbance

The same query state receives a different cost after the perceived environment changes. SDOB matches local cost patterns across successive feedback updates to estimate cost-field motion for future rollout evaluation.

Same state Different cost

State-effect compensation

From feedback mismatch to corrected prediction

Candidate disturbances are propagated through the forward model and scored against the completed feedback transition. Regularization and physical admissibility suppress implausible explanations, and the resulting estimate corrects the next state rollout.

Cost transport

From fixed queries to a horizon-aware cost

Local query matches estimate how the observed cost pattern moved. Interpolation makes that estimate available at arbitrary rollout states, allowing the controller to evaluate where the high-cost region is expected to move.

Control integration

Estimate, compensate, plan

  1. 1
    Observe

    Compare completed model predictions and cost queries with feedback.

  2. 2
    Estimate

    Fit task-defined state corrections and local cost-field motion using forward evaluations.

  3. 3
    Compensate

    Correct subsequent state propagation and the costs used to score candidate rollouts.

SDOB computation flow from task feedback through state-effect and cost disturbance observers into compensated sampling-based control.
State estimates modify forward dynamics; cost estimates modify rollout evaluation. The sampling-based controller then selects the next input.

Simulation

Five simulation studies across model interfaces

The studies compare complete controller configurations across analytical and learned models using 20 paired trials or seeds. They assess task performance under the specified disturbances.

Analytical model · State effect

Sinusoidal torque disturbance

SDOB estimates a matched torque bias from the completed state transition and compensates the next MPPI rollout.

Success
20/20
Disturbance RMSE
0.014 N·m
P95 planning
8.232 ms

Metrics report MPPI+SDOB unless a baseline is named. These comparisons evaluate complete controller configurations; they do not separately isolate each channel or correction prior.

Hardware

Three real-robot demonstrations

Vehicle and dual-arm obstacle avoidance demonstrate feedback compensation in closed loop. PointWorld box lifting uses initial feedback calibration followed by execution of a single plan.

Real-car navigation

Obstacle avoidance from local cost changes

MPPI+SDOB completed the representative route in 9.29 s, the shortest time among the shown trials.

Dual-arm manipulation

Collision-cost observation without velocity tracking

The obstacle pose updates the collision model, but its velocity is not supplied. SDOB estimates motion from changing collision-cost queries and begins avoidance earlier.

PointWorld box motion predictions and real-robot lifting results without and with SDOB feedback calibration.

PointWorld · Initial feedback calibration

Box lifting with a frozen point-cloud predictor

An initial 50 mm lift provides feedback to fit a task-specific correction to PointWorld predictions. MPPI then plans a 300 mm lift using the calibrated predictor. The selected action sequence executes without replanning or observer updates.

Paired trials
20
MPPI final task RMSE
55.26 mm
MPPI+SDOB final task RMSE
22.61 mm