BECKMANN WORLD MODELS 01 · RESEARCH NOTES

Building a one-step
world model

From Beckmann transport to action-conditioned video: the objective, the failed trials, and what improved.

A world model is useful when it can answer a practical question: given what I have seen, what happens if I take these actions? If we want to compare many candidate actions, the cost of imagining each future matters.

Our starting point is a direct endpoint map: noise, recent observations, and actions go in; a future video chunk comes out in one network call. This post explains how we adapted Beckmann Transport Models to that setting—and why the transport objective alone was only the beginning.

PushT

SPATIAL EPOCH 30
Ground truthBWM + spatial
EXAMPLE

Robomimic Can

SPATIAL EPOCH 30
Ground truthBWM + spatial
EXAMPLE
Figure 1. Final spatial-epoch-30 predictions. Each pair shows recorded video (left) and BWM (right), with 4 context frames followed by 60 generated frames. All eight saved examples per task are available. Playback is 10 fps; it does not indicate inference speed.
In this article
  1. From denoising to direct maps
  2. The endpoint-map idea
  3. Making the map a world model
  4. Trials that changed the recipe
  5. What 30 spatial epochs changed
  6. Looking beyond the averages
  7. What one step buys
  8. Prediction is not planning
  9. Bridge-V2 and RT-1 in latent space
  10. What we would carry forward
  11. References & experiment record

01From denoising to direct maps

Diffusion models learn to reverse a noise-corruption process. Sampling follows a sequence of model evaluations from noise toward data. A world model repeats this process whenever it needs another future chunk, multiplying the cost across time and candidate action sequences.

Faster samplers and consistency models shorten this path. Consistency training can be standalone or distilled; a pretrained teacher is not a requirement of every few-step approach. The broader question is whether we can train a useful one-call generator directly.

Drifting models take that route by evolving the generated distribution during training, using attraction toward data and repulsion among generated samples. DriftWorld brings this idea to action-conditioned video. It is an important one-step comparison: our speed advantage is over iterative diffusion, not over other one-call models.

DiffusionIterative sampling
→future
Few-stepShorter sampling path
→→future
Direct mapsDrifting / Beckmann
one network call →future
Figure 2. Sampling cost, schematically. Drifting and Beckmann maps share a one-call interface but use different training objectives.

Beckmann transport gives us a different organizing principle: learn the destination associated with an autonomous trajectory. We use that principle to train a conditional map, then add supervision aimed at the endpoint a world model actually produces.

02The endpoint-map idea

Fix the context c=(h,a)c=(h,a), where hh is the observed history and aa is a proposed action sequence. For an autonomous flow, the terminal map is constant along a trajectory. Its differential form and data boundary condition are:

X˙τ=bc(Xτ),JxTc(x)bc(x)=0,Tc(y)=y.\dot X_\tau=b_c(X_\tau),\qquad J_xT_c(x)b_c(x)=0,\qquad T_c(y)=y.
(1)
Noise / sourceData manifoldyxT(x) = y
Figure 3. Move the point along a schematic characteristic: its destination stays fixed. Transport time τ\tau is separate from physical video time. The population construction in BTM motivates the map; it is not a convergence guarantee for a trained video network.

For RGB video, our map is a residual U-Net. It receives the future tensor being transported, four history frames, and action embeddings. The inference rule is simply:

y^=Tθ(z,h,a)=z+Uθ(z,h,a),z∼N(0,1.352I).\widehat y=T_\theta(z,h,a)=z+U_\theta(z,h,a),\quad z\sim\mathcal N(0,1.35^2I).
(2)

One call predicts four future frames on PushT or two on Robomimic Can, at 96 × 96 resolution. We append the prediction to the history and repeat. The network has no interpolation-time input; this is a time-independent endpoint map.

CONTEXT
First observed PushT frameSecond observed frameThird observed frameFourth observed frame
4 observed frames + actions
GAUSSIAN NOISE ↓Endpoint mapResidual U-Net · 1 call
FUTURE CHUNK
First predicted future frameSecond predicted frameThird predicted frameFourth predicted frame
Append to history; predict again
Figure 4. The deployed RGB model. Frames are saved PushT observations and predictions; arrows describe the interface, not an action-vector visualization.

What the training update actually computes

We sample a straight interpolation between Gaussian source zz and recorded future yy. A Jacobian–vector product (JVP) measures how the endpoint map changes along the source-to-target chord, holding history and actions fixed:

xs=(1−s)z+sy,p=Tθ(xs,c),q=JxTθ(xs,c)(y−z).x_s=(1-s)z+sy,\quad p=T_\theta(x_s,c),\quad q=J_xT_\theta(x_s,c)(y-z).
(3)
u=sg⁡(p+0.25q),ei=mean⁡d(pi−ui)2,Ltr=1B∑ieisg⁡(ei)+0.01.u=\operatorname{sg}(p+0.25q),\quad e_i=\operatorname{mean}_{d}(p_i-u_i)^2,\quad\mathcal L_{\rm tr}=\frac1B\sum_i\frac{e_i}{\operatorname{sg}(e_i)+0.01}.
(4)

sg⁡\operatorname{sg} means stop-gradient. The target is detached, so optimization regresses the map toward a stopped update; it does not backpropagate through a squared JVP using second derivatives. The chord is a sampling device, not the actual autonomous characteristic, and individual sampled JVPs need not be zero.

LBTM=λLtr+E∥Tθ(y,c)−y∥2+E[1s>0.8∥p−y∥2].\mathcal L_{\rm BTM}=\lambda\mathcal L_{\rm tr}+\mathbb E\|T_\theta(y,c)-y\|^2+\mathbb E[\mathbf1_{s>0.8}\|p-y\|^2].
(5)

The last two terms anchor the data boundary and the neighborhood near data. Squared norms here use coordinate means; the near-data term is averaged over the whole batch. A detached coefficient λ\lambda balances transport and anchor gradient norms on the last spatial output layer. Implementation ↗

A training-step sketch
# h and a stay fixed inside the forward-mode JVP.
x = (1 - s) * z + s * y
p, q = jvp(lambda x: T(x, h, a), x, y - z)
target = stop_gradient(p + 0.25 * q)
transport = normalized_residual(p, target)

# Supervise the exact query used at deployment.
prediction = T(z, h, a)
loss = balanced_transport_and_anchors(transport)
loss += source_loss(prediction, y)
loss += motion_temporal_range_losses(prediction, y)
loss += spatial_feature_loss(prediction, y)
# Short detached-history recovery runs on scheduled updates.

Conceptual pseudocode. Ramps, weighting, and recovery scheduling are described below; the linked research code is authoritative.

03Making the map a world model

Transport geometry describes a map over the future tensor. World modeling adds two demands: the prediction must correspond to the supplied actions, and the model must remain useful when its history contains earlier predictions. We addressed these through the training recipe.

1. Put supervision on the pure-noise query

p0=Tθ(z,h,a),Lsrc=E∥p0−y∥2.p_0=T_\theta(z,h,a),\qquad\mathcal L_{\rm src}=\mathbb E\|p_0-y\|^2.
(6)

A correct boundary value T(y,c)≈yT(y,c)\approx y does not directly supervise T(z,c)T(z,c), the query used at inference. Source loss does. Its coefficient ramps to one over 500 updates, outside the transport-loss balancing rule.

There is a tradeoff: squared error rewards the conditional mean and penalizes variation across source draws. At fixed context and target, the error decomposes into mean error plus sampling variance:

Ez∥T(z,c)−y∥2=∥EzT(z,c)−y∥2+Ez∥T(z,c)−EzT(z,c)∥2.\mathbb E_z\|T(z,c)-y\|^2=\|\mathbb E_zT(z,c)-y\|^2+\mathbb E_z\|T(z,c)-\mathbb E_zT(z,c)\|^2.
(7)

Lower error therefore does not prove that a generator preserves all plausible futures. Conditional diversity remains an open evaluation question.

2. Emphasize motion and tolerate imperfect history

Most pixels in a manipulation scene can stay almost unchanged. We add a target-motion-weighted loss, a within-chunk temporal-difference loss, and a penalty for predictions outside the normalized RGB range:

Laux=0.1Lmove+0.05Ltemporal+0.01Lrange.\mathcal L_{\rm aux}=0.1\mathcal L_{\rm move}+0.05\mathcal L_{\rm temporal}+0.01\mathcal L_{\rm range}.
(8)

The moving-region denominator has a floor so sparse motion does not amplify the loss excessively. Every fourth update also uses short generated-history recovery: earlier predictions are detached, and only the last prediction receives gradients. Its weight ramps to 0.25 over 1,000 updates.

The recorded future remains the training target, so this is a local robustness regularizer, not supervision for arbitrary counterfactual histories. We retain unclipped RGB feedback during evaluation. The working continuation also uses a quarter of the earlier learning rate; the observed gains belong to the complete recipe.

3. Compare spatial features at the output

Pixel error alone gives a limited description of object structure. We add a loss between spatial VGG16 feature maps of the source prediction and target, using relu1_2, relu2_2, and relu3_3:

Lsp=13∑ℓ=13Ef,u,v ⁣[∥ϕˉℓ(p0)f,u,v−ϕˉℓ(y)f,u,v∥22].\mathcal L_{\rm sp}=\frac13\sum_{\ell=1}^3\mathbb E_{f,u,v}\!\left[\left\|\bar\phi_\ell(p_0)_{f,u,v}-\bar\phi_\ell(y)_{f,u,v}\right\|_2^2\right].
(9)

Features are normalized across channels at each location, then channel-squared differences are summed and averaged across locations, frames, and examples. The feature extractor is frozen, but gradients flow through its input into the prediction. The weight ramps to 0.05 over 1,000 updates.

This is output supervision using pretrained features; it is not the learned LPIPS evaluation metric. VGG is absent at inference. The RGB generator is trained without a pretrained generative teacher.

PushT

PushT LPIPS: source 0.0746, robust 0.0564, spatial epoch 5 0.0407, spatial epoch 30 0.0280

Robomimic Can

Can LPIPS: source 0.0225, robust 0.0090, spatial epoch 5 0.0073, spatial epoch 30 0.0056
Figure 5. Recorded recipe endpoints. Labels include cumulative optimizer updates. Each continuation adds training as well as recipe changes; these bars are not isolated, equal-budget component ablations.

04Trials that changed the recipe

The most useful experiments were not all wins. Two branches tested whether adding a learned transport direction or propagating gradients through connected predictions would improve the existing map.

A learned direction did not transfer reliably

We added a head to predict the chord direction and gradually substituted it into the JVP. On Can, 64-frame MSE rose from 0.0068 to 0.0203 and LPIPS from 0.0225 to 0.0854. PushT improved in this continuation, so the evidence argues against adopting it as a reliable default—not against every learned-direction method.

The Can map was already deteriorating before the learned tangent was enabled. The auxiliary head shared decoder features and a clipping group with the endpoint model. That makes auxiliary-task interference a plausible explanation; the logs do not establish its share of the damage or rule out optimizer-restart effects. [audit]

More connected gradients were not automatically better

The attached variant backpropagated through two connected predictions at the full learning rate. The detached recovery recipe used a lower learning rate and performed better on the final 64-frame metrics for both tasks. Because both the gradient path and optimization recipe changed, this comparison does not isolate detachment alone.

Trial endpoints: source, direction, and recovery
Training recipePushT MSE ↓PushT LPIPS ↓PushT full ↓Can MSE ↓Can LPIPS ↓Can full ↓
Source-supervised map0.02930.07460.04340.00680.02250.0940
+ learned direction0.02080.05010.03190.02030.08540.0759
Robust · detached / lower LR0.01990.05640.02990.00310.00900.0112
Robust · attached / full LR0.02250.08080.03380.00510.02150.0146

Table 1. Completed trial endpoints, rounded as recorded. Full = full-episode autoregressive MSE. Direction continues from the source checkpoint; robust variants branch from that source parent. These are single training runs with different cumulative budgets. Source logs ↗

Can also exposed a gap between short and long rollouts. The source model’s 64-frame MSE was 0.0068, but full-episode MSE was 0.0940, with a worst-episode value of 0.737. The robust recipe reduced full-episode MSE to 0.0112. Good short-window prediction had not been enough to make repeated feedback reliable.

Spatial supervision subsequently improved 64-frame quality, but not every number moved together: Can’s full-episode MSE was 0.0112 for the robust model and 0.0114 after five spatial epochs. We kept the full-episode measurements visible while extending training.

05What 30 spatial epochs changed

The completed spatial continuations produce the strongest final RGB prediction metrics in our local comparison. Relative to spatial epoch five, 64-frame MSE falls from 0.0155 to 0.0098 on PushT and from 0.0028 to 0.0022 on Can. The improvements require no additional inference steps.

PushT

Recorded PushT MSE during 30 spatial epochs, falling from 0.0155 at epoch 5 to 0.0098 at epoch 30

Robomimic Can

Recorded Can MSE during 30 spatial epochs, falling from 0.0028 at epoch 5 to 0.0022 at epoch 30
Figure 6. Recorded 64-frame error during the spatial stage. Curves use log precision; final tables use saved evaluations. The primary result is the final checkpoint, not the best point on the curve.

“30 epochs” refers to the spatial stage, not the model’s whole training history. Including inherited training, the final models use 301,028 updates on PushT and 45,000 on Can. Update counts also do not equal compute: batch sizes and per-update costs can differ.

Reading the tables

64 frames = 4 observed + 60 generated; RGB is normalized to [−1, 1]. PushT uses 388 eligible 64-frame videos and 400 full episodes; Can uses 30 for both. “Moving MSE” emphasizes the moving region. Lower is better except SSIM. Local baseline runs use five epochs; BWM rows include inherited training.

PushT rollout results
Model / checkpointCalls / chunkMSE ↓Moving MSE ↓LPIPS ↓SSIM ↑Full MSE ↓
BWM + spatialspatial epoch 30 · final10.00980.12470.02800.9750.0154
BWM + spatialspatial epoch 510.01550.19330.04070.9640.0231
BWMrobust recipe10.01990.24190.05640.9530.0299
DriftWorld5-epoch baseline10.03140.40930.07900.9270.0451
MSE U-Net5-epoch baseline10.03030.40080.08040.9430.0427
GPC diffusion5-epoch baseline120.06301.42410.16090.9040.0753
AVDC diffusion5-epoch baseline1000.06841.18400.18600.8950.0795

Table 2. PushT, held-out rollout evaluation. Calls are per four-frame future chunk.

Robomimic Can rollout results
Model / checkpointCalls / chunkMSE ↓Moving MSE ↓LPIPS ↓SSIM ↑Full MSE ↓
BWM + spatialspatial epoch 30 · final10.00220.02420.00560.9840.0082
BWM + spatialspatial epoch 510.00280.02980.00730.9790.0114
BWMrobust recipe10.00310.03160.00900.9740.0112
DriftWorld5-epoch baseline10.01040.07690.06820.8560.0256
MSE U-Net5-epoch baseline10.00510.05440.01040.9690.0138
GPC diffusion5-epoch baseline60.02760.31030.05890.9100.0380
AVDC diffusion5-epoch baseline1000.04000.26440.11940.8010.1768

Table 3. Robomimic Can, held-out rollout evaluation. Calls are per two-frame future chunk. Tables 2–3 compare deployed checkpoints, not equal total training compute. Evaluation sources ↗

For context, a released long-trained DriftWorld checkpoint on PushT records MSE 0.0085 and LPIPS 0.0163 after about 1.18 million updates—better than our final values. Its exclusion of these evaluation episodes is unverified, so we show it as a reference rather than a controlled ranking. The evidence supports an improving one-step recipe, not a universal state-of-the-art claim.

06Looking beyond the averages

Averages hide where a model loses the scene. The viewer below synchronizes all methods on the same saved episode. It comes from the earlier spatial-epoch-five comparison; Figure 1 contains the new epoch-30 predictions.

01 / 64
Figure 7. Spatial-epoch-five comparison. Select a task, episode, or frame; every visible panel uses the same video clock. All eight examples are included. Can has no saved final-checkpoint AVDC clip. Videos are examples, not a replacement for whole-cohort evaluation.

Even at epoch 30, some PushT predictions become visibly blurred or diverge from the recorded block configuration later in the rollout. The examples below make those remaining errors easy to inspect. We cannot attribute their cause from appearance alone.

PushT · example 06

Ground truthBWM · epoch 30

PushT · example 08

Ground truthBWM · epoch 30
Figure 8. Two less successful saved examples, including all 64 frames. Lower average error does not eliminate individual rollout failures.

07What one step buys

On a dedicated A800, the spatial-epoch-five RGB model takes 20.6 ms per PushT chunk and 23.9 ms per Can chunk. Accounting for four versus two output frames gives 193.9 and 83.8 predicted frames per second, respectively.

PushT · four frames / call

PushT throughput: BWM spatial 193.9 frames per second, about 6.8 times GPC and 101 times AVDC

Can · two frames / call

Can throughput: BWM spatial 83.8 frames per second, about 2.9 times GPC and 99 times AVDC
Figure 9. Measured model-only throughput: batch one, FP32, TF32 disabled. September 26 timings at spatial epoch five; no new timing run is implied for epoch 30. Environment stepping and planning overhead are excluded.

The measured gains are approximately 6.8× / 2.9× over GPC and 101× / 99× over AVDC for PushT / Can. DriftWorld and the regression U-Net are already one-call models with similar latency. For those comparisons, the question is prediction quality and training behavior, not fewer sampling steps. [timings]

08Prediction is not planning

PushT planning uses a frozen pose reader to score the last predicted frame for candidate action sequences. The planner needs to order nearby alternatives correctly. That is a different requirement from reducing average error along recorded trajectories.

The final BWM model scores 0.721 on GPC-RANK, versus 0.709 for the local DriftWorld baseline. The paired difference is +0.012 with a 95% interval of [−0.047, +0.071]. This does not establish a planning improvement.

Paired planning differences versus DriftWorld; the BWM epoch 30 confidence interval crosses zero
Figure 10. Paired uncertainty across 50 fixed evaluation seeds. These intervals describe this evaluation, not variation across independent training seeds.
PushT planning and policy-outcome prediction
ModelPlanning ↑Policy r ↑IoU MAE ↓
BWM + spatialspatial epoch 300.7210.9280.010
BWM + spatialspatial epoch 50.7080.9460.010
DriftWorld5-epoch baseline0.7090.8570.031
MSE U-Net5-epoch baseline0.6950.9150.012
GPC diffusion5-epoch baseline0.6150.3660.513

Table 4. Planning and policy-outcome prediction. Policy estimates use seven policies × 300 paired trials. MAE is the absolute error in predicted policy IoU; it is not a robot success rate. Values are rounded. Evaluation record ↗

Do the actions actually matter?

In a 64-window probe, swapping actions while holding real history fixed increases the spatial model’s moving-region error by 11.0× on PushT and 6.2× on Can. The predictions respond to the action input, rather than relying only on visual history.

PushT

PushT moving-region error with correct versus swapped actions

Robomimic Can

Can moving-region error with correct versus swapped actions
Figure 11. One-chunk action probe at spatial epoch five. Error is measured against the recorded future. The swapped action sequence has no corresponding ground-truth counterfactual here, so this tests sensitivity, not counterfactual correctness.

A useful next target is candidate-ranking error: whether the model prefers the action that really produces the better outcome. Neither a sharp frame nor a large action-swap penalty answers that question on its own.

09Bridge-V2 and RT-1 in latent space

We also transferred the recipe to 256-pixel robot video. The native model uses SD3 VAE latents with 16 channels at 32 × 32 resolution and a DriftWorld-style native U-Net: 175.4M parameters for Bridge-V2 and 160.3M for RT-1.

Each call predicts one latent frame, conditioned on four context frames and actions. The spatial and range losses operate through the decoded endpoint; the temporal term is zero for a one-frame output. The pretrained VAE remains part of this system.

4-frame context→SD3 latents→One-step map→Decoded frame

The completed subset runs use 2,078 eligible training episodes on Bridge and 2,041 on RT-1. Unlike the RGB continuations, these runs schedule map, source, robustness, and spatial stages within a single 30-epoch subset budget.

Perceptual error

LPIPS decreases across epochs 10, 20 and 30 on both native subsets

Structural similarity

SSIM increases across epochs 10, 20 and 30 on both native subsets
Figure 12. Completed native subset evaluations. Training stages change over these epochs, so the curves combine recipe progression and additional training.
Native video: completed subset runs, epoch 30
DatasetSSIM ↑PSNR ↑LPIPS ↓FID ↓FVD ↓Calls / frame
Bridge-V20.79622.130.15568.86592.241
RT-10.80223.150.16675.53535.591

Table 5. Final epoch-30 subset results. Bridge: 215 validation videos / 7,079 predicted frames. RT-1: 198 / 8,409. FVD uses 201 / 193 eligible videos and their first 16 generated frames.

Generator time is 16.3 ms per Bridge frame and 13.1 ms per RT-1 frame; VAE decoding adds about 26 ms per frame. The latest run table also records approximately 2.8 and 2.3 training GPU-hours for the subset runs, excluding evaluation.

At the September 29, 04:16 UTC source snapshot, full-data continuations were still training and matched-subset baselines were queued. There are no completed full-data results to report. Snapshot ↗ · Latest run table ↗

10What we would carry forward

The practical contribution is a direct conditional map whose output is supervised where it is used: at the Gaussian source, under imperfect history, and in spatial feature space. The strongest evidence is the completed improvement in video prediction while retaining one-call generation.

  1. Keep the inference query central. A transport-inspired objective still needs to train the operation deployed in the world model.
  2. Audit the entire training lineage. An epoch label for the last stage is not a total training budget.
  3. Keep the unsuccessful trials. A learned direction and more connected gradients did not reliably beat the simpler working recipe.
  4. Evaluate the downstream decision separately. Planning, conditional diversity, and counterfactual accuracy remain distinct from image quality.

The immediate open questions are whether these gains persist in matched native comparisons, whether useful conditional variation survives source supervision, and whether better endpoint structure improves action selection. The current measurements leave all three questions open.

References & experiment record

The September 25 manuscript predates the new results. This article uses saved evaluations and the September 29, 04:16 UTC repository snapshot for the updated experiment record.

  1. Denoising Diffusion Probabilistic Models. Background on iterative denoising.
  2. Consistency Models. One- and few-step generation, including standalone training.
  3. Generative Modeling via Drifting. Direct one-step generative training.
  4. DriftWorld: Fast World Modeling through Drifting. The one-step world-model comparison and native interfaces.
  5. Beckmann Transport Models: From Autonomous Flows to One-Step Maps. Terminal-map geometry and stopped-target training; our work adapts this foundation to world modeling.
  6. Trial records: Can source / direction, PushT source / direction, robust and spatial continuations.
  7. Direct-map causal audit and mathematical corrections. Mechanism hypotheses, gradient paths, and limits of the probes.
  8. RGB results: exact final metrics and checkpoint hashes, earlier baseline snapshot, spatial training record, comparison-video provenance.
  9. Dedicated inference-latency measurements. Hardware, precision, and model-call accounting.
  10. Downstream evaluations: policy evaluation, epoch-30 planning and policy results.
  11. Native Bridge-V2 / RT-1 snapshot. Completed subset evaluations, timing, and pending comparisons.

Beckmann World Models
Adam Lee · Shobhit Agarwal · JP Suh